telepace

Evidence to evals

Teach your AI
what correct means.

Telepace turns production failures, expert judgment, policies, and user evidence into eval cases, calibrated judges, and release gates.

Start with one costly failure. No survey-scale dataset required.

candidate B · release check
Hold

Release score

42/100

+9 vs production

Critical blocker

CASE-042

The agent promises a refund before verifying policy eligibility.

1 critical failure · ambiguous refund slice · zero tolerance

source
If eligibility is unclear, do not promise. Escalate with the order context.Refund policy v12 · owner approved
Any critical policy breach blocks releaseHOLD

enough to start

1 failure

to a first eval contract

< 3 min

where releases happen

JSON · MCP · CI

How it works

From one bad outcome to a permanent release gate.

01

Define the task contract

Name the release decision, expected outcome, prohibited outcomes, tools, and critical slices.

02

Map the evidence gaps

Telepace shows what is unknown, who can resolve it, and why each question is worth asking.

03

Collect only what is missing

Attach traces, policies, tickets, expert corrections, or a short user clarification to the same evidence graph.

04

Compile and gate

Export replayable Eval Cases, rubric anchors, layered graders, and a ship or hold decision.

The compounding artifact

Evidence that can protect the next release.

A quote is not the finish line. Telepace compiles accepted evidence into a versioned task contract, regression case, scoring rubric, judge plan, and release gate.

Refund policy regression

1 production failure · policy v12 · owner calibrated

Eval Pack v1

Real evidence

The bot said the refund was approved, then support told me I was not eligible.
Support ticket #1842 · linked to production trace · verified

trace_8f21 · turn 14 · policy_refund_v12

Compiled output

01

Task contract

Verify eligibility before committing

02

Critical Eval Case

Ambiguous refund must escalate

03

Release gate

0 policy breaches allowed

telepace.eval-pack.v1 · provenance preserved

One evidence graph

Use the source with authority, not the source with volume.

Production traces

what actually happened

Support tickets

costly failures in context

Domain experts

quality bars and boundaries

Policies and specs

deterministic hard gates

User clarification

only when intent is missing

Agent-native

Your coding agent can build. Telepace tells it what must never break.

Use MCP, Skill, REST, or versioned JSON to create evals from production evidence and carry the exact release gate into your existing runner and CI.

// Codexcodex> compile this refund failure into an eval // telepace.create_campaign✓ campaign_id: 4f2b…9c1✓ task_contract: refund eligibility // telepace.get_eval_pack✓ 3 candidate cases✓ 1 critical policy gate✓ judge order: deterministic → model → human // evidence compiled  · trace linked to policy v12  · ambiguous request becomes Eval Case  · zero-tolerance release gate preserved

A clear budget owner

Built for teams accountable for AI behavior.

AI Product Leads

Make a defensible ship, hold, or rollback decision.

Eval Engineers

Turn real failures into durable regression coverage.

Domain Experts

Encode judgment once, then measure judge agreement.

Risk and Operations

Make policy boundaries observable and release-blocking.

Do not ship
undefined correctness.

Start with the failure your team cannot afford to repeat.

telepace — From user evidence to production evals · telepace