Your benchmark is only as trustworthy as its verifier.
Tbench stress-tests the task, environment, reference solution and verifier before they enter training, evaluation or a marketplace. Start with one task or send the full set.
- 14-stage workflowSource lock through evidence seal.
- 47 registered rulesThe verifier is tested, not merely run.
- One task or a setThe unit changes; the standard does not.

- 01
- Published method
- 02
- Auditable record
- 03
- Clear refusal
A method. A record. A defensible result.
What the work is held to
Evidence has a standard.
Tbench is not a score overlay. It is a record of the method applied, the checks that ran, and the boundary between a result that can be defended and one that cannot.
A published methodology
Every check is named, weighted as required or advisory, and versioned. The result records the revision it was graded under.
A public identifier
Every completed audit carries an ID anyone can look up. A buyer does not have to take the vendor's word that a submission was audited.
A real expiry
The result lapses. A submission measured against a methodology superseded twice since is not the same claim it was on the day it was issued.
- stages from source lock to evidence seal
- 14
- registered static verifier rules
- 47
- explicit verdict states
- 5
- methodology revision, fixed per audit
- v1.0
The language of an audit
Every check returns one of five verdicts. Two of them mean nothing was verified.
- PassedThe check ran and the work satisfied it.
- FailedThe check ran and the work did not satisfy it.
- Needs reviewThe check could not decide. A person has to look.
- SkippedThe check did not run -- its toolchain is not installed. Nothing was verified, so this is not a failure.
- ErroredThe check crashed before it could decide. That says nothing about the work, so this is not a failure.
The process
One flow, from PR to evidence.
The customer journey is three decisions; the assurance engine underneath is fourteen stages. Model trials, toolchain availability and reviewer inputs remain explicit in the record.
You bind the source
Paste a pull request, repository, or archive and choose one task or the complete set. The audit record is bound to that source; Tbench grades what you submitted, not a version it produced.
- The submitted source is treated as read-only; remediation remains explicit
- A one-task audit is a first-class submission, not a minimum-size exception
- A submission ID is issued before anything runs
It runs against the methodology
Tbench Certified rests on a published fourteen-stage workflow: source lock, twelve task-level validation stages and one evidence seal. Forty-seven registered rules sit inside the static audit, while container and model stages retain their own evidence.
- Required checks decide the outcome; advisory ones are reported
- A check that did not run is never counted as a pass
- When configured, Megatron adds a second opinion; an unavailable review is recorded
You receive a result or a remedy
All required checks must run. A clean pass earns the full mark; required human review earns a conditional mark; a failed or unrun required check withholds the mark. Every outcome keeps the per-check evidence and names what happens next.
- Valid for 12 months, with the methodology revision on the result
- A refusal names the blocking checks, keyed so they can be fixed and re-submitted
- A later submission is measured against the methodology revision applied to that audit
The most useful thing we do is say no.
Most grading pipelines report a pass rate. A pass rate over a set where a quarter of the checks never ran flatters whoever submitted it, and it is how unverified data enters a training run with a green tick against it.
If even one required check cannot run, the audit withholds the mark. Checks that did run remain visible in the report, alongside the blocker and the evidence needed for a future submission.
- A check that did not run scores zero, never absent.
- Coverage of required checks has to be total before any mark is issued.
- A failed check and a check that never ran are reported as different things.
Required coverage incomplete
The report records what ran and what did not.
- Required coverage
- Incomplete
- Blocking threshold
- 1+
- Methodology
- v1.0
A required check that failed, errored, or did not run cannot be averaged away by the checks that passed. The audit keeps the evidence and withholds the mark.
Methodology rule — no fictional customer or certificate data.
Who sends us work
Three people, one decision each.
Frontier labs
Whether the benchmark data going into a run has been checked by someone other than the people who wrote it.
Data vendors
Whether a task batch can clear a buyer's procurement. Unverified verifier behavior has to be argued for; a defensible evidence bundle does not.
Eval and platform teams
Whether a benchmark number compares two systems on comparable work, or compares two different standards.
Know whether your verifier deserves to be trusted.
You get the report either way. A clean pass earns Tbench Certified, required human review produces a conditional mark, and a failed or unrun required check withholds the mark with the blocking evidence attached.
Approved organisations receive starter task credit. After that, the live per-task price is shown before a submission is admitted.
Tbench AI stress-tests agent tasks and their verifiers against a published methodology, built on the Infinity toolkit, with Megatron as an advisory reviewer.