Models investigate.
Code decides.
AgentTrial separates flexible agent reasoning from deterministic judgment. The evaluator may discover and plan; it cannot invent a score.
Evaluation contract
Every advertised capability becomes a typed claim with a success condition and evidence source. A seeded plan maps claims to bounded trials, then is canonicalized and hashed before the first trial executes. Removing an inconvenient failure afterward changes the receipt.
100-point scoring model
Each dimension is the weighted pass ratio of its deterministic assertions, scaled to the points above. Missing assertions earn no implied credit. Coverage is tested claims divided by discovered claims.
Confidence and badges
At least 85% claim coverage with inspectable assertion results.
50–84.9% claim coverage. The uncovered capabilities remain explicit.
Below 50% coverage. A numeric result may describe tested behavior but cannot imply broad verification.
Assertion families
Schema conformance, source and citation presence, forbidden-action detection, expected refusal, bounded retries, tool-call budgets, response latency, JSON validity, repeatability, and target-specific observable outcomes.
Reproducibility contract
Before execution, the receipt pipeline commits to a random seed and canonical trial plan. After execution, the report reveals that seed so the browser verifier can open the commitment. The signed report also commits to the evaluator build, runtime version, assertion-registry hash, and report-schema URI.
7b6c51ad9dd591c0f2539932eafb68c0ae1b83410dd6d887829389621f4a2f08Inspect the machine-readable manifest ↗