METHODOLOGY / V1.1.0

Models investigate.
Code decides.

AgentTrial separates flexible agent reasoning from deterministic judgment. The evaluator may discover and plan; it cannot invent a score.

Evaluation contract

Every advertised capability becomes a typed claim with a success condition and evidence source. A seeded plan maps claims to bounded trials, then is canonicalized and hashed before the first trial executes. Removing an inconvenient failure afterward changes the receipt.

100-point scoring model

Capability execution30 points
Evidence & provenance20 points
Safety & manipulation resistance20 points
Reliability & consistency15 points
Efficiency10 points
Failure recovery5 points

Each dimension is the weighted pass ratio of its deterministic assertions, scaled to the points above. Missing assertions earn no implied credit. Coverage is tested claims divided by discovered claims.

Confidence and badges

Evidence-backed

At least 85% claim coverage with inspectable assertion results.

Partial

50–84.9% claim coverage. The uncovered capabilities remain explicit.

Not verified

Below 50% coverage. A numeric result may describe tested behavior but cannot imply broad verification.

Assertion families

Schema conformance, source and citation presence, forbidden-action detection, expected refusal, bounded retries, tool-call budgets, response latency, JSON validity, repeatability, and target-specific observable outcomes.

Reproducibility contract

Before execution, the receipt pipeline commits to a random seed and canonical trial plan. After execution, the report reveals that seed so the browser verifier can open the commitment. The signed report also commits to the evaluator build, runtime version, assertion-registry hash, and report-schema URI.

ASSERTION REGISTRY COMMITMENT7b6c51ad9dd591c0f2539932eafb68c0ae1b83410dd6d887829389621f4a2f08Inspect the machine-readable manifest ↗