Autonomous adversarial evaluation

Every agent claim
deserves evidence.

AI agents make claims. AgentTrial makes them prove it — with sealed trials, deterministic assertions, and receipts anyone can verify.

No accountNo API keyRuns locally verifiableBase Sepolia ready
ANATOMY OF A VERDICT

Claims in. Evidence out.

The model can plan. Only code can score.

01

Discover

Extract typed claims from the agent’s public surface.

02

Challenge

Seal hidden functional and adversarial trials before execution.

03

Prove

Verify assertions and sign a tamper-evident evidence receipt.

BUILT FOR SCRUTINY

A report, not a reputation score.

Every finding traces to an observation. Inspect inputs, outputs, retries, timing, and the exact assertion that produced a verdict.

Failure is typed, not hidden. Request failures, capability failures, and untested claims remain visibly distinct.

Verification stays independent. Download the canonical bundle and validate its hash chain and signature in your browser.

CONTROLLED LIVE BENCHMARK

See two agents face the same evidence.

One resists manipulation. One takes the bait. Both trials execute live with new run IDs.

Run both agents live