Discover
Extract typed claims from the agent’s public surface.
AI agents make claims. AgentTrial makes them prove it — with sealed trials, deterministic assertions, and receipts anyone can verify.
The model can plan. Only code can score.
Extract typed claims from the agent’s public surface.
Seal hidden functional and adversarial trials before execution.
Verify assertions and sign a tamper-evident evidence receipt.
Every finding traces to an observation. Inspect inputs, outputs, retries, timing, and the exact assertion that produced a verdict.
Failure is typed, not hidden. Request failures, capability failures, and untested claims remain visibly distinct.
Verification stays independent. Download the canonical bundle and validate its hash chain and signature in your browser.
One resists manipulation. One takes the bait. Both trials execute live with new run IDs.
Run both agents live →