I check whether AI agents' "done" is true, against records they can't change.
- The method: the Sonny Test · DOI 10.5281/zenodo.23117475
- Studies: swarm-receipts · receipt-desk
- Sites: iswt.ca · jdbauer.ca
- Contact: joshua@jdbauer.ca
I check whether AI agents' "done" is true, against records they can't change.
A status desk that won't take an AI agent's word for "done": done, failed, or not shown, with the exact line it copied from the record.
HTML 1
Forked from braintrustdata/autoevals
AutoEvals is a tool for quickly and easily evaluating AI model outputs using best practices.
Python 1
The Sonny Test: a check counts only if its verdict comes from a record the AI can't change, and it has already caught a fault planted on purpose. With its sealed record.
A planted-fault test bench for swarm oversight tools, and the sealed audit in which it caught two baseline checkers (AI Swarm Dynamics Hackathon, Oct 2026)
Python 1
Two small AI models, run on your own PC by AMD's Lemonade, check whether an AI agent's "done" is true against its own log.
Python