BENCHMARKS / PUBLIC EVIDENCE
Scores, status, and evidence.
These results are self-reported, and each row states what evidence is public. The GAIA row is marked while its scrubbed trace capsule is pending. Open upstream submissions stay visible, and publication does not wait for a leaderboard maintainer to merge them.
| Benchmark | Model | Ouroboros | Comparison | Status and evidence |
|---|---|---|---|---|
| Terminal-Bench 2.1 | Claude Opus-5 high | 86.74% after zeroing one disclosed reward-hack trial (raw: 86.97%) | Claude Code + Fable 5: 83.8% | Self-reported with an open submission and a public run. |
| Terminal-Bench 2.1 | Claude Opus-4.8 high | 80.22% | Claude Code: 78.9% | Self-reported with a public run. |
| Terminal-Bench 2.1 | GPT-5.5 | 84.3% | Codex CLI: 83.1% | Self-reported with a public run. |
| Terminal-Bench 2.1 | Grok-4.5 | 84.94% | Cursor CLI: 79.3%. Hermes: 77.53%. | Self-reported after a reward-hack audit, with an open submission. |
| OSWorld-Verified | Claude Opus-5 | 90.69% | Previous best on the public board: 90.19% | Self-reported with public task evidence. |
| OSWorld-Verified | Claude Sonnet-4.6 | 83.27% | Pointer: 81.45% | Self-reported with public task evidence. |
| CL-Bench | Claude Sonnet-4.6 | 0.2301, rank 1 | Previous top: 0.1960 | Self-reported with an open submission and full traces. |
| SWE-bench Pro | GPT-5.6-luna | 58.2% | Codex CLI scored 59.4%, with no significant difference. | Self-reported with matched-pair traces. |
| GAIA | Claude Sonnet-5 | 129/165, 78.2% | Claude Code scored 131/165, 79.4%, while strict pass@1 was 128/165 for both. | Self-reported with public methodology; the scrubbed trace capsule is pending. |
How to read the table
Each comparison binds a model to a harness. A model score from one harness does not describe the same system as the score from another harness. The repository keeps adapters and methods under devtools/benchmarks.