Ouroboros

BENCHMARKS / PUBLIC EVIDENCE

Scores, status, and evidence.

These results are self-reported, and each row states what evidence is public. The GAIA row is marked while its scrubbed trace capsule is pending. Open upstream submissions stay visible, and publication does not wait for a leaderboard maintainer to merge them.

BenchmarkModelOuroborosComparisonStatus and evidence
Terminal-Bench 2.1Claude Opus-5 high86.74% after zeroing one disclosed reward-hack trial (raw: 86.97%)Claude Code + Fable 5: 83.8%Self-reported with an open submission and a public run.
Terminal-Bench 2.1Claude Opus-4.8 high80.22%Claude Code: 78.9%Self-reported with a public run.
Terminal-Bench 2.1GPT-5.584.3%Codex CLI: 83.1%Self-reported with a public run.
Terminal-Bench 2.1Grok-4.584.94%Cursor CLI: 79.3%. Hermes: 77.53%.Self-reported after a reward-hack audit, with an open submission.
OSWorld-VerifiedClaude Opus-590.69%Previous best on the public board: 90.19%Self-reported with public task evidence.
OSWorld-VerifiedClaude Sonnet-4.683.27%Pointer: 81.45%Self-reported with public task evidence.
CL-BenchClaude Sonnet-4.60.2301, rank 1Previous top: 0.1960Self-reported with an open submission and full traces.
SWE-bench ProGPT-5.6-luna58.2%Codex CLI scored 59.4%, with no significant difference.Self-reported with matched-pair traces.
GAIAClaude Sonnet-5129/165, 78.2%Claude Code scored 131/165, 79.4%, while strict pass@1 was 128/165 for both.Self-reported with public methodology; the scrubbed trace capsule is pending.

How to read the table

Each comparison binds a model to a harness. A model score from one harness does not describe the same system as the score from another harness. The repository keeps adapters and methods under devtools/benchmarks.