benchmarks · the runs
Benchmarks
Model evaluations run by the lab. Every run ships with the method, the limitations, the raw responses, and a command that recomputes the scores offline.
-
2026-09-03
AdventureBench
Updated 2026-09-16: AdventureBench now charts score against both cost and per-case latency across twelve language models, with raw evidence and offline score replay. · repository
-
2026-08-12
Regexes you could actually ship.
regexbench across 11 models on Re(gEx|DoS)Eval: does the pattern pass its tests, mean what the reference means, and survive a ReDoS screen. Reported in three groups by what a paired bootstrap resolves — not a ranking. · article · repository