research · the notebook
Research
Articles and analyses from the lab, with the whitepapers behind them. Articles link to methods and evidence where available. Each study states what can be reproduced and where records or checks are missing.
-
2026-10-02
Benchmarks are living things
A benchmark can keep returning scores long after it stops telling us much. I've been thinking about how to recognize that point, and what to do when the question we need it to answer changes. · pdf
-
2026-10-02
What I learned rebuilding Pi Audit Loop
A review and simplification loop found real bugs, but it also confused findings with repairs. I rebuilt it around one change, a checked revision and a final diff review. · pdf
-
2026-09-23
jev: three use cases, a ton of learnings
We spent a week pointing Jev at a Doom deathmatch, 370 enterprise documents, five text adventures and a laptop reimplementation. · pdf
-
2026-09-17
A model that answers in probabilities
TypeSafe's Jev returns typed, calibrated judgments instead of generating text. We ran its separately audited v2 runtime on AdventureBench: incredibly accurate, fast, and revealing on terse commands. · pdf
-
2026-09-11
The harness is the signal
As models improve, some of today's agent machinery will disappear. The parts that govern permissions, execution, recovery and accountability probably will not. · pdf
-
2026-08-30
What agent skills are actually made of
1.9 million agent skills on GitHub. Only 12 in 100 come with any code attached, and when they do it is usually Python, whatever the skill is about. · pdf
-
2026-08-27
A safety filter moved our benchmark by seven places
A provider content filter silently deleted 29.3% of one model's answers on our regex benchmark, and the blanked calls were not missing at random. Depending on how a harness handles them, the same model finishes 3rd, 10th or 4th out of 11. Running k=3 rather than a single sample is what caught it. · pdf
-
2026-08-24
What language are agent skills written in?
3.8 million AI agent instruction files on GitHub, read by a model that speaks every language. So why are 85.3% of them in English, and what is changing? · pdf
-
2026-08-21
a regex can pass its tests and still fail in use
A study of generated regular expressions, benchmark answer keys, and vulnerability screening across corpora. Passing tests is a limited guarantee; corpus differences do not establish what caused them. · pdf