Research, experiments, and reflections by Peter Sauer.
-
The Weights Kept Nothing. So I Built a Measuring Machine
I closed the last post with three fixes that would help more than any training run. One shipped, two died, and the week produced something better than all three.
-
The Weights Kept Nothing
I trained a small language model on 97,437 chunks of my own notes. It answered zero of 131 questions correctly. The real fix turned out to be a search index bug.
-
The Half-Life of a Loop
An unsupervised loop does not fail at some length. It decays on a fixed schedule, and the schedule has a name.
-
The Generator Was the Wrong Instinct
A slide generator that had worked for a year broke the moment I pointed it at an editable file, and why editability quietly turns a quality problem into a verification problem.
-
Proof Is a Technology
Babylon verified by application, the Greeks verified the inference. 2,500 years later, that fork decides how much you can trust an unsupervised AI agent.
-
Notation Is Compute
Multiplication in Roman numerals was expert work; in place-value it's a child's procedure. What Brahmagupta and al-Khwarizmi knew about spec design.
-
When Physics Was the Test Suite
Calculus ran 150 years on undefined foundations and nothing collapsed. The reason is the oldest test-coverage story in engineering.
-
The Most Productive Failure
Incompleteness killed Hilbert's program and somehow got blamed for far more. The limits fall on finding proofs, never on checking them, and that asymmetry is what carries Lean and Coq today.
-
The Untrusted Searcher
Twelve exhausted referees, a Lean kernel, and why an agent's autonomy budget scales with the hardness of its verifier.
-
Three Ledgers: How to Value AI Coding When the Metrics Disappoint
Uber capped its AI budget, METR lost its control group. Both stories say less about AI productivity than about how we account for value.