04 / MACHINES · FIELD NOTES
Automating Proof Doesn’t Automate Curiosity
KNOWING WHICH QUESTION
COMES NEXT.
STARHOUND STUDIO · 2026
AI is getting very good at answering questions.
That is not the same thing as knowing which question should come next.
There is a tendency to talk about mathematical and scientific automation as though the difficult part of research were gradually being reduced to computation: prove the theorem, solve the equation, search the parameter space, verify the result.
Those things matter.
But sometimes the most consequential moment in an experiment isn't when the system gives you the answer you expected.
It's when something tiny refuses to behave.
The 0.0002547% Problem
I have spent a considerable amount of time experimenting with a geometric computational system.
Early in that work, I was studying transitions in its semantic geometry and trying to understand what actually carried them.
The obvious place to look was the dominant structure.
And the dominant structure was very dominant.
In one experimental cell, the first principal component captured 99.7% of the variance.
If variance were a reasonable proxy for computational importance, the problem should have been essentially solved.
It wasn't.
Using that dominant component alone recovered zero of the two actual transitions and achieved only 62.5% agreement with the true boundary signs.
Then I put back a residual component.
It accounted for approximately
0.0002547%
of the total variance.
Boundary-sign agreement jumped from 62.5% to 87.5%.
Adding another small component recovered both true transitions.
The overwhelming majority of what was moving was not necessarily what was deciding.
That became one of the central findings of the research:
Variance importance and decision importance are different quantities.
But there is another lesson hiding inside that experiment.
The computer did not look at 99.7% and become uncomfortable.
I did.
The Question Wasn't in the Original Experiment
The interesting question became:
What is hiding in the residual?
And that question only existed because an expectation had failed.
There is an important distinction here.
A system can be exceptionally good at answering:
Given this question, what is the answer?
Scientific curiosity requires something different:
Given what I just observed, what question now deserves to exist?
And experimental science adds an even nastier requirement:
What experiment could make the explanation I currently prefer lose?
Those are not interchangeable abilities.
My experiments eventually became a chain of more than a hundred rounds. They did not proceed because I knew the destination in advance. They proceeded because each result constrained what question was intellectually defensible next.
A global explanation failed.
So I tested whether the effect was local.
A promising predictor appeared.
So I tried to manipulate it causally.
The manipulation failed.
A regime-like discontinuity appeared.
So I attacked the possibility that it was merely a normalization artifact.
A boundary effect appeared.
So I moved the boundary away from the endpoint to see whether the effect survived.
An estimator produced NaNs, numerical failures.
Instead of throwing them away as garbage, I asked what kind of failure they represented.
Again and again, the useful move wasn't simply obtaining another answer.
It was deciding what the answer had made suspicious.
That pattern became part of the experimental method itself: anomalies and reversals generated hypotheses, but they weren't allowed to become findings until subsequent experiments survived attempts to destroy them.
Failure Is Information, If You Let It Be
There is a version of scientific automation that worries me.
Not because the machines become too intelligent.
Because the workflow becomes too efficient.
Imagine an AI analyzing an experiment.
It produces a beautiful statistical report. It tests the registered hypothesis. It checks the assumptions. It calculates the confidence intervals. It writes the conclusion.
Everything is correct.
And we move on.
But what if the scientifically valuable observation was the ugly little thing sitting three columns over?
The subgroup that behaved backward.
The residual with almost no variance.
The control that nearly reproduced the treatment.
The estimator that failed only after crossing one particular geometric boundary.
The null result that quietly destroys the story everyone wanted to tell.
Those aren't necessarily errors to eliminate.
Sometimes they are invitations.
And accepting the invitation requires something beyond optimization toward the requested answer.
It requires the ability to become interested in the wrong thing.
Proof Is Not the Whole of Mathematics
Proof matters enormously.
But mathematics isn't a queue of propositions waiting to be processed.
Someone still has to decide which proposition matters, which restriction looks suspicious, which failed proof revealed something more interesting than the theorem it was trying to establish.
A machine capable of proving every proposition handed to it would be extraordinary.
But there are infinitely many propositions you could hand it.
The scarce resource may not be the proof.
It may be the question.
Curiosity Is Not Random Question Generation
Generating follow-up questions isn't curiosity. That's easy.
A consequential scientific question exists in relation to everything that has already survived. It recognizes which explanations have been eliminated, which controls have already been run, which observations are genuinely anomalous, and which remaining uncertainty deserves another experiment.
Curiosity without memory becomes novelty.
Curiosity without skepticism becomes confirmation bias.
Curiosity without experimental discipline becomes storytelling.
The interesting target isn't an AI that constantly asks more questions.
It's a system capable of recognizing when its model of the world has become inadequate, and constructing an experiment capable of proving its preferred explanation wrong.
That is much closer to science.
And none of this requires believing that machines can never do it.
Quite the opposite.
If we want machines that participate meaningfully in scientific discovery, then this is precisely the capability worth trying to build.
The problem is assuming that increasingly powerful answer generation, reasoning, or automated proof means we have already built it.
We may simply be measuring the wrong thing.
The Difference Between Knowing and Wanting to Know
We have become extraordinarily focused on making machines that know more.
Larger corpora. Longer contexts. Better retrieval. Better reasoning. Better proofs. Better answers.
All worthwhile.
But intelligence has another strange property that doesn't fit quite as neatly onto a benchmark:
It notices that something doesn't make sense.
Then it refuses to leave it alone.
The 99.7% result could have been the end of an experiment.
Instead, 0.0002547% became the interesting part.
That tiny residual helped redirect the investigation away from bulk variance and toward the geometry actually associated with decisions, a distinction that continued through the later experimental program.
That is why automating proof doesn't settle the future of mathematics.
And automating reasoning doesn't settle the future of science.
Because after the machine gives us the answer, there remains another problem:
Knowing when the answer should bother us.
Knowing which tiny contradiction deserves attention.
Knowing which beautiful explanation deserves to be attacked.
And knowing when to look at the thing everyone else has dismissed as noise and ask:
Why the hell did that happen?
— Starhound Studio