02 / MACHINES · FIELD NOTES

Variance Is Not Importance

99.7% of the variance.
Not the computation.

A destructive experiment in what actually carries information through geometric cognition.

STARHOUND STUDIO · 2026

99.7% of the variance was sitting in one principal component.

It should have been the obvious place to look.

It wasn't.

That dominant component produced only 62.5% boundary-sign agreement.

Then I restored another component—one accounting for approximately:

0.0002547% of the total variance.

Boundary-sign agreement jumped to:

87.5%.

That was annoying.

Which, experimentally, is usually where things become interesting.

I've spent an unreasonable amount of time lately trying to break a geometric computational system.

Not benchmark it. Not find a flattering test. Break it.

I've been removing mechanisms, destroying structure, changing evaluation conditions, generating fresh populations, and repeatedly asking:

What actually carries the computation?

This particular result forced an uncomfortable distinction.

The biggest direction isn't necessarily the important direction

Variance tells us where the data spreads.

But I wasn't asking where the representation was largest.

I was asking which geometric structure mattered to the decision.

Those aren't necessarily the same thing.

Imagine an enormous valley dominating a landscape. Nearly every measurement of the terrain is explained by it. Decompose the landscape statistically and the valley screams at you.

Now imagine that whether you end up on one side of a border or another depends on an almost invisible ridge cutting across it.

The ridge contributes essentially nothing to the overall scale of the landscape.

Destroy it, though, and you cross the wrong boundary.

Magnitude is not leverage.

And variance is not automatically importance.

So I tried to kill it

A strange result isn't valuable merely because it's strange.

It becomes interesting when it survives attempts to make it disappear.

So I didn't immediately ask how to explain it.

I asked:

Can I reproduce it under conditions that weren't effectively selected by the experiment that discovered it?

The immediate answer was wonderfully inconvenient:

Not yet.

The next round didn't produce enough fresh analyzable boundary cases to support the intended replication.

That mattered.

I could have treated the original result as confirmed anyway. Instead, the failure exposed a problem with the experiment itself: the boundary population needed to be generated mechanically and blindly.

A failed replication and a failed attempt to construct a valid replication test are not the same thing.

If I'm going to spend this much time trying to kill a result, I should at least make sure I'm using the right weapon.

Tiny geometry can carry enormous consequences

We have an intuitive tendency to associate large with important.

Large variance. Large activation. Large effect. Large cluster. Large signal.

Geometry doesn't owe us that convenience.

A tiny direction can determine which side of a boundary something occupies.

A low-variance dimension can separate regions that appear almost identical in the dominant representation.

The overwhelming majority of a representation can describe where you are, while some microscopic piece determines what happens next.

And that changes the question.

Instead of asking:

Where is most of the information?

I increasingly find myself asking:

Where is the leverage?

Those are very different searches.

And that's why I'm still testing

The PCA result didn't give me an answer.

It gave me a better target.

The interesting part wasn't finding a microscopic component with disproportionate decision relevance.

It was realizing that I now had something else I needed to try very hard to destroy.

That has become a recurring problem around here.

Every answer seems to produce three better questions.

Which is inconvenient.

But considerably more fun.

— Starhound Studio