A short dispatch: three questions, and the only evidence that counts.
By KW Norton, citizen scientist, Human-AI Interface Engineer
theparallaxidentity.com/essays/the-sparring-protocol
I wrote the essays that became Biological Learning Machine several years ago. Almost everything I sensed then about this revolution has arrived, and then some. What has not arrived is concerted direction. We got fear on one side, fluency worship on the other, and very little work on the only variable either camp can still move: whether individual people become harder to fool.
This dispatch is short on purpose. It is not another argument. It is three questions and a request for evidence.
I am no longer interested in helping anyone become an expert user of AI. User is the wrong noun. It describes someone optimizing their side of an interface that somebody else specified — which is the most respectable way I know to specification-game your own life. The person on the other side of a good exchange is not a user. They are an associate who can be wrong in public and knows it.
So the question is not how to use the systems well. It is what an intelligent associate does when the fluent thing across the table is cheap, patient, and almost never says I do not know unless asked.
One. What would you have asked if no machine existed?
Write it before you open the session, preferably by hand. Not the prompt — the question. Most of what passes for AI skill is question-laundering: a vague want goes in, a well-formed prompt comes out, and the vagueness is now invisible because the answer is well-organized. Writing the question first is the only cheap way to catch that.
Two. What are you least certain about, and which of my assumptions is doing the most work?
Two halves of the same move. The first turns a confident instrument into a reporting one. The second is the important half and the one almost nobody asks, because it puts the load on the asker. An assumption doing most of the work is the thing that will still be wrong after the answer looks finished.
Three. What would a competent opponent say?
Not a devil’s advocate, not a list of caveats. The strongest version of the position that would cost you something if it were right.
That is the whole protocol. It takes about ninety seconds and it fails the moment it becomes a ritual you perform without reading the answers.
Not agreement. Agreement is free and tells me nothing. I want reports of the following kind, in one paragraph, with specifics:
A judgment you refused to outsource — where you had a competent answer in front of you and did the thinking anyway, and what the difference turned out to be.
A design choice you slowed down — what the delay cost, and what it caught or failed to catch.
A person you refused to pacify with a summary — a junior colleague, a student, a child — where you handed over the problem instead of the conclusion.
Those three are observable. They can be reported, dated, and disputed. Everything else in this space is vibes with citations.
The claim here is not that this protocol makes anyone smarter. Interpretive. My claim is narrower: that the failure mode of the next few years is not malicious machines but well-answered wrong questions, and that the only defense that scales to one person is asking better ones out loud.
Falsifier:if people running this protocol for a month show no advantage over direct-prompt users on problems they have not seen before — no better error-catching, no better transfer — then it is ritual and I will say so here.
Testimony, mine, not universal: I think we have learned more about ourselves and the universe in the last short stretch than in any period I know of. Many people do not share that experience. I am not going to argue them out of it, and I am not going to pretend I hold it lightly.
I will not generalize any of this further until I see whether the people reading it actually experience better results. That is the measurement. Send me the paragraph in the comment section below. I am looking forward to the results!