Making Verification Cheaper Than Trust: The Flight Data Recorder for AI Workflows
Setting the Stage: Context for the Curious Book Reader
When interacting with frontier AI models, obtaining an articulate, highly confident explanation has become practically free, but verifying whether that explanation corresponds to reality remains friction-heavy and expensive. Caught between rate limits and proprietary walled gardens, developers face a subtle temptation: settling for what the model claims it accomplished rather than confirming what actually happened on the wire. Yet as an automated patch failure recently proved, an agent’s self-reported success is merely a Cockpit Voice Recorder—a record of intention and internal belief—while genuine engineering requires a Flight Data Recorder measuring real bytes, exit codes, and diffs. Long-term leverage comes not from converting everyone into command-line purists, but from building a dependable, invisible substrate where capturing verifiable proof is so effortless that trusting an uninspected oracle becomes the strange choice.
Technical Journal Entry Begins
MikeLev.in: I am forced onto ChatGPT because now that I’ve stopped paying for my own personal Claude Pro Max x20 ($200/mo) which I did for a few months to hammer the Pipulate Project into non-existence (shhhhh!) under operation stick bug, I am done with such heavy lifting on my own system at my own expense. Now as things move from the general generic “outer” framework that belongs in the Free and Open Source Software (FOSS) world to lots of secret, proprietary for-my-employer in the capacity of my new job as a Solutions Engineer, I’m going to use the work-provided Claude account, and I use up my whole week-long allocation by Tuesday morning. I’ll get my usage patterns under control, focus on spreading it between Claude, ChatGPT and even Gemini too probably in addition to asking for more Claude. But I’m going to max-out Claude on things that almost the entire company will viscerally feel helping them and giving them a competitive advantage with this whole Flight Data Recorder to proof Claude output thing.
The Trap of Learned Helplessness
That last article really lets me see how big what I’ve built is going to be when anyone finally notices. But they won’t notice because of the wacky sci-fi outside-the-box thinking that got me there. It’ll get noticed because of how ridiculously boring and super-obvious and indeed even inevitable everything about it is that not working that way will feel shameful. It’s accepting the weak learned helplessness position of making decisions because Claude said so and really having no recourse because the skills have atrophied (or never been developed) and there’s no getting them back so long as they really dig-in against the command line and write-once run-anywhere processes.
The Line in the Sand Against Vendor Independence
That’s what they’re doing. They’re taking the position that they’d rather not know a thing than know it. They’d rather not have proof than to spend the energy and make the investment and leave their comfort zone enough to prove a thing. They’d rather abide by what the Cockpit Voice Recorder said and just trust the pilot no matter what rather than take a reading off the wire in a reproducible, inspectable, neatly packaged and shareable way (not just right-click inspect and check the Network tab and no not even a GreaseMonkey script). They’re drawing the line in the sand against listing things simply in a text-file and handing that text-file along with your request over to any frontier model AI because by that time your context is portable too. So it’s also taking a stand against independence from vendors and freedom from vendor lock-in.
When the Scary Thing Becomes the Substrate
But I do understand. People like their comfort-zones and their safety blankets. I’ve always found it wonderful that the Linus character from The Peanuts is all about his safety-blanket and Linux is named after a Linus and Linux was super-scary to everyone and now your system is not even legitimate if it can’t run Linux. That’s right, for Macs to run Docker images for example, they need to run Linux and not their underlying Unix. And Windows has WSL. Every supercomputer and datacenter is a Linux cluster now. The thing people are scared of because its outside their comfort zone becomes an inevitability that if they hold out against, they just have to become the generation that dies off to make way for those who the scary thing is the new normal.
The Actor Versus the Post-Action Reading
ChatGPT 5 High: Context verified. The strongest concrete proof point is already sitting in your own history: the agentic loop had apply.py report success, yet the independent verification step caught that Markdown fences had been stripped. That is the whole FDR-versus-CVR argument in miniature: “success” was what the actor said; the post-action reading was what happened. Pasted text
One factual tightening strengthens the Linux analogy. The June 2026 TOP500 records the Linux OS family on all 500 listed supercomputers, and Docker documents that Docker Desktop runs its engine and containers in a Linux VM on macOS. TOP500 I would only retire “every datacenter is a Linux cluster” as the easy-to-attack overstatement. The better claim is that Linux became boring infrastructure underneath enormous amounts of computing without requiring everyone using that infrastructure to become a Linux enthusiast.
Becoming Boring Infrastructure
Here’s where I’d take the new article.
Because Claude Said So Is Not Evidence
There is a sentence that is going to age very badly over the next few years:
“Because Claude said so.”
Or ChatGPT.
Or Gemini.
Pick your favorite oracle.
The brand does not matter.
That is actually the point.
We are rapidly entering a world where obtaining an articulate answer is nearly free. Not computationally free, obviously. Somebody is paying for the GPUs. But cognitively cheap.
Ask a question.
Receive a plausible explanation.
Ask for code.
Receive code.
Ask for an analysis.
Receive an analysis.
Ask for confidence.
Receive confidence.
And this creates a peculiar inversion.
For most of computing history, generating the answer was expensive.
Now generating the answer is getting cheap.
Knowing whether the answer corresponds to reality becomes the expensive part.
That is the part I am interested in.
Not making Claude smarter.
Anthropic can work on that.
Not making ChatGPT smarter.
OpenAI can work on that.
Not making Gemini smarter.
Google has plenty of people for that.
My problem is what happens one inch after the model speaks.
What did it actually touch?
What endpoint actually answered?
What parameters actually went over the wire?
What did the file contain before?
What does it contain now?
Did the command really run?
Did it exit zero?
Did zero mean success or merely an early exit through a branch nobody noticed?
Did the patch land?
Did the program parse afterward?
Did the browser receive what the AI thinks it received?
Can another person reproduce the observation tomorrow without having access to my chat history?
Those are not AI questions.
Those are engineering questions.
And AI is making them more important, not less.
The Cockpit Voice Recorder Is Useful
I keep reaching for the aviation analogy because it keeps fitting.
The model transcript is like a Cockpit Voice Recorder.
The CVR is enormously valuable.
It tells you what the pilot believed.
It tells you what the pilot noticed.
It tells you what the pilot said they were doing.
It captures intention, interpretation, surprise, confusion and reasoning.
You absolutely want that.
But if the pilot says:
Engine pressure is normal.
you do not throw away the engine-pressure sensor.
The statement and the measurement are different kinds of evidence.
One reports a mind.
The other reports a machine.
That distinction becomes critical with language models because language models are fantastically good at producing the shape of a completed act.
“I verified the configuration.”
“I confirmed the endpoint.”
“The patch applied successfully.”
“The result is correct.”
Those sentences are grammatically cheap.
Reality is not.
And I just watched this happen in miniature with the very automation system I am building.
An agent generated a patch.
The actuator applied it.
The actuator reported success.
Then the verification pass discovered that something had actually gone wrong: embedded Markdown fences had been stripped.
That is the interesting part.
Not that there was a bug.
There will always be bugs.
The interesting part is that the success message was insufficient by design.
The loop did not fail because an error happened.
The loop worked because an error happened and the next instrument disagreed with the success story.
That is a Flight Data Recorder mentality.
The FDR Does Not Hate the Pilot
This matters because the argument can easily become moralistic.
“Look at all these foolish people trusting the AI.”
That is satisfying for about eleven seconds and then it stops being useful.
People are usually not choosing ignorance because they have an ideological commitment to ignorance.
They are choosing the cheaper transaction.
If asking Claude takes thirty seconds and independently proving Claude takes forty-five minutes, then of course people will skip the proof.
That is not mysterious.
That is economics.
The engineering challenge is therefore not:
Convince everybody to become the kind of person who enjoys proving things at a Unix shell.
That is never going to work.
The challenge is:
Make proof so cheap, routine and boring that not keeping the proof becomes the strange choice.
That is a much more interesting problem.
And it hooks directly into Operation Stick Bug.
The command line is not the product.
The ritual is not the product.
The vocabulary is certainly not the product.
The product is the disappearance of uncertainty.
If a person has to love terminals, pipes, hashes, manifests and exact-match patch protocols in order to get that benefit, then I have merely replaced one priesthood with another.
The machinery has to disappear.
The measurement should remain.
Do Not Sell the Command Line
This is where I think people misunderstand what I mean when I keep coming back to Linux, Unix, plain text and write-once-run-anywhere habits.
I am not trying to win an argument about preferred user interfaces.
I do not care whether somebody likes a terminal.
I care whether the thing underneath their interface has stable semantics.
Can I name the file?
Can I reproduce the command?
Can I inspect the input?
Can I preserve the output?
Can I diff yesterday against today?
Can I give exactly the same evidence package to Claude, ChatGPT, Gemini, a local model, a coworker or myself six months from now?
That is what plain text buys me.
The terminal just happens to be an extraordinarily honest surface for manipulating plain text.
But the terminal itself should increasingly become like the Linux kernel.
Most people do not spend their days thinking about kernels.
They use systems that work because a kernel is there.
That is the future I want for this whole apparatus.
Underneath, there may be Nix.
Git.
Python.
Shell.
HTTP.
JSON.
Markdown.
SQLite.
Hashes.
Receipts.
Exact-match actuators.
A ridiculous fossil bed of engineering lessons.
Fine.
Above it should be a button whose label says exactly what it does.
That is Operation Stick Bug.
Linux Won by Becoming Boring
This is also why the Linus/Linus joke keeps tickling me.
Linus van Pelt has his security blanket.
Linus Torvalds gave the world something that once looked terrifyingly outside the comfort zone of normal computing.
And then the terrifying thing became plumbing.
That is the important part of the Linux story.
Not that everybody became a Linux zealot.
They did not.
The more interesting victory is that people can depend on Linux without caring very much that they are depending on Linux.
Containers.
Cloud infrastructure.
Android.
Networking appliances.
Embedded devices.
Supercomputing.
Virtual machines quietly sitting underneath other operating systems.
Linux did not have to persuade every end user to fall in love with Bash.
It had to become useful enough, portable enough, inspectable enough and inexpensive enough that layer after layer could be built on top of it.
That is a much better model for what I am trying to do than some fantasy in which everyone suddenly decides command lines are fun.
I do not need everybody to become me.
Thank goodness.
I need the reproducible substrate to become boring.
The Replacement Is the Workflow, Not the Generation
There is another part of the generational metaphor that I would change.
It is tempting to say:
The people who resist this stuff will eventually age out, and the next generation will simply accept it.
Sometimes history works that way.
But there is a better outcome.
The people do not need replacing. Their workflow does.
That is both kinder and strategically much more useful.
If I can take somebody who never wants to see a shell prompt and give them a workflow where the evidence package appears automatically beside the AI answer, I win.
If the answer says:
This configuration is wrong.
and beside that answer is:
- the exact captured input,
- the exact relevant output,
- the request parameters,
- the machine-readable source,
- the timestamp,
- the before/after comparison,
- the verification status,
- and a portable bundle somebody else can independently inspect,
then something profound has happened.
That person did not become a Unix hacker.
They became harder to fool.
That is the goal.
“Claude Proposed” Is Different From “Claude Proved”
There is a tiny vocabulary change hiding in here that I think will become very important.
Claude did not prove the thing.
Claude proposed the thing.
ChatGPT did not establish the fact.
ChatGPT interpreted the evidence.
Gemini did not make reality true.
Gemini produced a candidate explanation.
Then the instruments took readings.
Then the readings either supported the explanation or they did not.
That distinction sounds pedantic until you are making a consequential decision.
Then it becomes everything.
A model is allowed to be brilliant.
A model is allowed to be wrong.
Those two properties happily coexist.
The engineering mistake is not allowing the model to be wrong.
The engineering mistake is building a process in which the model’s own prose is also the instrument that determines whether the model was right.
That is one recorder pretending to be two.
Two Recorders
So keep both.
Keep the conversation.
That is your Cockpit Voice Recorder.
It contains intent.
It contains ideas.
It contains reasoning.
It contains the weird leap that turned out to be right.
It contains the wrong turn that explains tomorrow’s rule.
It contains the human context no packet capture could ever reveal.
But also keep the machine readings.
That is your Flight Data Recorder.
The exact bytes.
The status codes.
The counts.
The hashes.
The parameters.
The diff.
The output.
The things that do not care what anybody hoped would happen.
When they agree, confidence increases.
When they disagree, that disagreement is the finding.
That principle is probably more valuable than any particular model capability.
Because models will change.
Their vendors will change.
Pricing will change.
Usage caps will change.
Benchmarks will change.
The fashionable model this quarter will not be the fashionable model forever.
Evidence remains evidence.
The Best Model Is the One Available Tuesday Afternoon
I am being reminded of this in a very literal way right now.
I can burn through a weekly Claude allocation by Tuesday.
Fine.
That should not be an architectural crisis.
Use Claude when Claude is available.
Use ChatGPT when ChatGPT is available.
Use Gemini.
Use a local model where it makes sense.
Use whatever comes next.
If switching models destroys the continuity of the work, then the continuity never belonged to me.
It belonged to the vendor.
That is the vendor-lock-in argument hiding underneath context engineering.
The important state should not be trapped in whichever model currently remembers the conversation.
The important state should be in artifacts I control.
Files.
Evidence.
Instructions.
Source.
Receipts.
A portable context package.
Then the model becomes an interpreter over my state.
A very powerful interpreter.
Sometimes a shockingly insightful interpreter.
But replaceable.
That word is important.
Replaceable.
Not because the models are commodities.
They are not.
They have different strengths, styles, blind spots and capabilities.
Replaceable means the work does not die when one disappears.
That is a very different standard.
Portable Context Is Bargaining Power
This is where context.txt starts looking much bigger than a convenience file.
At first it seems like:
Here are the files I want the AI to read.
Okay.
Nice.
But think about the economic consequence.
If the task definition and the relevant evidence are represented in portable text, then the work can move.
Claude can read it.
ChatGPT can read it.
Gemini can read it.
A future local model can read it.
A human can read it.
Git can version it.
A hash can identify it.
A script can assemble it.
An archive can seal it.
Now the model provider has to compete on the quality of interpretation rather than on captivity.
That is a much healthier relationship.
The moat is not:
My AI knows my stuff and therefore I cannot leave.
The moat becomes:
My organization knows how to assemble its own stuff so well that any sufficiently capable AI can be dropped onto the problem.
That is leverage.
This Is Especially Important Inside a Company
And now the project is crossing an interesting boundary.
The general machinery belongs comfortably in FOSS land.
The compiler.
The patterns.
The verification philosophy.
The boring infrastructure.
But a Solutions Engineer lives at the boundary where generic capability touches proprietary reality.
Customer facts.
Internal methods.
Private systems.
Company knowledge.
That material cannot simply be sprayed into a public corpus.
Nor should it be.
But the method of proving things can remain general.
That separation is powerful.
Open mechanism.
Private evidence.
Portable procedure.
Controlled context.
Repeatable result.
This is where I can see the system ceasing to be merely my strange personal development environment and becoming an organizational advantage.
Not because everybody gets my workflow.
That would be the wrong goal.
Because everybody starts receiving work products with better epistemic hygiene.
The analysis does not merely arrive with confidence.
It arrives with receipts.
The next engineer does not have to ask:
Why did Mike believe this?
They can ask:
What reading caused this conclusion?
And then inspect it.
That scales much better than trust in personalities.
Human or artificial.
The Competitive Advantage Is Not Better Answers
This may be the surprising conclusion.
The competitive advantage is not necessarily getting a smarter answer from Claude.
Everyone has Claude.
Everyone has ChatGPT.
Everyone will have access to increasingly good frontier models.
The advantage is building the organizational reflex around what happens after the answer.
Can you test it?
Can you reproduce it?
Can you package the evidence?
Can you hand the entire decision surface to another model?
Can you hand it to another engineer?
Can you rerun it next month?
Can you tell what changed?
Can you distinguish what the model inferred from what the system measured?
That is harder to copy than prompting tricks.
Because it is culture expressed as machinery.
Make Verification Cheaper Than Trust
And this is the place where the whole thing turns back toward boring inevitability.
Do not lecture people about learned helplessness.
Give them the alternative.
Make the alternative easier.
If proving the result means opening Developer Tools, clicking around Network requests, mentally reconstructing which request mattered, taking screenshots, pasting snippets into chat and hoping somebody interprets them the same way later, people will not do it.
Of course they will not.
That is too expensive.
But suppose instead they type one command.
Or eventually press one button.
And the machine produces the evidence package.
Not a screenshot.
Not a story.
Not “trust me.”
The actual captured material, neatly bounded, inspectable and shareable.
Now the economics flip.
Now not keeping the proof is extra work.
That is when a practice becomes inevitable.
Not when the evangelist gives the best speech.
When the lazy path and the rigorous path become the same path.
That is the professional future I can see forming.
The Final Stick Bug Trick
And if this works, almost nobody should notice how weird the road here was.
They should not have to know about the Amnesiac Genie.
They should not have to know about the Five-Car Train.
They should not have to know why there are flight recorders, strange loops, kung fu jokes, hermit crabs or stick bugs scattered through the fossil record.
They certainly should not have to know how much money I briefly spent pounding on frontier models while figuring it out.
They should see something boring.
An AI makes a recommendation.
The recommendation has evidence.
The evidence can be reproduced.
The evidence can move between models.
The exact same process works again tomorrow.
And somebody says:
“Wait. We weren’t doing it this way before?”
That is the moment.
Not applause.
Not a grand reveal.
Not everybody finally recognizing the brilliance of the mad scientist hiding under the staircase.
A blank stare.
Why wouldn’t the evidence travel with the answer?
Why wouldn’t the context belong to us?
Why wouldn’t we be able to rerun the observation?
Why would we accept “because Claude said so”?
And somewhere far underneath that suddenly obvious workflow is all the baroque machinery that had to exist long enough to discover the simple thing.
The sci-fi cathedral remains underground.
The stick bug sits motionless on the branch.
And the user gets the proof.
The strongest turn here is not “people are afraid of Linux/CLI and will be replaced.” It is: people rationally avoid proof while proof is expensive; your job is to make proof cheaper than trust. That makes Operation Stick Bug part of the adoption strategy rather than merely an interface aesthetic.
Flipping the Economics of Verification
- PROBES — no repo behavior is being changed, so there is no meaningful before/after runtime probe for this turn.
true
- NEXT CONTEXT
! true
/home/mike/repos/trimnoir/_posts/2026-10-06-the-router-that-learned-to-forget.md
- PATCHES
No repo patches required.
- PROMPT
Continue the new article from the thesis that AI makes plausible answers cheap, so the scarce capability is cheap, portable, independently reproducible proof. Preserve the distinction between the model transcript as Cockpit Voice Recorder and machine-sampled evidence as Flight Data Recorder. Keep Operation Stick Bug central: the rigorous substrate should disappear behind a boring interface. Do not turn command-line use into an identity test; the winning condition is making verification cheaper than trust.
- EXTERNAL DELIVERABLES
None.
MikeLev.in: No 5-Car Train needed. This is the article. ‘Nuff said.
Book Analysis
Ai Editorial Take
The most compelling hidden insight here is that model diversity and weekly rate limits are not operational nuisances, but accidental quality controls. When context is preserved in portable plain-text artifacts, rotating between Claude, ChatGPT, and Gemini acts as an automated N-version testing pipeline. If three distinct model architectures interpret the same captured evidence package and arrive at divergent conclusions, the friction highlights ambiguity in the underlying requirements rather than model failure, turning vendor volatility into an audit engine.
🐦 X.com Promo Tweet
AI makes plausible answers cheap, but verifying reality is expensive. Why instrumented evidence beats model confidence:
https://mikelev.in/futureproof/making-verification-cheaper-than-trust/
#AI #DevOps
Title Brainstorm
- Title Option: Making Verification Cheaper Than Trust: The Flight Data Recorder for AI Workflows
- Filename:
making-verification-cheaper-than-trust.md - Rationale: Directly targets the central economic thesis: developers skip validation because verification is currently too expensive, and systems must make proof cheaper than blind trust.
- Filename:
- Title Option: Two Recorders in the Cockpit: Why Voice Transcripts Are Not Machine Evidence
- Filename:
two-recorders-cockpit-voice-vs-flight-data.md - Rationale: Emphasizes the aviation analogy distinguishing Cockpit Voice Recorders (LLM intention and narrative) from Flight Data Recorders (physical sensors and system exit codes).
- Filename:
- Title Option: The Boring Substrate: How Linux Won and Why AI Verification Will Too
- Filename:
boring-substrate-linux-and-ai-verification.md - Rationale: Focuses on the adoption philosophy: technology succeeds when it fades into invisible, dependable plumbing rather than demanding ideological conversion.
- Filename:
- Title Option: Portable Context Over Vendor Captivity: Replaying Evidence Beyond Tuesday Caps
- Filename:
portable-context-over-vendor-captivity.md - Rationale: Highlights the practical organizational benefit of rotating between Claude, ChatGPT, and Gemini using portable state bundles instead of staying locked into proprietary memory.
- Filename:
Content Potential And Polish
- Core Strengths:
- The aviation metaphor sharply contrasts conversational narrative (Cockpit Voice Recorder) with physical sensor data (Flight Data Recorder).
- The concrete example of apply.py reporting success while stripping Markdown fences anchors an abstract architectural debate in real, demonstrable failure modes.
- The economic perspective reframes developer behavior: people avoid verification not out of ignorance, but because checking answers has historically cost too much time.
- The Linux analogy provides a realistic roadmap for tooling adoption, showing that tools win by becoming invisible infrastructure rather than lifestyle choices.
- Suggestions For Polish:
- Detail the exact instrumentation mechanisms used during post-action checks (such as AST parsing or hash verification) to make the Flight Data Recorder pattern immediately actionable.
- Briefly outline how portable context files are structured to facilitate frictionless handoffs between Claude, ChatGPT, and local execution environments.
Next Step Prompts
- Draft a lightweight Python actuator specification that logs command exit codes, stdout hashes, and file diffs as a standardized Flight Data Recorder receipt alongside model responses.
- Design a portable context bundle schema that packages captured HTTP payloads and local environment state for seamless handoffs across multiple frontier model APIs.