An agent that investigates a fault must decide what to check next and when to stop. In this research both decisions stay out of the language model. The model reads the raw output of each check, while code computes the belief and the choice, so every run can be replayed from its log.
A Kafka lab (one 4.0 broker, one producer, three consumers) receives one of six injected faults. Seven checks can be run, each costing between 1 and 8 minutes, and a wrong diagnosis costs 60 minutes. The loss of an episode is the time spent on checks plus 60 minutes if the concluded cause is wrong, and every method is scored on that same loss.
raw output ──► reader LLM or parser "is the signal present?" → p
│
▼
belief 6 causes Bayes, tables counted on past episodes
│
▼
choice next check or stop value of information per minute
stop at λ = 12 minutes per bit
the reader is the only part that calls a language model
The study was pre-registered on OSF on 5 October 2026, before any reading of the series. It ran 120 episodes, 20 per cause, with two-fold cross-fitting so that every episode is scored by tables that never saw it. Intervals are paired-bootstrap at 98.33 % (Bonferroni for three hypotheses), in minutes of loss.
The second study was pre-registered after the analysis of the first and before any new episode, on 120 new faults. Checks whose verdict can be read off the text go to a judgment model; checks that need a comparison of numbers go to extraction, followed by a stated comparison in code. Intervals are at 97.5 % (Bonferroni for two).
The task sits near its ceiling, since the best arms conclude correctly on 120 faults out of 120. The studies therefore measure cost more than accuracy.
On a lab with fixed output formats, the routed reader behaves like the hand-written reference parser, partly by construction.
Both studies use one lab and one fault at a time. Drift in the history, asymmetric error costs and simultaneous causes are planned separately, each with its own pre-registration.
The model alone and the hybrid's readers do not receive the same context, because a narrow reader sees one output and one question.
Asking a language model for the likelihood of an observation and applying Bayes outside the model was published by Amin (2026, arXiv:2601.01522). Stopping when the next observation is no longer worth its cost goes back to Wald's sequential analysis. The baseline controller is HESP (Liu and Wang 2026, arXiv:2609.33446), and Jev is the decision model of TypeSafe. The contribution of this work is empirical: pre-registered measurements on real faults, with the reading stage measured apart from the decision.
A paper is in preparation. The next study replaces the price per bit with the price of the error itself and compares the result with the exact optimum on ambiguous faults. A second testbed then brings several agents into one world, where decision theory meets game theory: AI personas that have to decide what to believe about each other before they act.