flux7.art / research
● ongoing · paper in preparation

Believe & act

Decision research for LLM agents

An agent that investigates a fault must decide what to check next and when to stop. In this research both decisions stay out of the language model. The model reads the raw output of each check, while code computes the belief and the choice, so every run can be replayed from its log.

The setting

A Kafka lab (one 4.0 broker, one producer, three consumers) receives one of six injected faults. Seven checks can be run, each costing between 1 and 8 minutes, and a wrong diagnosis costs 60 minutes. The loss of an episode is the time spent on checks plus 60 minutes if the concluded cause is wrong, and every method is scored on that same loss.

  raw output ──► reader      LLM or parser        "is the signal present?"  →  p
                    │
                    ▼
               belief      6 causes             Bayes, tables counted on past episodes
                    │
                    ▼
               choice      next check or stop   value of information per minute
                                                  stop at λ = 12 minutes per bit

               the reader is the only part that calls a language model

Study 1: 120 faults, three hypotheses

The study was pre-registered on OSF on 5 October 2026, before any reading of the series. It ran 120 episodes, 20 per cause, with two-fold cross-fitting so that every episode is scored by tables that never saw it. Intervals are paired-bootstrap at 98.33 % (Bonferroni for three hypotheses), in minutes of loss.

H1 · confirmed
Likelihood tables corrected by counted incidents beat the tables written by the language model: 19.11 against 24.80 minutes (difference −5.69, interval [−8.89, −2.72]).
H2 · confirmed
Stopping at the price of information beats the stopping rule of the HESP controller: 19.73 against 23.10 (−3.37, [−4.24, −2.52]).
H3 · contradicted
Claude Sonnet 5.5 alone, reading the same outputs with the whole picture, beat the hybrid that used Jev as its reader: 8.56 against 19.11 (+10.55, [+6.00, +15.75]). Most of the gap came from a single cause, on which the reader returned probabilities near 0.5 for a rate ten times the usual level.

Study 2: a reader routed by the kind of signal

The second study was pre-registered after the analysis of the first and before any new episode, on 120 new faults. Checks whose verdict can be read off the text go to a judgment model; checks that need a comparison of numbers go to extraction, followed by a stated comparison in code. Intervals are at 97.5 % (Bonferroni for two).

H4 · confirmed
The routed reader beats the Jev reader: 9.33 against 18.68 minutes (−9.35, [−14.46, −4.85]).
H5 · confirmed
The hybrid with the routed reader is not worse than Sonnet alone by more than the pre-registered margin of 2 minutes: 9.33 against 8.63 (+0.70, [+0.34, +1.08]). Sonnet alone therefore remains slightly better.
Cost
$0.0036 per diagnosis for the routed hybrid, against $0.0226 for Sonnet alone. The routed hybrid concluded correctly on all 120 faults, with no failed extraction out of 1,440 readings.
Support
13.3 % of Sonnet's conclusions were right without being supported by its own path: replaying the checks it ran through the same tables does not bring the concluded cause to a posterior of 0.8.

Limits

01

The task sits near its ceiling, since the best arms conclude correctly on 120 faults out of 120. The studies therefore measure cost more than accuracy.

02

On a lab with fixed output formats, the routed reader behaves like the hand-written reference parser, partly by construction.

03

Both studies use one lab and one fault at a time. Drift in the history, asymmetric error costs and simultaneous causes are planned separately, each with its own pre-registration.

04

The model alone and the hybrid's readers do not receive the same context, because a narrow reader sees one output and one question.

Prior work

Asking a language model for the likelihood of an observation and applying Bayes outside the model was published by Amin (2026, arXiv:2601.01522). Stopping when the next observation is no longer worth its cost goes back to Wald's sequential analysis. The baseline controller is HESP (Liu and Wang 2026, arXiv:2609.33446), and Jev is the decision model of TypeSafe. The contribution of this work is empirical: pre-registered measurements on real faults, with the reading stage measured apart from the decision.

Next

A paper is in preparation. The next study replaces the price per bit with the price of the error itself and compares the result with the exact optimum on ambiguous faults. A second testbed then brings several agents into one world, where decision theory meets game theory: AI personas that have to decide what to believe about each other before they act.

Bayesian inference value of information optimal stopping pre-registration LLM agents Kafka Jev (TypeSafe) Claude Sonnet 5.5
← Back to the home page