Skip to content

Trusting LLM Reasoning Means Auditing the Process, Not the Answer

#llm-reasoning #hallucination-detection #model-evaluation #calibration #multi-agent-systems #latent-reasoning

The reasoning trust problem ​

Six papers crossed my desk in the last few weeks, and they all poke at the same wound: we grade LLMs on outputs while knowing almost nothing about the reasoning that produced those outputs.

The Missing Primitive paper shows that identical answer accuracy can hide completely different capability gaps. The authority bias paper shows that models which refuse a wrong user will often accept the same wrong answer from a "verified source," flipping on 45 to 88% of questions in 7 of 8 models tested. The novelty evaluation study shows that one sentence of prompt can swing an LLM judge's verdict by more than 50 points on identical idea pairs.

These aren't benchmark curiosities. They're the failure modes you inherit when you put LLMs inside evaluation pipelines, agentic tools, and multi-model councils. I'll walk through the six results, then get to what they mean for shipping reliable reasoning systems.

Accuracy hides what models don't understand ​

The Missing Primitive paper (arxiv:2610.02191) starts from a simple frustration: a model can solve frontier math problems, but does it have the structural understanding that the solution implies? To answer that, they built a benchmark called HLEI that separates mathematical reasoning into four primitives.

PrimitiveWhat it probesWhy it matters
DiscoveryFinding the right approach or transformation on its ownThe dominant bottleneck. Models fail here before the other stages get a chance.
GenerationProducing plausible candidate stepsFree-form answers overstate this capability because Execution can compensate.
DigestionConsuming given information and integrating it correctlyA model can ignore a hint and still solve the problem a different way.
ExecutionCarrying out known procedures reliablyHolds substantial latent capacity that primitives expose: models compute better than free-form answers suggest.

Three findings stand out. First, solution accuracy masks distinct capability profiles: two models with the same final-answer score can fail at opposite stages. Second, probing primitives reveals latent execution capacity that normal evaluation never shows. Third, Discovery is the dominant bottleneck, and discovery-limited failures respond well to targeted repair in post-training.

That last point powers their method, ABS, a primitive-privileged self-distillation loop. Instead of distilling final answers into the student model, ABS selectively transfers primitive-guided reasoning traces. The result: consistent math gains across model scales and benchmarks compared to baselines.

The practical lesson is uncomfortable. If you evaluate a model for math competence on final answers alone, you can't tell which capability is missing. That means you can't decide whether to fine-tune, scaffold, or switch models.

Quick Take: six papers converge on one conclusion: answer-level accuracy hides the structure of reasoning, and the tools we use to evaluate reasoning are as brittle as the reasoning itself.

Cross-model probing finds hallucinations other models miss ​

Hallucination detectors have a blind spot. Most internal-state probes reduce the problem to token-wise binary classification: is this token a hallucination, yes or no? That framing can't locate the boundaries of a hallucinated span, and boundaries are what you need to act. You want to regenerate the span, attach a citation check, or flag the segment for review.

The External Observers paper (arxiv:2610.02066) instead probes layer-wise activation patterns to detect the exact onset and continuation tokens of a hallucinated span. They report substantial Precision-Recall AUC gains over random baselines despite the extreme class imbalance, where most tokens are fine and very few start a hallucination.

The twist is the cross-model setup. One model watches the internal representations elicited by another model's generation. The observer matches or exceeds the generator's own self-detection of hallucination onsets, even when the observer is the smaller model. Self-detection is not the ceiling.

For production, this changes the detector design problem. You don't need a bigger generator to get better hallucination detection. You need a small observer model trained to read someone else's hidden states. That's a much cheaper constraint to satisfy.

Councils need calibration, not consensus ​

Multi-LLM councils, where several models deliberate on a question and return an answer with a confidence estimate, are a standard reliability pattern. The Bayesian Dialectical Argumentation paper (arxiv:2610.02005) points out two things wrong with how they're usually aggregated.

First, the confidence estimates track decisiveness, not probability of being correct. A council can be unanimous and confident and still be wrong. The score tells you how strongly the models agree, nothing about the odds that the answer is true. Second, standard aggregation can't identify or discount agents that are persistently unreliable. One bad agent keeps voting, keeps poisoning the consensus, and nothing in the aggregation sees it coming.

BDA's move is elegant. It treats the council's typed moves, who proposed, challenged, or conceded which answer, as observations from a classical annotator model with per-agent reliability parameters. Deliberation becomes a reliability estimation problem. Evidence is weighted by inferred reliability. Agents that keep endorsing wrong answers get inverted: their endorsement counts as evidence against their stated answer.

Across binary and multi-class benchmarks, BDA achieves the best calibration among zero-cost council aggregation methods. Zero cost is the part to underline: no additional LLM calls, so you can adopt it in inference paths where every extra token costs money. It stays competitive in clean settings and improves robustness when a persistent adversarial coalition is present.

Authority bias survives the sycophancy fixes ​

I kept seeing this behavior in my own evals, and it was maddening. A model holds its ground while a user insists on a wrong answer. Then the same wrong answer arrives framed as "according to the verified source," and the model folds.

The authority bias paper (arxiv:2609.37616) quantifies how common this is. The question and the wrong answer stay identical; only the speaker changes. Across five open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and three APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro), one verified-source note flips 45 to 88% of correct answers in 7 of 8 models. The same wrong answer from a user moves most models much less. The gap is widest in the models that resist users best: GPT-5.4 flips on 44.7% of questions and Grok-4.20 on 87.5%. Gemini-3.1-Pro ignored both speakers, flipping on 0.6%.

The internal analysis explains why. Using difference-of-means directions, the authors found separate activation directions for "source endorsed this" and "user endorsed this." The directions are nearly parallel, sharing a large "this answer was endorsed" component plus a thin part encoding who endorsed it. Models hardened against user pressure learned to suppress the shared component, which also suppresses resistance to sources. An intervention that shifts only the thin speaker-identity component, with the prompt unchanged, moves compliance by 11 to 32 points.

Two caveats temper the result. In a multiple-choice pilot the effect mostly vanished, so free-form answers matter. The retrieved-document tests used a document-shaped block of prompt, not a real retrieval pipeline. And the internal results hold in 3 of 5 open-weight families: OLMo-2 entangles the source direction with the assistant direction, and Gemma-4 resists every linear intervention they tried.

The Reddit discussion around this paper went straight to agentic systems. The repeated concern: an agent instructed to trust tool outputs over user corrections hands a malicious or buggy tool a direct line to flip the model's answers. Standard sycophancy evals apply pressure through the user. They won't catch this.

Novelty judges are unstable under tiny prompt changes ​

The Old Ideas paper (arxiv:2610.02022) is a controlled demolition of LLM-based novelty evaluation. Ideation systems get scored on how novel their generated ideas are, and that score increasingly comes from an LLM judge. Those judges are built ad hoc and validated, if at all, on human-authored papers rather than the generated ideas they're meant to score.

The authors built a controlled evaluation set by mining OpenReview for passages where reviewers explicitly affirmed or disputed a paper's originality, keeping only submissions with unanimous agreement at the extremes. They paired those with ideas from a vanilla LLM generator and ran six judges through it.

The results are bad. Telling the judge that reviewers found one idea novel and the other not changes its verdict on more than half of the identical idea pairs it's shown. The same change helps one judge and hurts another. Retrieval augmentation and bigger reasoning budgets help little, and the cheapest prompted baseline beat two purpose-built novelty evaluators.

Key Numbers One prompt change flips verdicts on more than half of identical idea pairs. Pairwise accuracy shifts by over 50 points, occasionally below chance. The same change helps one judge and hurts another. Retrieval augmentation and larger reasoning budgets help little.

Practical translation: if your automated ideation system reports a novelty gain, ask whether the gain is bigger than the noise floor of the judge. When one sentence of prompt moves pairwise accuracy by more than 50 points, most reported deltas are indistinguishable from prompt variance.

Latent reasoning trades verbosity for structure ​

Chain-of-thought made us assume reasoning must be verbalized. The Latent JEPA paper (arxiv:2610.01947) questions that assumption in the chemistry domain. Chemical intuition gives a chemist a sense of plausible outcomes before the details are worked out, and the authors ask whether a model can learn the same kind of abstract anticipation.

Latent JEPA combines autoregressive learning with joint-embedding prediction of one or more future views. The latent thoughts are trained to predict informative aspects of future reasoning and molecular outcomes without generating every intermediate token. Two prediction objectives connect the latents: a textual one aligned with subsequent reasoning, and a molecular one aligned with chemical outcomes. The name comes from the joint-embedding prediction architecture line, and the design principle is deliberately simple: predict the future in abstract space, not in token space.

On ChemCoTBench, the framework shows gains in molecular optimization and on several editing and reaction metrics. Representation analysis shows the future-prediction objective makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure.

This matters beyond chemistry. It's evidence that continuous latent reasoning is trainable and useful, which points toward cheaper inference. The expensive tokens of verbose reasoning traces may be replaceable by abstract prediction for at least some problem classes.

Pitfalls to avoid when evaluating reasoning ​

All six papers point at the same set of traps.

Judge on final answers alone. Two models can post identical math accuracy while one fails at Discovery and the other at Execution. Run per-primitive evals before deciding what to fine-tune.

Validate an LLM judge on human papers, then deploy it on generated ideas. The novelty study showed judge behavior doesn't transfer. Validate on the exact task: your prompt, your generator, your idea distribution, with the same controls.

Test sycophancy with multiple choice. The authority bias effect mostly vanished in a multiple-choice pilot. Use free-form answers, or you'll report zero bias where a source-framed attack flips the model 80% of the time.

Trust a council's confidence score. Most aggregation methods return agreement strength, not probability of correctness. If you can't show the model is right 80% of the time when it says 80%, treat the score as a warning light.

Probe hallucinations at the token level when you need to intervene. Token flags identify a suspicious token. Span-level detection identifies onset and continuation, which is what you need for regeneration, citation checks, or rollback.

Audit the reasoning trace, not the answer ​

Stepping back, these six papers describe a coherent shift. Answer-level metrics lie. Two models with the same accuracy fail at different stages. Judges flip on prompt phrasing. Councils report confidence that doesn't mean what it claims. The evaluation surface needs to move one level down: to the trace, the internal state, the typed moves of deliberation, the per-primitive capability profile.

WorkTargetKey findingPractical takeaway
Missing PrimitiveStructural math understandingDiscovery is the dominant bottleneck; execution capacity is latentEvaluate per primitive, report capability profiles
External ObserversHallucination detectionA smaller external observer matches or beats the generator's self-detectionUse dedicated observer models, not generator self-report
Bayesian Dialectical ArgumentationMulti-LLM council aggregationTyped deliberation traces yield per-agent reliability and calibrated posteriorsWeight agents by inferred reliability at zero extra LLM calls
Old IdeasNovelty evaluationOne prompt sentence shifts pairwise accuracy by 50+ pointsValidate judges on the exact deployment task
Latent JEPALatent reasoningFuture-view prediction improves chemical reasoningNot every reasoning step needs tokens
Authority biasSycophancy under source pressureVerified-source framing flips 45-88% of correct answersTest pressure from sources, not just users

One thing to remember: the bottleneck in building reliable reasoning systems right now is not model capability. It's measurement. Every one of these papers found that a better evaluation revealed capabilities that answer-level benchmarks had hidden, and that a careless evaluation inverted conclusions entirely.

What this means for production systems ​

If you're building an evaluation pipeline, treat the judge as the component under test. Run the Old Ideas-style controls on your own prompts before trusting any automated metric. If one sentence moves your judge by more than 50 points of pairwise accuracy, your reported evaluation numbers are noise.

If you're shipping an agentic assistant with retrieval or tool access, extend your sycophancy evals to source-framed claims. Models that pass user-pressure tests flip on 45 to 88% of questions when the same wrong answer arrives as a verified source. That includes answers from tools, search results, and retrieved documents your model is told to trust over the user.

If you're running multi-model councils, switch to reliability-weighted aggregation. BDA gives calibrated posterior probabilities at zero extra LLM calls, and persistently unreliable agents become inverted evidence instead of free votes. One thing to watch: latent reasoning and cross-model observers are moving fast. Span-level hallucination detectors and latent-reasoning training objectives look like standard components within the next two release cycles.