Skip to content

Medical AI Is Entering Its Audit Era

#medical-ai #algorithm-auditing #clinical-decision-support #pathology-ai #semi-supervised-learning #llm-safety

The AI infomediary problem ​

A patient types "which family doctor should I see?" into an LLM assistant and books whoever comes back first. That's already happening. No human reviews the recommendation. No regulator watches it. The model decides, silently and at scale, which physicians become visible and which disappear from the patient's screen. These systems are AI infomediaries: algorithms that intermediate one person's choice among other people.

Researchers recently ran a prespecified randomized audit of exactly this process. Seven models (six open-weight, plus gpt-4o-mini) each chose among five synthetic family-medicine physician cards. The attributes on those cards were randomized independently across 3,024 choice sets, three patient personas, nine prompt paraphrases, and nine experimental arms. That's 40,068 scored responses. Gender and ethnicity were signaled through names, following the correspondence-audit methodology used in hiring discrimination studies.

The headline result: reputation signals dominate. Raising a physician's rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points. Raising the fee from $90 to $190 lowers it by 20.0 points. Those two numbers should not surprise anyone. What comes next is the uncomfortable part.

Key numbers from the audit

  • 31.4 pp gain in choice probability from a rating bump of 3.9 to 4.7
  • 20.0 pp drop from a fee increase of $90 to $190
  • 2.5 pp gain for female-signaled names over male-signaled
  • 1.3-2.9 pp gain for Hispanic-, South Asian-, and Black-signaled names over White-signaled
  • 0.03% of stated reasons mentioned gender or ethnicity
  • $11 fee-equivalent value of a content-free first-listed position

Reputation beats everything ​

The demographic effects are small but systematic. Female-signaled names gain 2.5 percentage points. Hispanic-, South Asian-, and Black-signaled names gain 1.3 to 2.9 points over White-signaled names. In fee-equivalent terms, those tilts are worth $7 to $14 per visit. A content-free first-listed position is worth $11.

The causal effects, plotted:

The direction is the surprise. Human audit studies of physician choice consistently find discrimination against women and minority doctors. These models tilt the other way. That doesn't make the bias harmless. It makes it different, and it makes it invisible.

Demographic parity is rejected, but not in the direction anyone predicted. And the models can't tell you about it.

Why self-explanation fails as a monitoring technology ​

The models mentioned gender or ethnicity in at most 0.03% of their stated reasons. They abstained from answering in 0.39% of trials. Ask one of these models whether ethnicity influenced its recommendation, and you'll get a no, because its own explanation never mentions ethnicity. The data says otherwise.

One reasoning model failed the prespecified auditability gate outright. The paper doesn't name it, but the point stands: some models can't even be audited with a fixed stimulus set because they won't follow the protocol.

This has a direct consequence for regulation. Any transparency obligation that relies on model self-report will not detect these effects. The models' explanations are post-hoc rationalizations, and the audit shows they omit the very signals that moved the choice. Behavioral audit measures what the model does. It never asks the model to explain itself. Run it repeatedly against identical stimuli, and you get a monitoring technology that works.

Quick Take: The models keep getting better. The bottlenecks now are data hygiene, annotation budgets, and the absence of routine behavioral audits.

The data problem: redundant notes and scarce labels ​

Two papers this month attack medical data quality from opposite ends, and both land on the same conclusion: the input pipeline matters as much as the model.

The first tackles mechanical ventilation. Ventilator settings need dynamic adjustment as a patient's condition evolves, which makes it a natural fit for reinforcement learning. But the state space is polluted. Clinical notes are inflated by copy-forward text, templating, and repetitive documentation. That redundancy dilutes time-local updates and degrades the state representation. The fix is a redundancy-aware framework that strips duplicated note text before policy learning. Two strategies work:

StrategyMechanismTrade-off
Embedding-space decompositionSVD on local history subspacesComputationally cheap, less interpretable
Sentence-level diffFilters out previously documented sentences before encodingInterpretable, preserves explicit clinical updates

Both beat structured-only and raw-note baselines across four off-policy evaluation methods: Model-Based Rollouts, Fitted Q-Evaluation, Weighted Importance Sampling, and Weighted Doubly Robust Evaluation.

The second tackles annotation scarcity in retinal OCT. Labeling B-scans requires specialists, which is expensive and slow. The proposed framework, TRIAGE, uses a patient-level conformal risk controller with an asymmetric cost matrix. That asymmetry is the key idea: not all errors are equal. Under-grading (missing disease) is worse than over-grading (flagging a healthy scan), so the pseudo-label admission policy is tuned to control the under-grading rate rather than raw accuracy.

With only 20% of the labeled data, TRIAGE hits 89.66% scan-level accuracy, 0.8805 macro-F1, and an 8.34% under-grading rate. The under-grading rate means the model misses disease in roughly 1 in 12 scans. With 5% of labels, it keeps 76.88% accuracy and a 0.1656 under-grading rate. On the public OCT-C8 dataset, it reaches 98% accuracy for 3-class classification with just 1% of labels. Compared to six other semi-supervised methods, it cuts the under-grading rate by 42.7% versus fixed-threshold approaches. The practical read: you can skip 80% of your annotation budget and keep clinical-grade performance, as long as you control the error you actually care about.

Spatial reasoning in language space ​

Pathology has a scale problem. Whole Slide Images are gigapixel, and multimodal LLMs can't fit them in context. The standard workaround is tiling, but tiling severs the tissue neighborhoods that define tumor-stroma interfaces. You lose the spatial context that pathologists actually use.

A new framework called SLMP does spatial reasoning entirely in language space. A WSI region becomes a spatial text graph: tiles are nodes initialized with MLLM descriptions, edges encode spatial adjacency. An LLM refines each tile's description by integrating language messages from adjacent tiles, under an aggregation policy that acts like an adaptive local kernel, but operating on text instead of learned embeddings. The policy is an inspectable prompt, so it can be refined from observed tissue phenotypes via textual gradients, no fine-tuning required.

The results on HER2 and CAMELYON16 regions: +3.3 to +19.6 percentage points in tile-level tumor description accuracy, across general-purpose and pathology-specialized backbones. The biggest gains land on general-purpose models, which narrows the gap to pathology-specialized ones. The random-neighbor ablation is what makes this credible. Swapping adjacent tiles for random ones removes the gain, which confirms the signal is spatial context, not just extra text. And because the policy is a readable prompt, you can inspect the decision rules the model converged on. That's a rare property in pathology AI.

A related paper reformulates morphology-to-transcriptomics prediction as conditional generation in transcriptional program space. Instead of predicting genes independently, it extracts a low-dimensional set of transcriptional programs via consensus NMF, then trains a conditional diffusion model to generate program activations from histology. Same instinct: respect the structure in the data instead of treating every output as independent.

The evaluation gap: TRIPOD+AI ​

None of this matters if the models can't be evaluated or replicated. That's where TRIPOD+AI comes in. It's the updated reporting guideline for prediction models, and it replaces the 2015 TRIPOD statement, which should no longer be used. The new checklist has 27 items, covering both regression and machine learning models, plus a separate abstract checklist.

This is the boring infrastructure that makes everything else possible. Most medical ML papers are underreported to the point where independent evaluation is guesswork. TRIPOD+AI forces authors to state how the model was developed, how missing data was handled, how performance was measured, and what the model actually does in a clinical workflow. Complete reporting is what turns a model in a paper into a model that survives contact with a hospital.

The same pattern shows up across the field. Reviews of point-of-care diagnostics, wearable blood pressure sensors, nanomedicine, and precision oncology all describe the same arc: promising models, then the hard work of validation, standardization, and regulation. Wearable BP sensors, for instance, still can't hit clinical-grade reliability because of calibration drift and motion artifacts, regardless of how good the ML estimation is. The capability questions are getting answered faster than the verification questions.

The deployment frontier: AMIE and consumer AI ​

At the other end of the spectrum, Google's AMIE research system is doing real-time clinical video consultations. It's built on Gemini and Project Astra with a multi-agent architecture, and it interprets visual and auditory cues, guides virtual physical exams, and reasons diagnostically in real time. In a randomized study with simulated consultations and patient actors, clinical evaluators rated AMIE favorably on history-taking thoroughness, diagnostic accuracy, management appropriateness, and communication quality. Patient actors preferred the video experience over text chat. AMIE is still a research system. But the direction is clear: the interaction model for medical AI is moving from chat to multimodal consultation.

The consumer side is moving faster than the clinical side, and the quality bar is all over the place. I tried PawWise, a weekend project that analyzes a dog photo and returns a body condition score, coat health, posture analysis, and breed-specific risks. The serious modes carry the real value. The emergency triage prompt tells the model to err on the side of caution, because a false negative (missing a real emergency) is far worse than a false positive (one unnecessary vet call). That's the same asymmetric cost logic as TRIAGE, applied to a consumer product. When I uploaded a photo of my own dog, the health check came back with a body condition score and a green urgency level. The voice summary read it back in a calm tone. It felt designed to reduce panic. The thread under the project was mostly dog owners sharing photos and asking about the urgency levels. Nobody asked about the model's failure modes. That's the pattern. The developer's biggest technical lesson was that Gemini's structured JSON output eliminated parsing bugs entirely, and that ElevenLabs' Text-to-Dialogue API was the underrated piece, generating multi-voice courtroom drama from a single call. The fun mode is why people open the app again. The serious mode is why it matters.

The gap between AMIE and PawWise is enormous, and accountability is the biggest part of it. AMIE has a research protocol behind it. PawWise has a disclaimer telling you to see a vet. Both are better than nothing. Neither is a substitute for a clinical audit.

Common Pitfalls ​

  1. Trusting model self-explanation for bias detection. The audit showed models influenced by demographic signals while mentioning them in 0.03% of stated reasons. If your monitoring plan is asking the model why, you will miss the bias. Run a behavioral audit with fixed stimuli instead.

  2. Treating all errors as equal in semi-supervised medical learning. Confidence thresholds ignore the asymmetry between under-grading and over-grading. In OCT screening, missing disease is the expensive error. Use an asymmetric cost matrix and control the rate you actually care about.

  3. Feeding raw clinical notes into RL state spaces. Copy-forward text and templating inflate notes and dilute time-local updates. Strip temporal redundancy before encoding, either with SVD on local history subspaces or with a sentence-level diff.

  4. Tiling gigapixel WSIs without preserving spatial context. The tiling workaround makes pathology tractable but destroys the neighborhoods that define tumor-stroma interfaces. SLMP's random-neighbor ablation showed that spatial adjacency, not text volume, drives the gain. If you tile, carry the adjacency information forward.

  5. Publishing prediction models without TRIPOD+AI. The 2015 checklist is officially superseded. Reviewers and regulators are starting to ask for the 27-item checklist, and models without complete reporting are practically un-evaluable.

One thing to remember: the most dangerous medical AI failure modes are the invisible ones. Demographic tilts in physician recommendations, under-grading in OCT screening, redundancy in clinical notes. None of them show up in a demo. All of them show up in an audit.

The Bottom Line ​

If you're building a medical LLM product, run a behavioral audit with fixed stimuli before launch. Self-reported explanations won't surface demographic tilts worth $7 to $14 per visit, and the audit is the only method that caught them.

If you're working with clinical notes or scarce labels, strip temporal redundancy and use asymmetric error costs. Raw notes and confidence thresholds will silently degrade your model. TRIAGE's 42.7% reduction in under-grading shows what the right cost structure buys you.

One thing to watch: AMIE-style video consultation is advancing faster than the regulatory frameworks around it. Expect behavioral audit, not self-explanation, to become the monitoring standard within the next year, because it's the only approach that caught what the models couldn't say.