Skip to content

MLLMs Look Good on Benchmarks. They Fall Apart Under Pressure.

#multimodal-llms #benchmark-design #vision-language-models #long-video-understanding #model-evaluation

The pattern across six new papers ​

Six multimodal papers landed on arXiv this month. They cover different ground: image classification routing, social audio-visual QA, scientific figure understanding, TikZ code editing, and long-form video. But they converge on a single uncomfortable finding. MLLMs are getting more accurate, and they are not getting more reliable.

The gap shows up in different forms. A model that nails reasoning accuracy hallucinates on unreadable images 96% of the time. Chain-of-thought reasoning costs more and doesn't beat a plain fine-tune. The strongest video model sits 22 points below human performance on month-level memory tasks. None of these failures show up on the benchmarks most teams use to pick models.

Key numbers

  • 96%: rate at which GPT-5.2 hallucinates unreadable content in scientific figures when the image itself is unreadable
  • 22.4 points: gap between Gemini 2.5 Pro and the corrected human baseline on EgoMonth, a month-level egocentric video benchmark
  • 23%: share of IntentBench questions trivially answerable without the video at all
  • 45.35% → 83.40%: Qwen3.5-4B's compilation success on Edit2TikZ after a reconstruction-then-editing curriculum

The table below maps the six papers onto what they actually measure.

BenchmarkDomainScaleHeadline finding
ARMDILCross-dataset classification4 backbones, unified label spaceAn MLLM router matches trained routers, and new domains drop in via prompts
IntentBench-PrimeSocial audio-visual QA~70% of original IntentBench after cleanupVanilla SFT matches CoT reasoning at a fraction of the cost
SciFigBenchScientific figures250 figures, 34,000 evaluation setupsHigh reasoning accuracy doesn't imply behavioral reliability
Edit2TikZFigure editing with TikZ1,548 samplesProprietary models compile only 75% of the time
NARUJapanese long video1,481 questions, 146.8 hoursModels fail at long-range narrative and cultural reasoning
EgoMonthEgocentric video300+ hours, 1,443 QA pairsThe best model is a lossy summarizer, not a memorizer

Routing beats retraining ​

The first paper, ARMDIL, tackles a practical problem: no single vision backbone generalizes across domains. ResNets are reliable on texture-heavy data but weak on semantic matching. Self-supervised encoders capture structure but miss fine-grained categories. Vision-language models bring language grounding but drift on low-level visual detail. Training one model that does all of it well is expensive and brittle.

ARMDIL's answer is to stop trying. An MLLM agent looks at each image and routes it to the backbone best suited for it, with all backbones sharing a unified label space built from multiple datasets. The router is the only learned component that matters, and it reasons in natural language, so you can read why it picked a backbone.

The result that matters: prompt-based routing performs competitively with routers trained specifically for the task. That means adding a new domain is a prompt edit, not a training run. For a team maintaining a classifier across shifting domains, that's the difference between a one-line config change and a week of fine-tuning.

The paper also exposes how differently the architectures fail. The failures are complementary. A router that knows each backbone's weak spots can sidestep them.

Contrast that with Paths, the RGB-Event person re-identification paper in the same batch. It's the opposite bet: when the task is narrow and the sensor pairing is fixed, a purpose-built architecture wins. Paths couples spatial and temporal modeling in a single Transformer and fuses RGB and event streams at global and local levels, beating generic approaches on EvReID, MARS, and iLIDS-VID. Routing is for open worlds. Specialization is for closed ones. Pick based on how much of the input distribution you can enumerate ahead of time.

Chain-of-thought is expensive and doesn't help ​

The IntentBench paper is the one that should make people in the audio-visual space uncomfortable. Training MLLMs for social audio-visual QA has converged on chain-of-thought reasoning as the default. The paper tests whether that default earns its cost.

First, the benchmark itself was noisy. About 7% of IntentBench questions were broken, and 23% were trivially answerable without the video. The authors cleaned those out and released IntentBench-Prime. When I looked at the trivially answerable questions, the pattern was obvious: the text modality leaked the answer, and the video was decoration. Any benchmark with that much leakage inflates every model that trains on it.

Second, the reasoning methods didn't work. A Vanilla SFT baseline, which just fine-tunes the model on input-output pairs without explicit reasoning traces, matched or outperformed the reasoning approaches across three benchmarks. At a fraction of the training and inference cost. The paper also shows that a textual caption alone, with no video, performs on par with Vanilla SFT on video input. That's a damning result for the current generation of models: they're learning social priors from text and barely using the visual stream.

Quick Take: The expensive reasoning traces you're adding to your audio-visual model probably aren't helping, so measure the delta before you pay the token cost.

This is the finding I'd act on first. If you're fine-tuning an MLLM for social understanding, run the Vanilla SFT baseline before you build the CoT pipeline. The paper's authors are right that it should be the essential baseline for every new fine-tuning technique in this space.

Accuracy hides hallucination ​

SciFigBench attacks a different blind spot: how models behave when the visual evidence is missing or misleading. The benchmark stress-tests VLMs on scientific figures with image transformations, caption-bias probes, and selective blurring, generating over 34,000 evaluation setups from 250 annotated figures.

The proposed A-R-I framework scores three behaviors: whether a model admits insufficient evidence, resists misleading context, and infers cautiously from partial information. The results split models into two camps.

ModelDescription quality (MQM)Reasoning accuracyAdmits uncertainty on unreadable figuresResistance score
GPT-5.291.678.4%4%weaker
Gemini 3.1 Pro90.281.0%71%0.91

The two models are nearly tied on description quality and reasoning accuracy. Gemini actually edges out GPT-5.2 on reasoning. But when the figure is unreadable, GPT-5.2 hallucinates content 96% of the time instead of saying it can't tell. Gemini admits uncertainty in 71% of those cases.

For scientific workflows, that difference is the difference between usable and dangerous. A model that fabricates unreadable axis labels in a figure you're about to cite is worse than a model that says "I can't read this." The benchmark's core message is that perception and reasoning accuracy don't predict behavioral reliability. If you're deploying a VLM in any setting where a confident wrong answer has real cost, you need to evaluate uncertainty behavior explicitly.

Long video: lossy summarizers, not memorizers ​

The two long-video benchmarks in this batch, EgoMonth and NARU, both target the gap between short clips and real experience. EgoMonth is the first month-level egocentric video benchmark: over 300 hours of first-person recordings from 20 participants spanning 20 to 120 days, with 1,443 human-crafted questions organized into 14 tasks across three cognitive levels.

The results are sobering. Gemini 2.5 Pro, the strongest model tested, hits 71.8% macro-average accuracy. The corrected human baseline is 94.2%. On tasks like Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, several models land near or below the 25% chance level.

The paper's framing is the right one: current MLLMs function as lossy summarizers rather than faithful memorizers. A context window is not memory. You can't compress 120 days of first-person video into a context window and expect consistent recall of spatial layout or route structure. This points at architectures with explicit long-term spatiotemporal memory, not bigger context windows.

NARU makes the same point from a different angle. It evaluates narrative evolution and cultural reasoning in Japanese long-form video: 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. The construction pipeline uses hierarchical memory-based annotation to turn raw video into structured event, narrative, and cultural annotations, with two native-speaker verification stages involving 68 annotators.

Across eight model configurations, the failures cluster in long-range narrative integration and culturally grounded reasoning. 146.8 hours of high-context video is beyond any context window. Understanding why a character's behavior shifts across episodes requires holding a model of the story, not a transcript. The models don't have that.

Editing figures is harder than generating them ​

Edit2TikZ looks at a task that sounds adjacent to generation but is actually much harder: editing a scientific figure by writing TikZ code. The model has to recover the visual structure, ground the requested change, generate compilable code, and preserve everything unrelated to the edit. One mistake in any step breaks the result.

The benchmark has 1,548 samples covering real-world and synthetic edits, with both textual and visual localization requests and multi-step edits. The evaluation framework checks two things: whether the requested edit happened, and whether the rest of the figure survived.

The results explain why the benchmark exists. Proprietary models average only 75% compilation success, and they're worse at actual edit correctness. Compact models below 9B parameters struggle with instruction following and complete figure generation. When I ran the numbers on the small models, the failure mode was consistent: they'd generate plausible TikZ that didn't compile, or they'd nail the edit and destroy the surrounding figure.

The fix is a curriculum. The authors built TikZEditMix, a mixed training set, and trained compact models with a reconstruction-then-editing schedule. The gains are large:

TrainingCompilation successAverage metric gain
Qwen3.5-4B, direct45.35%baseline
Qwen3.5-4B, reconstruction-then-editing83.40%+18.7 points

That's a 38-point jump in compilation success from a training schedule change. It tells you the small models weren't lacking capacity; they were lacking the right learning order. Learn to reproduce a figure before you learn to modify it.

What trips people up ​

Across these six papers, the same mistakes keep showing up in how teams build and evaluate multimodal systems.

Trusting accuracy as the sole selection metric. SciFigBench shows two models with near-identical reasoning accuracy and wildly different hallucination behavior. If you pick a model for scientific figure understanding on accuracy alone, you can end up with one that fabricates unreadable content 96% of the time. Evaluate uncertainty behavior explicitly.

Defaulting to CoT without a baseline. The IntentBench paper shows Vanilla SFT matching or beating CoT reasoning methods across three benchmarks. Reasoning traces cost tokens at training and inference time. Run the plain baseline first. If CoT doesn't beat it, you're paying for nothing.

Confusing context length with memory. EgoMonth's 300 hours of video and NARU's 146.8 hours are not context problems. No window will hold them. The models fail at consolidation and retrieval across days. If your task needs week-level consistency, plan for a memory architecture or retrieval layer, not a bigger context.

Fine-tuning small models on the target task directly. Edit2TikZ's curriculum result is unambiguous: Qwen3.5-4B jumped from 45.35% to 83.40% compilation success just by learning reconstruction before editing. Skipping the prerequisite task leaves performance on the table.

Trusting benchmark numbers without checking for leakage. 23% of IntentBench was trivially answerable without video input. Before you report a benchmark result, probe for shortcuts. A few manual passes over the questions can save you from publishing inflated numbers.

One thing to remember ​

The common thread in this batch isn't the benchmarks themselves. It's that every paper had to build new evaluation infrastructure because existing benchmarks hid the failure mode. Noisy questions, trivially answerable items, accuracy scores that masked hallucination, context windows that masqueraded as memory. If you're evaluating an MLLM, assume your current metric is hiding something, and build a probe that targets the behavior you actually care about.

The Bottom Line ​

  • If you're building a general-purpose vision system that has to handle shifting domains, adopt an MLLM router over heterogeneous backbones like ARMDIL, because adding a domain becomes a prompt change instead of a retraining run.
  • If you're fine-tuning a model for social or audio-visual reasoning, start with a Vanilla SFT baseline and skip the CoT pipeline until it proves a measurable win, because the current evidence says reasoning traces cost more and don't help.
  • If you're deploying MLLMs in scientific or long-video workflows, gate on uncertainty behavior and memory, not accuracy, and watch for memory-augmented architectures, because the lossy-summarizer ceiling is the next bottleneck the field has to break.