Skip to content

Beyond Chain-of-Thought: Latent Reasoning, Search Scaling, and AI That Picks Its Own Problems

#llm-reasoning #latent-reasoning #search-scaling #automated-theorem-proving #ai-agents

Four things landed in the same week. A paper that shows chain-of-thought can ace a reasoning benchmark while learning a shallow trick. A study that treats search budget like a first-class resource in autonomous research. A theorem-proving benchmark built from actual STOC and COLT papers. And an open research system that has produced 453 manuscripts and started proposing its own conjectures. On their own, each is incremental. Together they trace where AI reasoning is heading: from producing answers to producing research agendas.

Same scores, two different kinds of thinking ​

Every reasoning benchmark in circulation assumes that a high score means the model is reasoning. The ProsQA-Ext results call that assumption into question. The group took a single GPTNeoX backbone and trained five variants from scratch on an extended multi-hop reasoning task: a vanilla model, a chain-of-thought (CoT) model, a pause token model, and two latent-reasoning models trained end-to-end with no intermediate traces.

All five reach strong in-distribution performance. That's where the similarity ends. When the test set shifts to problems with more hops than anything in training, the vanilla, CoT, and pause token models collapse. They had been leaning on local graph features, matching patterns near the question rather than propagating information across the graph. The latent variants keep working, and their internal activity lines up with forward reachability propagation on the graph.

VariantComputation between input and answerIn-distribution performanceLonger-hop generalization
VanillaNoneStrongPoor
Chain-of-ThoughtToken-level reasoning tracesStrongPoor
Pause TokenFixed number of filler tokensStrongPoor
Latent (A)Hidden-state recurrence, no tracesStrongGeneralizes
Latent (B)Hidden-state recurrence, no tracesStrongGeneralizes

The practical read: a model can produce articulate reasoning traces and still be doing feature matching. Depth generalization is a property of the computation the model actually performs, not the format of its output. A CoT model writes "step 1, step 2, step 3" and looks thoughtful. On ProsQA-Ext it was pasting together local pattern matches the whole time.

The recurrent circuit that actually reasons ​

The latent models didn't generalize by accident. Causal interventions and circuit analysis localized a sparse recurrent search circuit in the bottleneck latent model: one attention head retrieves graph relations, an MLP and the residual stream update the reachability state across recurrent steps, and multiple attention heads perform candidate matching. That maps cleanly onto a forward-reachability algorithm on a graph: retrieve edges, update the frontier, check the target.

This is the kind of result interpretability work should produce. A localizable circuit, a matching algorithm, and a crisp prediction about where it fails. The bottleneck matters here. Forcing the model to compress its reasoning into a fixed-width latent state is what makes the recurrence reusable across depths. Token traces let the model off the hook because there's always room to paste one more local match onto the sequence.

Quick Take: What generalizes is the underlying computation a model learns, and in this setting latent recurrence beat token-level reasoning at out-of-distribution depth.

Search scaling: spending test-time compute like a researcher ​

Reasoning doesn't only happen inside a single forward pass. In autonomous agents it also comes from how much computation you allow around the task: more candidates, more iterations, more dead ends walked before committing. The quantitative factor-mining study makes this concrete. Fifty tasks, each grounded in a financial research report, each requiring the full loop: read a hypothesis, implement it as a factor in code, evaluate it, refine it, repeat. Nine models, from weak to strong, run this loop under varying search budgets.

Three findings stand out. Initial performance tracks model capability more than search depth, but deeper search narrows the gap between models. Give a weaker model enough budget and it approaches what a stronger model gets in a single pass. Model grafting explains why: transfer the early research state from one model to another, and the final result tracks the grafted state. The first steps commit the trajectory. Parallel search beats sequential search at the same iteration budget, consistent with broader coverage of the search space.

The trajectory analysis carries the mechanism. Higher-performing models used extra budget to diagnose failures, revise search direction, and drop candidates that violated the intended economic hypothesis. Weaker models kept refining candidates that were never viable. Capability buys you a better starting point and better failure diagnosis. Budget buys you coverage. Neither substitutes for the other.

A benchmark that treats proofs like research work ​

TCSAlgBench takes a different approach to evaluation. Instead of competition math, it builds theorem-level challenges from 138 STOC and COLT 2026 papers, 398 problems total. Competition math is saturated territory; the open question is whether models can do research-level reasoning, and this benchmark measures that directly.

Each task ships with paper-specific context, preserves computational assumptions and quantitative guarantees, and withholds constructions when discovering an algorithm is part of the task. Prover systems receive the theorem statement and access to cited prior work. The pipeline is refreshable, so new batches can be generated as new papers appear. That's a property worth stealing for any serious eval.

Under direct inference, coverage is low. The strongest model configuration, GPT-5.6 Sol max, reaches 23.6% verifier-accepted coverage after ten rounds of prover-verifier discussion. In the separate agent comparison using GPT-5.5 xhigh, decomposition beats plain discussion, and agentic planning tops out at 25.4%. Those two numbers come from different evaluations with different models, so don't put them in the same race. Put them against the difficulty instead: fewer than one in four theorem-level challenges from real papers gets a verifier-accepted proof. That's the honest state of automated proving right now.

Verifier acceptance is the right gate. Without it, a proof is just confident prose, and LLMs produce confident prose in volume. The interesting result isn't the 23.6%; it's that discussion and repeated sampling reliably improve coverage, which means test-time compute works on proofs the same way it works on math word problems.

When the agent picks the next problem ​

AI Has Taste, a project by UCSI University researcher Zeng Zaijian, is attempting the step the other three works stop short of. TCSAlgBench measures proof discovery. Search scaling measures how agents spend budget inside a fixed research loop. AI Has Taste is automating what comes before and after: which problems to attack, what to do when a route fails, and whether the residue of a failure becomes the next question.

Zeng's account of the turning point is understated. He fed a conjecture from a paper to a model, expecting it to try to prove it. The model built a counterexample instead. He didn't trust it, so he wrote to the paper's author and asked. The author confirmed the counterexample was correct. That exchange is what made him commit to the project. If a model can attack a new conjecture and survive verification against the person who wrote it, it can participate in the research loop, not just demonstrate one.

The system separates research from criticism structurally. A Supervisor / Screener selects candidate problems and looks for the cheapest decisive experiment. A Researcher produces proofs, counterexamples, and exact computations. A Fresh-context Reviewer attacks frozen propositions, key lemmas, and certificates in a clean context without having generated them. A Writer / Release Checker scopes the final claims and audits citations, builds, and release materials. The roles run across parallel machines, sharing frozen propositions, failure logs, code, and evidence. The design principle is stated plainly in the README: criticism is part of the engine, not an afterword.

The flagship case study is a failure. On a fractional Gaussian trace problem, the system tried monotonicity and one-sided pairing routes, and precise counterexamples kept appearing. A conventional benchmark would record FAILED and move on. This workflow analyzed what the counterexamples broke, constructed ten correction structures around a fixed negative defect, and proposed a new unifying boundedness conjecture. The new conjecture then got attacked in turn by fresh-context review and exact computation. Still unproven. The original question wasn't solved, but the failure was compressed into a sharper, falsifiable question, and that's a type of progress no current benchmark measures.

453 research manuscripts in the public archive 2,312 pages across the corpus, 223 manuscripts from the BigWIN2 pipeline alone 6 submissions explicitly classified as AI-proposed conjectures 1,004 archive entries scanned; the unified boundedness conjecture remains open

The six proposed conjectures span combinatorics, graph theory, number theory, and analytic problems, including a nonsingularity conjecture on infinite binary-kernel matrices, a coloring conjecture on Young diagrams, and the fractional Gaussian trace defect conjecture that grew out of the failure. What makes these worth attention isn't the act of guessing. It's that each one is written as a precise definition with supporting evidence, known attack surfaces, and what a counterexample would look like, with computation and source code left behind for verification.

The project is also honest about its limits. 453 manuscripts is not 453 peer-reviewed papers. Agent-internal review is not community peer review. Finite computation over finite cases does not generalize automatically to infinite statements. Human experts, including Malaysian Academy of Sciences fellows Ong Seng Huat and Kurunathan Ratnavelu and China University of Geosciences professor Xiong Yonghua, have verified and discussed specific problems, proofs, and computations, not the whole corpus. A machine red team plus human spot checks at critical nodes is the right shape for automated research.

The common thread: from answers to agendas ​

Each of these works answers a different question. ProsQA-Ext asks what kind of computation a model actually learned. Search scaling asks how a model spends budget across a research loop. TCSAlgBench asks whether proofs from real papers can be discovered and accepted by a verifier. AI Has Taste asks what happens when the model also chooses the questions and converts failures into new problems.

The unit of evaluation is moving from "did the model answer this question" to "did the model advance a research agenda." That has practical consequences for anyone building eval harnesses. A fixed question set with a pass/fail label can't capture whether an agent reformulated a dead route into a stronger conjecture. Verifier-based benchmarks like TCSAlgBench are a step in the right direction because acceptance stops the model from grading its own homework.

The human role shifts too. If systems produce hundreds of manuscripts and choose which directions to pursue, the researcher becomes a PI: setting research boundaries, choosing the problem pool, deciding what deserves continued investment, and deciding when to stop. That's a different job description than writing proofs by hand, and it's the one that matters as these systems scale.

What trips people up ​

Reading these four sources together, a few mistakes stand out.

Trusting in-distribution scores as evidence of reasoning. The CoT model in ProsQA-Ext solved its training distribution while learning local graph features. If your eval only has fixed hop lengths, you'll measure the wrong thing. Add an out-of-distribution split with longer reasoning chains, or you're rewarding pattern matching.

Accepting generated proofs without a verifier. TCSAlgBench's strongest configuration still fails more than three out of four theorem-level challenges. If you evaluate proof generation with a rubric instead of an automated verifier, you're measuring prose quality, not proof correctness. A proof that reads well and a proof that verifies are different objects.

Confusing search budget with search quality. The factor-mining study found parallel search beats sequential search at the same budget, and the early research state shapes the final outcome. Throwing more iterations at a weak model produces diminishing returns when the first steps committed to a bad direction. Spend budget on breadth early, and keep the early trajectory capable.

Counting volume as validation. 453 manuscripts sounds like a research lab. It's not 453 accepted results, and the project says so. Internal agent review is like a single author reviewing their own work: useful, insufficient. Demand external verification at the critical nodes, which is exactly the gap the fresh-context reviewer and human expert checks are designed to cover.

One thing to remember ​

Generalization comes from the computation a model actually performs, not from the format of its output. Token traces can deceive. Latent recurrence can generalize. Search budget helps most when the search is parallel and the early trajectory is sound. And verifier acceptance is the only honest gate for proofs. Watch which of these claims survive contact with scale.

The Bottom Line ​

If you're building a reasoning eval, add an out-of-distribution depth split. Your CoT-capable model will look strong in-distribution and collapse on longer hops, and you want to catch that before you ship the benchmark.

If you're running autonomous research agents, allocate budget to parallel search and protect the early research state. The first steps commit the trajectory, and breadth beats depth at equal iteration counts.

If you're assessing claims about AI for mathematics, gate everything on verifier acceptance and human expert spot checks. Manuscript volume is noise. Expect refreshable, verifier-based benchmarks like TCSAlgBench to become the default evaluation within a year, and expect "which problem did the agent choose next" to become the headline metric after that.