1. Harness Learning
An agent is defined by its model and its harness, the program that organizes model calls, tool use, and information flow. Researchers at CMU train a proposer model to revise the harness based on execution feedback, while the solver model’s weights never change.
Edits as updates: The authors frame harness adaptation as meta-learning over executable programs, where each harness revision plays the role of a weight update in gradient-based adaptation. The proposer reads a task, the current harness, and an execution report, then writes a code edit to the harness.
RL on harness scores: The proposer is trained with reinforcement learning, and the reward is the revised harness's task score. At test time, it refines the harness on a new task through successive executions, without updating parameters.
Small proposer, broad transfer: A trained 4B proposer beats its 35B teacher at single-step revision on Reasoning Gym, including task families it never saw in training. A proposer trained on HotpotQA keeps improving harnesses on MuSiQue and 2WikiMultihopQA, and policies trained on single revisions continue improving harnesses over multiple rounds.
Why it matters: Improving agents by editing harness code instead of weights is becoming its own AI engineering skill, and this work shows the skill can be learned and transferred to unseen tasks. It points toward agents that turn accumulated experience into general improvements without retraining the underlying model.
2. CorpusMap
Agents that search large document collections usually see the corpus as a flat set of files, so a relevant document gives no hint of how it connects to others. Researchers from Microsoft and colleagues introduce CorpusMap, a navigation layer built around the entities that recur across a collection.
Entity pages: CorpusMap resolves mentions of the same entity across documents in advance and gives each recurring entity a page that aggregates what is known about it and links to every document that refers to it. The original documents stay in place.
Built once, reused per query: The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again. The links are built offline, so they are shared across queries instead of rediscovered at inference time. The map can be built without LLM calls and updated as new documents arrive.
Better answers, fewer tokens: Across 7 models and three benchmarks, answer quality rises 6.4 to 11.7 points over raw-corpus agentic search while input tokens drop 34% to 57%. CorpusMap also beats an LLM Wiki layer and three other navigation layers.
Why it matters: Enterprise questions often need evidence spread across emails, reports, and contracts in different folders. A precomputed entity graph gives agents structure to follow, which reduces both missed evidence and repeated searches.
3. VeriHarness
A common practice with long-horizon agents is to sample several rollouts and trust the answers they agree on. Researchers at Google show that agreement can hide shared errors, while disagreement often points to the correct alternative, and they built VeriHarness to verify rollouts with the same base model.
Verifier from the generator’s model: VeriHarness gives the model the generator uses a workspace, evidence tools, and reusable verification skills. It needs no reference answers or grading rubrics at test time.
Two checks: A disagreement resolver checks competing claims against environment evidence. A consensus challenger tests claims that every rollout shares and looks for requirements they all missed. Their findings guide which artifact to select and how to revise it.
Results: Across five long-horizon workspace benchmarks and two frontier models, VeriHarness reaches the highest selection scores among the baselines tested. With evidence-backed revision, it adds 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. Its verification skills also improve from feedback on failures.
Why it matters: Majority voting discards the minority answer, which is sometimes correct. The authors release about 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support further work on agentic verification.
4. HERMES
Software engineering agents on long-horizon tasks keep reconstructing program state spread across source files, configurations, tests, and dependencies, which leads to long interaction histories and context explosion. HERMES is a harness that turns repository components into active participants in the task.
Dev-Primitives: Each repository component is paired with a resident LLM that knows its own implementation and dependencies. The component gets an agent-native interface for natural-language reasoning, communication with other components, and localized self-modification.
Activation and diagnosis: A dependency-aware step decides which components to activate for a task, and a diagnosis step maps test failures and execution evidence back to the components that need changes.
Same model, different harness: On whole-repository migration, GPT-5.6 Sol goes from 6.5% to 31.0% when Codex is replaced with HERMES, at the same model and effort setting. Across four software engineering benchmarks, HERMES beats matched baseline harnesses by 12.4% on average.
Why it matters: With strong activation and diagnosis models, HERMES running Qwen3-8B Dev-Primitives stays within 4.5% of an all-GPT-5.6 Sol configuration on all four benchmarks and cuts inference cost on Terminal-Bench 4.0 by 26.2%. Most of the migration gain comes from the harness, so a harness built around your own codebase is worth the engineering effort.
Message from the Editor
Our new DAIR Academy lab, Build a Website with a Pi Coding Agent, teaches you to build and refine a responsive one-page website by directing a Pi coding agent.
5. MASS
Recursive self-improvement on open-ended tasks like research runs into a supervision problem because the outputs can exceed what even human experts can reliably assess, and there is no external checker. Researchers at Sakana AI propose Multi-Agent Self-Supervision (MASS), in which a single base model acts as the agent, evaluator, and optimizer.
Workflows as the search space: The base model proposes multi-agent workflows, runs them, and grades the results. An evolutionary search constrained by structural guardrails keeps the best-scoring workflows and settles on distinct roles and information routing for each task.
Alternating cycles: Each cycle pairs that workflow optimization with supervised fine-tuning on the model’s own multi-agent traces. The improved model then starts the next cycle as a better optimizer and grader, so the loop can keep bootstrapping the model’s capabilities.
Results: Two MASS cycles on Qwen3.6-27B give 1.2 to 1.6x higher performance per output token on four open-ended public benchmarks. A student trained on multi-agent traces also beats a single-agent student trained on 1.4x more tokens.
Why it matters: Self-improvement loops usually depend on an external verifier, so open-ended tasks with no checker are left out. MASS shows that one model can supply its own supervision signal by learning orchestration and bounded subagent execution from multi-agent trajectories.
6. Sharpening Tax
A common hypothesis holds that RL post-training only sharpens behaviors the base model already has, raising single-shot accuracy at the cost of solution coverage. Researchers from Meta Superintelligence Labs test whether this holds for agentic tasks, where multi-turn tool use might need capabilities learned during post-training.
Base models as agents: Pre-trained LLMs with a light inference harness work as capable agents. Post-trained models win on pass@1, but at large K the base models often solve tasks their post-trained versions never solve on BFCL v4 multi-turn, ACEBench, and WebShop.
Why coverage drops: Post-training pushes each task toward always solved or never solved. Consistency and sampling efficiency increase, while coverage decreases.
Measuring the cost: The authors propose the Sharpening Tax, a metric for the test-time scalability lost after post-training. Across 42 base and post-trained cases from four model families and three benchmarks, the tax shows up in most settings, grows with model size, and can be estimated from a few rollouts.
Why it matters: Their fix, posterior-tempered group sampling (PTGS), sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax than a fixed-temperature baseline and also raises pass@1. If your agent relies on repeated sampling at test time, a base model may cover tasks that your post-trained model has lost.
7. VERA
Long-horizon agents need environments that can be resumed at any stage and scored from observable evidence, yet most environments score only the final outcome. Researchers from NVIDIA present VERA, which builds such environments at scale and uses them to update both the model and its harness.
Verifiable environments from trajectories: VERA turns initial trajectories into more than 9,000 restartable sandboxes. An agent writes rubrics and executable checks, a judge verifies each sandbox, and only environments that run and can be scored from observable evidence enter the training bank.
Two update paths: VERA alternates between training the model with rubric rewards and editing the harness skills. A harness edit is kept only if it passes self-tests and adds at least 5 points on the development set, and a model checkpoint is rejected if its score drops by more than 20%.
Results: At 9B, the co-evolved agent beats the strongest single-axis baseline by 10.3 points on AutoCoWorkBench and 13.0 points on AutoMedBench. At 27B, it scores 71.6 on AutoCoWorkBench and 80.7 on AutoMedBench, transfers to unseen workflows, and keeps its general capabilities.
Why it matters: Training the harness alongside the model is a big part of owning the intelligence stack, and many frontier AI companies have started doing it. The open-source corpus of environments gives others a starting point for this kind of co-evolution on long-horizon work.
8. Agent Plasticity
Researchers from Meta Superintelligence Labs introduce agent plasticity, a measure of how efficiently an agent converts experience into gains on held-out tasks per dollar spent on learning, with model weights frozen and each run starting from a fresh context while inheriting artifacts written by earlier runs. The model with the best final score is often a different model from the one that learns most efficiently. In chess, Go, and Hex, Claude Fable 5 reaches the highest final score, while GPT-5.6 Sol gains the most per dollar; in NetHack, only Claude Opus 5.5 improves significantly, by 66 normalized points for about $1,073 of learning. Slow learners often ignore artifacts they already wrote, while faster learners reuse their artifacts and still fail when an artifact is low quality.
9. Continuous Memory Machines
Sakana AI’s Continuous Memory Machine gives a recurrent model two matrix-valued memories: a short-term store that tracks recent neuron activity and a long-term store for information needed later, with a Transformer reading and writing both at every step. Built on the Continuous Thought Machine, it beats LSTM, DNC, RMC, and CTM baselines on copy, associative recall, sorting, few-shot regression, and maze solving, and it generalizes to longer inputs than earlier memory-augmented networks. Its attention maps show it uses long-term memory for algorithmic and in-context tasks and skips it when a task does not need it.
10. MIMESIS
Agent RL setups usually let an assistant LLM play the user, which makes the simulated user overly cooperative and explicit, and a fixed GPT-5.5 agent finds tau-bench tasks easier with these users than with real people. Researchers from Meta Superintelligence Labs train MIMESIS, a 9B user simulator built from human conversations with explicit reasoning supervision and 13 behavioral patterns observed in real users, which beats Claude Opus 5 on behavioral fidelity by 13.4 points. Agents trained against the frozen simulator with multi-turn RL outperform agents trained against GPT-5.5 under all nine unseen user simulators, and a coaching step that turns the simulator’s private reasoning into token-level feedback adds further gains.








