1. Context Language Models
Agent harnesses usually manage the model’s context through fixed rules such as summarization or compaction. Researchers at Meta and collaborators propose Context Language Models (CLMs), which treat the live context as a file the model can edit freely, deciding what to keep, rewrite, or remove.
Context as a read-write file: The model edits its own context with code, writing loops that prune irrelevant results, defining functions it reuses, and tracking subagents. Recursive language models place a large input in an external variable the model can read, but they do not let it edit its live interaction context. Because several agent contexts can coexist as files, the approach extends to multi-agent systems.
Works zero-shot: Built from existing models with no training, CLMs reach 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus than state-of-the-art context-management strategies, 5% higher scores with 59% fewer FLOPs on the 12-hour EdgeBench, and 65% more improvement for the same compute on a 24-hour agent-swarm task across six repositories.
Learnable in context and in weights: Natural-language instructions evolved through a skill-optimization loop raise held-out accuracy by up to 35.9 points on a context-management task. An online RL method improves Qwen3.5-9B on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs.
Why it matters: Edits in the middle of the context break prefix caching, so the authors co-design Suffix Cache Reuse, which cuts server-side compute by 35% against standard SGLang at matched performance. Context management can move from the harness into the model without raising serving cost.
2. JAZ
Long-term memory and self-improvement are usually built as separate systems around the agent loop. Researchers from MIT CSAIL built JAZ to test how far a minimal harness, little more than the agent loop itself, can go on the tasks those systems target.
One primitive: JAZ exposes a single LLM-based primitive, invoke, which acts as a function whose implementation the LLM writes at runtime each time it is called. Built-in hooks let the programmer apply constraints and monitor the run.
Everything is a variable: The LLM can write arbitrary executable code, including recursive calls to invoke, so subagents are the default. All inputs to invoke and the full interaction history live as variables in the code environment, where the model can read and transform them with code.
Beats specialized systems: With prompting only and no manually designed tools, memory system, or file system, JAZ outperforms Letta (MemGPT) by 8% at half its cost on the recall-heavy portion of StuLife, a long-horizon benchmark that requires recall far beyond the context window. On AppWorld, it beats ACE, a continual self-improvement method, by 4% at lower cost.
Why it matters: Memory and self-improvement become code the agent writes inside its own loop. If you build custom harnesses, this is evidence that a smaller surface can match heavier infrastructure on these workloads.
3. Agensh
Multi-agent harnesses usually depend on a central orchestrator that assigns tasks and coordinates workers, and that orchestrator limits how many agents the system can use. Microsoft Research introduces Agensh, a self-organized multi-agent harness with no central orchestrator, and scales it to 1,024 coding agents.
A shared cooperation loop: Every worker runs the same loop asynchronously. It gathers context, claims and self-assigns a sub-task, takes action and shares findings, verifies the result, and merges progress.
Organization infrastructure: Three components support the loop. A shared workspace holds proposed, ongoing, and completed work, a message interface lets workers communicate, and a shared context retains reusable findings and work intentions.
Scaling the number of agents: On the five hardest ProgramBench tasks with GPT-5.6-sol, going from 1 to 128 agents raises the mean final test-pass rate from 19.31% to 28.78%, about a 49% relative improvement, and larger teams reach a given pass rate sooner. On building pandoc from scratch under a 6-hour budget, going from 1 to 1,024 agents raises the test-pass rate from 33.89% to 55.06%.
Why it matters: Worker trajectories show forms of cooperation that the agents start on their own and that become standard practice as the organization grows. The authors treat the number of agents as a new scaling dimension, which is most useful for complex tasks under hard latency limits or time budgets.
4. AutoGym
Training agents with RL requires a gym, meaning a task, an executable environment to attempt it in, and a verifier that reliably separates success from failure. These gyms are still built by hand, saturate as models improve, and get exposed to contamination. Researchers from Amazon AGI present AutoGym, which generates complete gyms from a minimal domain seed or from prior model trajectories.
Blueprint first: AutoGym specifies the valid solution space, environment requirements, and verification criteria before it builds the environment. Solvability is settled during construction, so no unreliable LLM judge has to decide correctness afterward.
Difficulty you can steer: Explicit generation parameters control task topology, interaction depth, capability axes, question obfuscation, and distractor composition. Single-pass synthetic tasks tend to be hard only in their phrasing, and these parameters target actual difficulty.
Active curriculum: Performance-informed calibration shifts the distribution over those parameters as model capabilities change, and failure analysis feeds recurring capability gaps back into the generation parameters.
Why it matters: Across productivity and temporal-reasoning settings, AutoGym produces gyms spanning the capability spectrum, including instances that challenge frontier models. Building RL environments is becoming a core AI engineering skill, and this gives a recipe for gyms that keep pace with the models trained on them.
Message from the Editor
Our new DAIR Academy lab, Build a Website with a Pi Coding Agent, teaches you to build and refine a responsive one-page website by directing a Pi coding agent. Across 5 labs, you add a practical feature to the same project each time, from the first hero section to responsive styling, an interaction, accessibility, and final polish.
5. CASD
Search-based prompt optimizers such as GEPA propose edits, run fresh rollouts, score them, and keep only the edits that improve a validation metric. Researchers from Microsoft show that this loop may be unnecessary when you already have a corpus of agent trajectories.
One pass over the logs: Coding-Agent Skill Distillation (CASD) gives an off-the-shelf coding agent a static corpus of trajectories. The agent writes and runs analysis code to compute corpus-wide statistics, find systematic failure modes, inspect representative episodes, and distill the findings into behavioral rules in one prompt. It needs no environment access and no validation data.
Reflection scope: Search-based optimizers reflect on small batches of trajectories at each step, so failures that only show up across the whole corpus stay invisible to them. CASD reads the whole corpus, so it can find those failures.
Better and cheaper: Across ALFWorld, tau2-bench retail and telecom, and SpreadsheetBench-Verified, a single CASD pass improves the unoptimized baseline by 16.6 points on average, against 10.9 for GEPA and 5.3 for SkillOpt. Each optimized prompt costs about $1.60, over 22x cheaper than validation-gated search.
Why it matters: Even when GEPA and SkillOpt also get validation data and unrestricted environment access, CASD stays ahead on two of four benchmarks. If your agents already produce logs, pointing a coding agent at them is a cheap first step before running a search-based optimizer.
6. Taste-Bench
On long-horizon tasks, decisions such as which hypothesis to test or which implementation to build on determine how the whole run turns out. Researchers at Microsoft and collaborators call the ability to make these decisions well an agent’s taste, and they built Taste-Bench to measure it.
Decision forks: Each question shows a point in a trajectory where several directions are open, and one leads to a better outcome. The model picks a direction without seeing what happens after the fork.
Mined automatically: Forks come from engineering and research runs, either from parallel attempts at the same task or from detours inside a single trajectory. Filters drop forks that are guessable from the options alone and forks that cannot be decided from the context, so no human annotation is needed.
Just under 60%: The best model answers 59.7% of questions correctly. Forks whose deciding evidence appears later in the trajectory are much harder for every model, and a larger reasoning budget does not improve accuracy.
Why it matters: Taste can be trained. The authors distill the judgment of a teacher that has seen the outcome into a student model, which makes better decisions on unseen tasks and improves end-to-end success on held-out SWE-bench Pro tasks. Benchmarks that score only end-to-end success miss this capability, and for recursive self-improvement, an agent’s choice of direction also matters.
7. Jev-Mem
Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller.
Three planes: A System-One control plane makes the fast decisions, a structured multi-relational memory plane stores semantic, temporal, causal, and entity relations, and a System-Two reasoning plane handles slower, deliberative reasoning.
The controller runs memory: When storing, the controller assigns memory types and relations. When reading, it handles query routing, retrieval-budget allocation, graph traversal, candidate scoring, and adaptive stopping. The LLM is called only for complex reasoning and writing the final answer.
Faster and more accurate: On LoCoMo, Jev-Mem scores 0.777 overall with an LLM judge, an 11.0% relative improvement over the strongest baseline. Memory construction takes 158 seconds, 6.6x faster than the fastest competing memory system, and average query latency drops 36.7% to 0.93 seconds.
Why it matters: Memory latency adds up across every step of a long-horizon agent. Moving routine memory decisions to a small controller lowers that cost while improving answer quality, and the same split could apply to other harness decisions that a frontier model currently makes.
8. Critical-State RL
Salesforce AI Research’s Critical-State RL finds the one call in a multi-turn tool-use interaction where training helps, since reward variation that depends on later turns often reflects downstream randomness instead of the current action. It uses nested sampling to separate action-dependent reward variation from continuation noise, then trains only the selected call with contextual-bandit updates. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points, while training the alternative turn leaves accuracy flat or worse.
9. SkillGym
SkillGym turns human-written agent skills into 2,756 verifiable training environments across 12 categories, each with code-based checkers, and collects 8,364 successful trajectories for fine-tuning. Under Claude Code, fine-tuning Qwen3.5-35B-A3B adds 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, reaching 51.47%, above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model still beats the base model with skills in context.
10. AutoCompact
AutoCompact trains a coding agent to decide when to compact its context, what working state to keep, and how to continue afterward, as part of its own policy. A judge reviews the base agent’s compaction decisions and replaces flawed ones before they execute, and the corrected trajectories are used for SFT and then for RL that optimizes coding and compaction together on task success. Pass rates rise by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified, and the gains hold with both a 256K window that never overflows and a 16K window that falls back to forced compaction.








