Survey
Beyond the Agent Harness: Why do we need continual learning Agent Companion Models?
We believe the next bottleneck for AI agents is not building general-purpose systems that achieve a moderate success rate, but developing agents that match human professionals’ capabilities across domains and continuously improve after deployment, which can consistently approach ~ 99% reliability as human preferences and task environments evolve.
The AI-Agents’ large-scale deployment is still uncommon, and the agents’ task completion quality is a frequently named barrier. In a 2025 Capgemini survey of 1,500 executives, 2% of organizations reported agent deployment at scale [1]. Not surprisingly, trust in fully autonomous agents fell from 43% to 27% in one year. And, in LangChain’s survey of 1,340 practitioners, 32% named quality of work as their main barrier to production [2]. Reliable AI-Agent post-deployment operation is a critical issue we cannot dismiss now.
Source: Capgemini Research Institute, 2025
How does an agent learn after deployment today?
Today, the standard way to build an AI agent now seems simple:
Agent = Base Model + Harness [3].
The model, which is a pretrained and/or post-trained foundation model, reasons over its context window and decides which tools to call.
The harness, which is all the code and configuration around the model, runs the agent loop and provides the system prompt, context management (retrieval, memory and compaction), tool execution, agent orchestration (subagents and handoffs) and safety guardrails (input and output checks, permission and approval gates). The model’s parameters are frozen at deployment.
Then how can the AI-Agents keep improving under this agent builder setup?
Unless engineers change the harness, an agent built this way can use what happened in earlier tasks only by placing it in the context window.
Alternatively, the pre-built harness can iteratively generate new model architecture, training codes to improve the base model.
All together, the agents’ architecture, prompt configs, memory updating modules, tools, and even the base model itself can be improved based on post-deployment data (agent-env interaction trajectories with a feedback).
Sometimes new researchers like to call this Recursive-Self-Improvement (RSI), which is category of methodology that let the AI agents improve according to a predefined goal (there could be cases we do not need to precisely define the goal).
The 3rd component: Agent-Companion Models (Agent-CMs)
Despite the (Base model + Harness) way of building AI-Agents, we observed that many real-world deployed agentic systems also contain a third component, a parametric model that runs alongside the base model. We call it an agent-companion model (Agent-CM).
In this post, we report this observation and argue the following:
Post-deployment experience can raise the agent’s performance ceiling.
If we continually update the agent’s companion model (Agent-CM) using its post-deployment experience, the model can internalize what the agent has learned rather than merely using that experience as temporary context. This creates the potential for a higher performance ceiling on real-world tasks.But continual updating creates a generalization VS. runtime efficiency dilemma.
A larger and more general model may learn and retain more, but repeatedly training and running such a model can quickly become too expensive. This is where online plasticity becomes important: the Agent-CM needs to efficiently adapt to new experience while preserving broad capabilities and remaining practical at runtime.The key open problem is how to achieve this plasticity efficiently.
How can we continually update Agent-CMs with minimal training compute and minimal inference overhead, while retaining the capabilities of today’s strong companion models? This remains largely unsolved. Recent continual-learning approaches increasingly explore alternatives to conventional gradient-based updates—or even to backpropagation itself—but progress toward models that can match the capabilities of today’s widely used Agent-CMs remains slow.
By post-deployment experience, we mean the trajectories of earlier tasks (model outputs, tool calls and tool results), their outcomes, and user feedback/preference/corrections.
In summary:
Parametric continual learning has a higher potential ceiling than in-context continual learning.
Agents that continually update their Agent-CMs can internalize post-deployment experience, rather than relying on an ever-growing context to reuse past experience.The key bottleneck is the trade-off between generalization and runtime efficiency.
More capable Agent-CMs offer stronger generalization, but are also more expensive to continually update and run.Online plasticity of Agent-CMs may be the key to resolving this trade-off.
So far, what have we found?
| Role | Work | What is learned | Parameters updated after deployment |
|---|---|---|---|
| Memory manager | MemoPilot [4] | Policy for updating a frozen player’s memory | No; the policy edits external memory with fixed weights |
| Verifier | Agentic Verifier [5] | Execution-based tests and candidate selection | Not reported |
| Reward model | MagicGUI-RMS [6] | Step-level evaluation and correction of GUI actions | Partially; retrained over successive rounds of collected trajectories |
| Model router | RouteLLM [7] | Strong/weak model selection from preference data | Not reported |
| Critic | Rubric-Supervised Critic [8] | Rubric-based judgments from sparse real-world outcomes | Not reported |
| World model | Web Agents with World Models [9] | Web-state changes following candidate actions | Not reported |
This table lists parametric models that run alongside the base model in recent agentic systems. The list is not exhaustive, and as you can see, different model devs are actually driven by different problems that an agentic system may encounter in real world tasks:
memory: how to consolidate knowledge/skills and how to reuse to improve new tasks.
verifiers and critics/reward models: provide correction signals to improve the agentic system
model routing: which model best fits the task and costs less.
guardrail models: handle safety and compliance to constraints.
agent world models: predicts help for a better planner.
In all of them, a function that a harness could implement with prompts or hand-written rules is implemented by a trained model (the companion model), and the base foundation model’s weights are not changed.
This agrees with the compound AI systems view [10], in which a deployed system combines several models and other components.
A companion model can be trained for one domain, and an agent can use more than one.
Do these agent companion models learn after deployment?
Mostly they do not.
A foundation model is trained only on data collected before its training cutoff. Likewise, a companion model that is trained once is trained only on data collected before deployment. In both cases the parameters are not updated on deployment data.
Will in-context learning be the rescue?
When the model’s parameters and the harness code are both fixed, an agent acts on task as follows:
1
where is the input of task (the user’s request and the current environment state), is the deployment experience from earlier tasks, is retrieval over the agent’s long-term memory, is the harness, which assembles the context window and runs the agent loop, and denotes the actions the agent takes on task . Learning of this kind is in-context learning.
Experience can also be used by engineers, who edit the system prompt or add a tool, a hook or an eval after a failure. This is harness engineering, and it costs forward deployment engineering (FDEs) time for each failure.
For example, consider an agent that processes expense reports. A user corrects it once: receipts above a certain amount need a second approval, and the harness writes the correction to long-term memory. On a later task we can ask “what if this memory entry is not retrieved?” Under Equation (1) the agent repeats the error, even though the correction is in .
CL-Bench [11] measures continual learning of AI-Agents. It runs a stateful system through a sequence of instances within a task and compares it with a stateless baseline, which is the same system reset between instances. It reports a normalized gain, which is the difference between mean stateful and mean stateless reward divided by the difference between maximum reward and mean stateless reward. On the public leaderboard accessed on September 24, 2026 [12], the highest average normalized gain over the six tasks is 0.241 out of a maximum of 1, shared by ICL with the full conversation history and by Claude Code, both using Claude Sonnet 4.6. Entries that add a memory layer, a notepad or a context playbook (Mem0, ICL Notepad, ACE) range from 0.077 to 0.224. The authors report that the evaluated systems over-rely on recent instances, fail to reuse relevant earlier experience, and retain early errors after corrective feedback. Likewise, related studies report that agent memory systems can introduce incorrect generalizations and outdated beliefs [13], [14] and can lose constraints during compaction [15].
As we can see here, in-context learning does improve performance over a sequence of instances, but the best reported gain is about a quarter of the maximum.
What changes if the agent-companion model keeps learning?
Adding a companion model gives two further cases:
2
3
Equation (2) describes most entries in Table 1. In Equation (3) the companion’s parameters are updated with deployment experience, which is hard in practice and we will discuss it in the next section.
Periodic fine-tuning on production traces and updates after every task are both cases of Equation (3); we use online plasticity for any update that uses deployment experience while the agent is in service (online).
Let denote performance on later tasks, and let , and denote the sets of agents of the form Equation (1), Equation (2) and Equation (3).
By upper bound we mean , given unlimited deployment experience and a fixed inference budget of tokens per task, which includes the companion’s inference.
We first compare and . In Equation (1), the agent can read over several steps of the agent loop, through retrieval, tool calls and subagents, and every token it reads counts against the budget. So depends on only through at most tokens.
In Equation (3), also depends on through . An agent in whose companion output is constant is an agent in . Therefore , and
4
The inequality is strict for some deployments. Specifically, suppose the correct actions on task depend on information in that cannot be written in tokens. No agent in takes the correct actions in this case, while an agent in can do so if its companion model has enough parameters.
We then compare with . In Equation (2), is trained only on data collected before deployment, so depends on only through at most tokens, as in Equation (1). A frozen companion model can improve performance at the start of deployment.
This argument shows that better agents exist in .
Why is it hard to make Agent Companion models keep updating?
If a learning agent-companion model has a higher upper bound, we next ask what model it should be. We consider three properties.
Task breadth is measured by performance on tasks not seen in training.
Online plasticity means that the parameters are updated with deployment trajectories and feedback, and it is measured by the gain on later tasks without loss on earlier ones, which is the stability–plasticity dilemma [16].
Runtime efficiency is measured by tokens, compute, memory and inference latency.
We call the trade-off among the three the agent companion-model trilemma.
Existing solutions satisfy two of the three.
A frozen frontier model, which is accessed through a provider’s API, performs well across many domains and requires no training by the deployer, but its parameters are not updated on the deployer’s data.
A large locally deployed model can be fine-tuned periodically and also performs well across domains; but the training procedure, version evaluation are tedious to perform.
A small fine-tuned specialist is the most common solution now, but it is cheap to serve and to update online. It is trained for one task, however, so each additional task requires its own data pipeline, training objective and evals, which adds to the maintenance cost.
Can online plasticity resolve the trilemma?
The trilemma holds under one assumption: a model’s task breadth is fixed at deployment time and comes from parameter size and pretraining dataset size. Under this assumption a model needs more parameters to perform more tasks, and more parameters cost more compute.
This suggests another possibility: a small companion model could learn additional related tasks through deployment experience, rather than require a larger pretrained model or a separately designed model for each task. It need not perform every task in its domain at deployment time. It needs a learning procedure that can reuse experience across those tasks.
Online plasticity may therefore change the breadth–efficiency tradeoff. Whether it does depends on how much additional task coverage can be learned at a fixed parameter budget, how well earlier learning is preserved, and how much computation the updates require.
ICL remains necessary. A new instruction or tool schema can affect the next action before any parameter update is available. The two mechanisms can serve different purposes:
Fast adaptation through context: incorporate new instructions, current state, and recent feedback immediately.
Slower consolidation into parameters: learn reusable patterns from selected deployment experience, so later tasks benefit without requiring the same source records in every context. These updates can be stored as customer-controlled parameter deltas or checkpoints, with versioning, regression tests, audit records, and rollback.
The distinction between rapid acquisition and gradual integration has a precedent in complementary learning systems theory [17]. In agent systems, Letta’s memory-model agenda explores a related combination: a trained model creates reusable context that downstream models consume through ICL, with parameter distillation discussed as a further possibility [19].
The remaining question is how to train a companion model for continued learning, rather than only for performance at deployment. This requires studying the architecture, training objective, and deployment data together, including what to consolidate, when to update, and what to revise or forget. That’s what we will start at TrueIntrinsics.