We believe the next bottleneck for AI agents is not building general-purpose systems that achieve a moderate success rate, but developing agents that match human professionals’ capabilities across domains and continuously improve after deployment, which can consistently approach ~ 99% reliability as human preferences and task environments evolve.

The AI-Agents’ large-scale deployment is still uncommon, and the agents’ task completion quality is a frequently named barrier. In a 2025 Capgemini survey of 1,500 executives, 2% of organizations reported agent deployment at scale [1]. Not surprisingly, trust in fully autonomous agents fell from 43% to 27% in one year. And, in LangChain’s survey of 1,340 practitioners, 32% named quality of work as their main barrier to production [2]. Reliable AI-Agent post-deployment operation is a critical issue we cannot dismiss now.

Scaled AI-Agent deployment remains rare
Figure 1. Scaled AI-Agent deployment remains rare

Source: Capgemini Research Institute, 2025

How does an agent learn after deployment today?

Today, the standard way to build an AI agent now seems simple:

Agent = Base Model + Harness [3].

  • The model, which is a pretrained and/or post-trained foundation model, reasons over its context window and decides which tools to call.

  • The harness, which is all the code and configuration around the model, runs the agent loop and provides the system prompt, context management (retrieval, memory and compaction), tool execution, agent orchestration (subagents and handoffs) and safety guardrails (input and output checks, permission and approval gates). The model’s parameters are frozen at deployment.

Then how can the AI-Agents keep improving under this agent builder setup?

  • Unless engineers change the harness, an agent built this way can use what happened in earlier tasks only by placing it in the context window.

  • Alternatively, the pre-built harness can iteratively generate new model architecture, training codes to improve the base model.

  • All together, the agents’ architecture, prompt configs, memory updating modules, tools, and even the base model itself can be improved based on post-deployment data (agent-env interaction trajectories with a feedback).

  • Sometimes new researchers like to call this Recursive-Self-Improvement (RSI), which is category of methodology that let the AI agents improve according to a predefined goal (there could be cases we do not need to precisely define the goal).

The 3rd component: Agent-Companion Models (Agent-CMs)

Despite the (Base model + Harness) way of building AI-Agents, we observed that many real-world deployed agentic systems also contain a third component, a parametric model that runs alongside the base model. We call it an agent-companion model (Agent-CM).

In this post, we report this observation and argue the following:

  • Post-deployment experience can raise the agent’s performance ceiling.
    If we continually update the agent’s companion model (Agent-CM) using its post-deployment experience, the model can internalize what the agent has learned rather than merely using that experience as temporary context. This creates the potential for a higher performance ceiling on real-world tasks.

  • But continual updating creates a generalization VS. runtime efficiency dilemma.
    A larger and more general model may learn and retain more, but repeatedly training and running such a model can quickly become too expensive. This is where online plasticity becomes important: the Agent-CM needs to efficiently adapt to new experience while preserving broad capabilities and remaining practical at runtime.

  • The key open problem is how to achieve this plasticity efficiently.
    How can we continually update Agent-CMs with minimal training compute and minimal inference overhead, while retaining the capabilities of today’s strong companion models? This remains largely unsolved. Recent continual-learning approaches increasingly explore alternatives to conventional gradient-based updates—or even to backpropagation itself—but progress toward models that can match the capabilities of today’s widely used Agent-CMs remains slow.

By post-deployment experience, we mean the trajectories of earlier tasks (model outputs, tool calls and tool results), their outcomes, and user feedback/preference/corrections.

In summary:

  • Parametric continual learning has a higher potential ceiling than in-context continual learning.
    Agents that continually update their Agent-CMs can internalize post-deployment experience, rather than relying on an ever-growing context to reuse past experience.

  • The key bottleneck is the trade-off between generalization and runtime efficiency.
    More capable Agent-CMs offer stronger generalization, but are also more expensive to continually update and run.

  • Online plasticity of Agent-CMs may be the key to resolving this trade-off.

So far, what have we found?

Table 1: an overview of existing methods
Role Work What is learned Parameters updated after deployment
Memory manager MemoPilot [4] Policy for updating a frozen player’s memory No; the policy edits external memory with fixed weights
Verifier Agentic Verifier [5] Execution-based tests and candidate selection Not reported
Reward model MagicGUI-RMS [6] Step-level evaluation and correction of GUI actions Partially; retrained over successive rounds of collected trajectories
Model router RouteLLM [7] Strong/weak model selection from preference data Not reported
Critic Rubric-Supervised Critic [8] Rubric-based judgments from sparse real-world outcomes Not reported
World model Web Agents with World Models [9] Web-state changes following candidate actions Not reported

This table lists parametric models that run alongside the base model in recent agentic systems. The list is not exhaustive, and as you can see, different model devs are actually driven by different problems that an agentic system may encounter in real world tasks:

  • memory: how to consolidate knowledge/skills and how to reuse to improve new tasks.

  • verifiers and critics/reward models: provide correction signals to improve the agentic system

  • model routing: which model best fits the task and costs less.

  • guardrail models: handle safety and compliance to constraints.

  • agent world models: predicts help for a better planner.

In all of them, a function that a harness could implement with prompts or hand-written rules is implemented by a trained model (the companion model), and the base foundation model’s weights are not changed.

This agrees with the compound AI systems view [10], in which a deployed system combines several models and other components.

A companion model can be trained for one domain, and an agent can use more than one.

Learned models alongside the main agent
Figure 2. Learned models alongside the main agent

Do these agent companion models learn after deployment?

Mostly they do not.

A foundation model is trained only on data collected before its training cutoff. Likewise, a companion model that is trained once is trained only on data collected before deployment. In both cases the parameters are not updated on deployment data.

Will in-context learning be the rescue?

When the model’s parameters and the harness code are both fixed, an agent acts on task tt as follows:

at=FθF(H(xt,R(E<t))),θF fixed a_t = F_{\theta_F}\left(H\left(x_t, R(E_{<t})\right)\right), \qquad \theta_F \text{ fixed}

1

where xtx_t is the input of task tt (the user’s request and the current environment state), E<tE_{<t} is the deployment experience from earlier tasks, RR is retrieval over the agent’s long-term memory, HH is the harness, which assembles the context window and runs the agent loop, and ata_t denotes the actions the agent takes on task tt. Learning of this kind is in-context learning.

Experience can also be used by engineers, who edit the system prompt or add a tool, a hook or an eval after a failure. This is harness engineering, and it costs forward deployment engineering (FDEs) time for each failure.

For example, consider an agent that processes expense reports. A user corrects it once: receipts above a certain amount need a second approval, and the harness writes the correction to long-term memory. On a later task we can ask “what if this memory entry is not retrieved?” Under Equation (1) the agent repeats the error, even though the correction is in E<tE_{<t}.

CL-Bench [11] measures continual learning of AI-Agents. It runs a stateful system through a sequence of instances within a task and compares it with a stateless baseline, which is the same system reset between instances. It reports a normalized gain, which is the difference between mean stateful and mean stateless reward divided by the difference between maximum reward and mean stateless reward. On the public leaderboard accessed on September 24, 2026 [12], the highest average normalized gain over the six tasks is 0.241 out of a maximum of 1, shared by ICL with the full conversation history and by Claude Code, both using Claude Sonnet 4.6. Entries that add a memory layer, a notepad or a context playbook (Mem0, ICL Notepad, ACE) range from 0.077 to 0.224. The authors report that the evaluated systems over-rely on recent instances, fail to reuse relevant earlier experience, and retain early errors after corrective feedback. Likewise, related studies report that agent memory systems can introduce incorrect generalizations and outdated beliefs [13], [14] and can lose constraints during compaction [15].

As we can see here, in-context learning does improve performance over a sequence of instances, but the best reported gain is about a quarter of the maximum.

What changes if the agent-companion model keeps learning?

Adding a companion model CC gives two further cases:

at=FθF(H(xt,R(E<t),Cϕ(xt))),ϕ fixed a_t = F_{\theta_F}\left(H\left(x_t, R(E_{<t}), C_{\phi}(x_t)\right)\right), \qquad \phi \text{ fixed}

2

at=FθF(H(xt,R(E<t),Cϕt(xt))),ϕt=ϕt(E<t) a_t = F_{\theta_F}\left(H\left(x_t, R(E_{<t}), C_{\phi_t}(x_t)\right)\right), \qquad \phi_t = \phi_t(E_{<t})

3

Equation (2) describes most entries in Table 1. In Equation (3) the companion’s parameters are updated with deployment experience, which is hard in practice and we will discuss it in the next section.

Periodic fine-tuning on production traces and updates after every task are both cases of Equation (3); we use online plasticity for any update that uses deployment experience while the agent is in service (online).

Let JJ denote performance on later tasks, and let Π1\Pi_1, Π2\Pi_2 and Π3\Pi_3 denote the sets of agents of the form Equation (1), Equation (2) and Equation (3).

By upper bound we mean sup⁡π∈ΠiJ(π)\sup_{\pi \in \Pi_i} J(\pi), given unlimited deployment experience and a fixed inference budget of BB tokens per task, which includes the companion’s inference.

We first compare Π1\Pi_1 and Π3\Pi_3. In Equation (1), the agent can read E<tE_{<t} over several steps of the agent loop, through retrieval, tool calls and subagents, and every token it reads counts against the budget. So ata_t depends on E<tE_{<t} only through at most BB tokens.

In Equation (3), ata_t also depends on E<tE_{<t} through ϕt\phi_t. An agent in Π3\Pi_3 whose companion output is constant is an agent in Π1\Pi_1. Therefore Π1⊆Π3\Pi_1 \subseteq \Pi_3, and

sup⁡π∈Π3J(π)  ≥  sup⁡π∈Π1J(π). \sup_{\pi \in \Pi_3} J(\pi) \;\ge\; \sup_{\pi \in \Pi_1} J(\pi).

4

The inequality is strict for some deployments. Specifically, suppose the correct actions on task tt depend on information in E<tE_{<t} that cannot be written in BB tokens. No agent in Π1\Pi_1 takes the correct actions in this case, while an agent in Π3\Pi_3 can do so if its companion model has enough parameters.

We then compare Π2\Pi_2 with Π1\Pi_1. In Equation (2), ϕ\phi is trained only on data collected before deployment, so ata_t depends on E<tE_{<t} only through at most BB tokens, as in Equation (1). A frozen companion model can improve performance at the start of deployment.

This argument shows that better agents exist in Π3\Pi_3.

Why is it hard to make Agent Companion models keep updating?

If a learning agent-companion model has a higher upper bound, we next ask what model it should be. We consider three properties.

  • Task breadth is measured by performance on tasks not seen in training.

  • Online plasticity means that the parameters are updated with deployment trajectories and feedback, and it is measured by the gain on later tasks without loss on earlier ones, which is the stability–plasticity dilemma [16].

  • Runtime efficiency is measured by tokens, compute, memory and inference latency.

We call the trade-off among the three the agent companion-model trilemma.

Companion Model Trilemma
Figure 3. Companion Model Trilemma

Existing solutions satisfy two of the three.

  • A frozen frontier model, which is accessed through a provider’s API, performs well across many domains and requires no training by the deployer, but its parameters are not updated on the deployer’s data.

  • A large locally deployed model can be fine-tuned periodically and also performs well across domains; but the training procedure, version evaluation are tedious to perform.

  • A small fine-tuned specialist is the most common solution now, but it is cheap to serve and to update online. It is trained for one task, however, so each additional task requires its own data pipeline, training objective and evals, which adds to the maintenance cost.

Can online plasticity resolve the trilemma?

The trilemma holds under one assumption: a model’s task breadth is fixed at deployment time and comes from parameter size and pretraining dataset size. Under this assumption a model needs more parameters to perform more tasks, and more parameters cost more compute.

This suggests another possibility: a small companion model could learn additional related tasks through deployment experience, rather than require a larger pretrained model or a separately designed model for each task. It need not perform every task in its domain at deployment time. It needs a learning procedure that can reuse experience across those tasks.

Online plasticity may therefore change the breadth–efficiency tradeoff. Whether it does depends on how much additional task coverage can be learned at a fixed parameter budget, how well earlier learning is preserved, and how much computation the updates require.

ICL remains necessary. A new instruction or tool schema can affect the next action before any parameter update is available. The two mechanisms can serve different purposes:

  • Fast adaptation through context: incorporate new instructions, current state, and recent feedback immediately.

  • Slower consolidation into parameters: learn reusable patterns from selected deployment experience, so later tasks benefit without requiring the same source records in every context. These updates can be stored as customer-controlled parameter deltas or checkpoints, with versioning, regression tests, audit records, and rollback.

Fast adaptation and slow consolidation
Figure 4. Fast adaptation and slow consolidation

[17], [18], [19].

The distinction between rapid acquisition and gradual integration has a precedent in complementary learning systems theory [17]. In agent systems, Letta’s memory-model agenda explores a related combination: a trained model creates reusable context that downstream models consume through ICL, with parameter distillation discussed as a further possibility [19].

The remaining question is how to train a companion model for continued learning, rather than only for performance at deployment. This requires studying the architecture, training objective, and deployment data together, including what to consolidate, when to update, and what to revise or forget. That’s what we will start at TrueIntrinsics.

References

[1]
Capgemini Research Institute. 2025. “Trust and Human–AI Collaboration Set to Define the Next Era of Agentic AI, Unlocking $450 Billion Opportunity by 2028.” July 16, 2025. Official release. ↩︎
[2]
LangChain. 2026. “State of Agent Engineering.” June 12, 2026; voluntary survey of 1,340 respondents conducted November 18–December 2, 2025. Original report. ↩︎
[3]
Trivedy, Vivek. 2026. “The Anatomy of an Agent Harness.” LangChain, March 10, 2026. Original post. ↩︎
[4]
Cai, Yishuo, Xingyu Guo, Xuancheng Huang, et al. 2026. “From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory.” Accepted at ICML 2026. arXiv:2606.08656. doi:10.48550/arXiv.2606.08656. ↩︎
[5]
Ma, Zeyao, Jing Zhang, Xiaokang Zhang, et al. 2026. “Scaling Agentic Verifier for Competitive Coding.” arXiv:2602.04254. doi:10.48550/arXiv.2602.04254. ↩︎
[6]
Li, Zecheng, Zhihui Cao, Wenke Huang, et al. 2026. “MagicGUI-RMS: A Multi-Agent Reward Model System for Self-Evolving GUI Agents via Automated Feedback Reflux.” Preprint. arXiv:2601.13060. doi:10.48550/arXiv.2601.13060. ↩︎
[7]
Ong, Isaac, Amjad Almahairi, Vincent Wu, et al. 2024. “RouteLLM: Learning to Route LLMs with Preference Data.” arXiv:2406.18665. doi:10.48550/arXiv.2406.18665. ↩︎
[8]
Wang, Xingyao, Valerie Chen, Heng Ji, and Graham Neubig. 2026. “A Rubric-Supervised Critic from Sparse Real-World Outcomes.” Preprint. arXiv:2603.03800. doi:10.48550/arXiv.2603.03800. ↩︎
[9]
Chae, Hyungjoo, Namyoung Kim, Kai Tzu-iunn Ong, et al. 2025. “Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation.” ICLR 2025. doi:10.48550/arXiv.2410.13232. ↩︎
[10]
Zaharia, Matei, Omar Khattab, Lingjiao Chen, et al. 2024. “The Shift from Models to Compound AI Systems.” Berkeley Artificial Intelligence Research, February 18, 2024. Original post. ↩︎
[11]
Asawa, Parth, Christopher M. Glaze, Gabriel Orlanski, Ramya Ramakrishnan, Benji Xu, Asim Biswal, Vincent Sunn Chen, Frederic Sala, Matei Zaharia, and Joseph E. Gonzalez. 2026. “Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments.” Preprint. arXiv:2606.05661. doi:10.48550/arXiv.2606.05661. ↩︎
[12]
CL-Bench. 2026. “Continual Learning Bench Public Leaderboard.” Official leaderboard. The article reports the author’s September 24, 2026 snapshot; the live page may change. ↩︎
[13]
Hu, Qisheng, Quanyu Long, and Wenya Wang. 2026. “When Continual Learning Moves to Memory: A Study of Experience Reuse in LLM Agents.” arXiv:2604.27003. doi:10.48550/arXiv.2604.27003. ↩︎
[14]
Xiong, Zidi, Yuping Lin, Wenya Xie, et al. 2025. “How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior.” arXiv:2505.16067. doi:10.48550/arXiv.2505.16067. ↩︎
[15]
Zerhoudi, Saber, Jelena Mitrovic, and Michael Granitzer. 2026. “The Compaction Cliff in Long-Running AI Agent Memory.” arXiv:2608.22752. doi:10.48550/arXiv.2608.22752. ↩︎
[16]
Mermillod, Martial, Aurélia Bugaiska, and Patrick Bonin. 2013. “The Stability–Plasticity Dilemma: Investigating the Continuum from Catastrophic Forgetting to Age-Limited Learning Effects.” Frontiers in Psychology 4:504. doi:10.3389/fpsyg.2013.00504. ↩︎
[17]
McClelland, James L., Bruce L. McNaughton, and Randall C. O’Reilly. 1995. “Why There Are Complementary Learning Systems in the Hippocampus and Neocortex: Insights from the Successes and Failures of Connectionist Models of Learning and Memory.” Psychological Review 102(3):419–457. doi:10.1037/0033-295X.102.3.419. ↩︎
[18]
Ba, Jimmy, Geoffrey Hinton, Volodymyr Mnih, Joel Z. Leibo, and Catalin Ionescu. 2016. “Using Fast Weights to Attend to the Recent Past.” arXiv:1610.06258. doi:10.48550/arXiv.1610.06258. ↩︎
[19]
Letta. 2026. “Memory Models: Towards Agents That Learn.” June 25, 2026. Research post. ↩︎
[20]
Behrouz, Ali, Peilin Zhong, and Vahab Mirrokni. 2024. “Titans: Learning to Memorize at Test Time.” arXiv:2501.00663. doi:10.48550/arXiv.2501.00663.
[21]
Hübotter, Jonas, Frederike Lübeck, Lejs Behric, et al. 2026. “Reinforcement Learning via Self-Distillation.” arXiv:2601.20802. doi:10.48550/arXiv.2601.20802.
[22]
Inan, Hakan, Kartikeya Upasani, Jianfeng Chi, et al. 2023. “Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.” arXiv:2312.06674. doi:10.48550/arXiv.2312.06674.
[23]
Meta. 2025. “Llama Prompt Guard 2 Model Card.” Official model card for the 86M and 22M prompt-attack classifiers, accessed October 5, 2026. Model card.
[24]
Sculley, D., Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Dennison. 2015. “Hidden Technical Debt in Machine Learning Systems.” Advances in Neural Information Processing Systems 28. Conference paper.
[25]
Sun, Yu, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, and Carlos Guestrin. 2024. “Learning to (Learn at Test Time): RNNs with Expressive Hidden States.” arXiv:2407.04620. doi:10.48550/arXiv.2407.04620.
[26]
Wang, Meng, Haohan Zhao, Wenzhuo Liu, et al. 2026. “Denser ≠ Better: Limits of On-Policy Self-Distillation for Continual Post-Training.” arXiv:2607.01763. doi:10.48550/arXiv.2607.01763.