ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Agentic Self-distillation for Cross-task EvolutioN at Test-time
Abstract
A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their own trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text at the harness level, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on the agent's own execution trajectories during deployment. The agent executes each task once in one pass over the task stream, and the executed trajectory with its verification result is the only and immediate learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy.
We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verifier-accepted trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns from the trajectories, ASCENT can distill enhanced privileged experience for more effective and efficient execution. We characterize the population target of this distillation and the limits of sparse outcome selection.
Across ALFWorld, WebShop, and AppWorld on varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms the compared online adaptation methods, and transfers to held-out scenes, showing that a deployed agent can consolidate its own verified experience into its policy without a separate training phase or memory retrieval at inference.
Setting
Online Agentic Test-Time Training: learn from one attempt per task
At deployment time, an agent meets a one-pass stream of related tasks. OaTTT trains the LLM's weights on the agent's own execution trajectories, from one attempt per task. We ask:
Can a deployed agent turn its own verified execution trajectories, one per task, into persistent weight updates that improve its success on subsequent tasks?
In-context adaptation
- Stores reflections, memories, or skills from earlier episodes as text at the harness level and keeps the model frozen.
- Reuse depends on retrieving the right experience for the current task.
- It also depends on the frozen policy executing that experience over a long interaction.
OaTTT: evolving the weights
- The LLM's weights adapt online, at test time, from the agent's own execution trajectories.
- Like in-context adaptation agents, the agent learns from each trajectory once. It accumulates that learning in the LLM's weights, whereas those agents use textual harness-level memory.
- One pass over the task stream and one attempt per task, a multi-turn episode with a sparse episode-level verification outcome.
- Each outcome is recorded before the task's own update, so later success reflects how much earlier experience helps.
- Learning occurs entirely at test time, without task-specific or offline training. The updated weights carry across tasks.
Existing parametric adaptation also updates the LLM's weights. Some test-time training methods adapt them within a single episode and reset them afterward. Others need pre-deployment training of the adaptation behavior or several attempts per task, and RL post-training optimizes the policy before deployment, typically with many attempts per task or through single-attempt RL with a learned critic. None of them adapts a long-horizon agent's weights at test time, in a single pass over the task stream with one attempt per task, while carrying updates across tasks.
Episode \(i\) is the agent's single attempt at task \(x_i\), and \(v_i\) is its sparse, episode-level verification outcome, recorded before any update. Observations describe the resulting states and give no turn-level reward or causal credit. In-context adaptation methods evolve the memory \(\mathcal M_i\), parametric methods evolve the weights \(\theta_i\), joint methods may evolve both, and frozen methods such as the base evolve neither. All online methods, in-context or parametric, follow this OaTTT protocol.
OaTTT via direct imitation or reinforcement of experience
Directly imitating or reinforcing the generated tokens of a single attempt
A straightforward way to perform OaTTT is to directly imitate or reinforce the agent's own interaction trajectories. This is fragile, because each task yields only one attempt by the current policy. We formulate five such baselines in two families. Every method uses LoRA fast weights on the frozen base, carried across episodes, and all five learn only from the tokens the agent generated.
Imitation and reinforcement on the generated trajectories
The two families are direct imitation of the generated tokens and single-attempt policy gradient, which reinforces or suppresses them with a signed advantage. Since group-relative advantages need several attempts per task, Online REINFORCE uses a cross-task baseline and Online REINFORCE++ normalizes advantages over the stream.
- Direct imitation
- Online ungated imitation\(\omega=1\)
- Online RFT\(\omega=v_i\)
- Online RFT+KL\(\omega=v_i\), plus KL to the frozen base
- Single-attempt policy gradient
- Online REINFORCE\(\omega=v_i-\bar v_{\le i}\)
- Online REINFORCE++running-normalized advantage, clipped surrogate
\(y_{i,t,j}\): the \(j\)-th generated token of turn \(t\). \(c_{i,t,j}\): the student prefix taken from the single attempt and reused unchanged during the update. \(\mathcal I_i\) indexes the turns with nonempty responses, and \(\alpha_{i,t,j}\) averages tokens within each turn and then across turns. The weights \(\omega\) make each generated token more or less likely and say nothing about the other tokens at that position. Lowering a token moves its probability to tokens the policy already favors, and the KL anchor of Online RFT+KL pulls toward a base model that has not seen the episode. Every update therefore stays centered on the generated tokens, including incidental reasoning, formatting, and detours.
Scroll sideways or tap to open the figure at full size.
Direct online imitation or reinforcement degrades agent execution. An online update rule can become unstable when, along the stream, its policy drifts ever further from the initial model and degrades while its task success falls below that of the base. All five increasingly produce invalid actions, mostly in long responses, and over the second half of the stream they succeed less often than the base.
Online ungated imitation and Online RFT become nearly deterministic at the first response token while drifting from the initial model. Online RFT+KL drifts more slowly but still substantially. For the policy-gradient methods, successes soon stop, so nearly every update comes from a failed episode. Online REINFORCE ends collapsed onto first response tokens that the base finds very unlikely, and the first-response-token entropy of Online REINFORCE++ swings above and below the base as it drifts.
In OaTTT, we aim to enable single-attempt, per-sample online learning (batch size 1) of LLM weights, analogous to how in-context adaptation agents evolve. In these baselines, each update relies on a single trajectory, so its sample-specific noise is not averaged out. Because later attempts come from the updated policy, imitating or reinforcing the generated tokens can amplify this noise into overfitting, policy drift, and eventual degradation. ASCENT addresses this by distilling experience with an on-policy self-distillation (OPSD) objective, using next-token distributions from an anchored self-teacher conditioned on hindsight from verified trajectories as targets.
Method
ASCENT self-distills verified experience for stable online test-time training
ASCENT stably distills verified experience from the hindsight-informed predictions of the frozen initial model into LoRA fast weights, evolving the agent during OaTTT. When an attempt is verified, the frozen model reads privileged information built from the filtered trajectory and gives full next-token distributions at the agent's own student prefixes. Failed attempts skip the update. The updated agent produces better trajectories, enabling online self-evolution along the stream.
1 · AttemptOnline, at test time: task 98 arrives once, and the agent makes one attempt, turn by turn.
The view follows the action. Scroll sideways to look around.
Scroll sideways or tap to open the figure at full size.
1Gate the update on the verification outcome
Each episode yields one trajectory \(\tau_i\) and an episode-level verification outcome \(v_i\), recorded before any update. Only accepted episodes trigger a self-distillation update. Otherwise, the fast weights stay unchanged.
Verification selects trajectories for adaptation. It assigns no credit to individual turns. Updating only on accepted episodes acts as rejection sampling, so the teacher's hindsight always comes from a trajectory that passed verification. With binary rewards, maximizing the likelihood of successful trajectories gives an unbiased policy-gradient estimate. ASCENT thus learns only from verified successes and needs no value model or critic under sparse rewards.
\(\phi_i\): the persistent LoRA fast weights entering episode \(i\), carried across tasks. The base \(p_{\theta_0}\) stays fixed, and \(\mathsf{Evolve}\) is the operator from the setting, here a few optimizer steps on the ASCENT objective.
2Hindsight from the complete trajectory, with the action-validity filter
At a student prefix, the student has not yet seen the rest of its response, the later turns and observations, or the verification outcome. This later part, together with the verified outcome, is the hindsight. It shows how the episode continued from each earlier decision to verified success. Although it does not identify which actions caused success, it provides the student with the successful experience as privileged information via the self-teacher. Even an accepted trajectory can contain invalid-action turns, whose response yields no parseable action or whose action the environment does not execute, such as an unavailable command or code that raises an error. Invalid-action signals arise naturally during execution, requiring neither additional cost nor credit assignment. The action-validity filter marks these turns with a validity indicator \(d_{i,t}\) read from the environment's response at each turn. ASCENT builds the privileged information \(z_i\) from the task goal, a success statement, and the reasoning–action responses of valid turns. By default, \(z_i\) omits observation text.
Filtering changes only the teacher input. The filter records only whether an action was executed. It does not judge whether a valid action helps the task, gives no reward, and assigns no credit. The verified trajectory thus plays two roles. Its valid turns build \(z_i\), and all its nonempty responses, including those of invalid-action turns, still supply the student prefixes.
\(w_{i,t}\): the reasoning–action response at turn \(t\), from which the parser \(\rho\) extracts the action \(a_{i,t}=\rho(w_{i,t})\). \(o_{i,t+1}\): the observation returned after that action. \(z_i\) is built from the agent's own attempt, without an external reference solution.
3On-policy self-distillation at shared student prefixes
On-policy distillation trains on student-generated prefixes with a teacher's next-token distributions, and OPSD conditions a self-teacher on privileged information that the student lacks. ASCENT applies this asymmetry to each verifier-accepted episode, with the frozen base model \(p_{\theta_0}\) as the teacher. At every student prefix \(c_{i,t,j}\) of the agent's own attempt, the teacher reads \(z_i\mathbin{\oplus}c_{i,t,j}\) and gives a detached next-token distribution. The student reads \(c_{i,t,j}\) alone. ASCENT minimizes the forward KL, averaged over tokens within each turn and then across turns.
The teacher stays fixed at \(p_{\theta_0}\), so self-distillation injects each verified trajectory's experience into the student's fast weights through a stable version of the LLM. Student updates never change the teacher's targets for a given trajectory, which helps mitigate student drift during online learning. Because \(z_i\) covers the full horizon of the verified trajectory, the teacher sees at each turn how the episode continued to verified success, which gives the student better-informed turn-level guidance.
The turn set, token lengths, and weights \(\alpha_{i,t,j}\) match those of the common loss above, so ASCENT trains at the same student prefixes as direct imitation and replaces its target.
4Distill experience online stably through full-vocabulary distribution matching
With a signed weight \(\omega\), a term of the common loss performs hard-target matching on the generated token \(y\). A positive weight makes the student more confident in \(y\), a negative weight hands \(y\)'s probability to tokens the student already favors, and both tend to sharpen the distribution around the student's own outputs. Forward KL instead performs full-vocabulary distribution matching and stops only when \(p^{\phi}=q\). Alternatives the hindsight teacher prefers rise, and \(y\) falls when the teacher disagrees, for example at an invalid-action turn omitted from \(z_i\).
As a stable version of the LLM, the frozen initial copy does not follow the student, which prevents the target from sharpening with it. In Figure 1, ASCENT stays closest to the initial model while keeping its valid-action rate, with a relatively more stable dynamic behavior. Under a fixed collection distribution, the population target at each student prefix averages the teacher distributions of verifier-accepted trajectories through that prefix. The paper characterizes this target and the limits of sparse outcome selection.
\(\boldsymbol\ell^{\phi}\): the student's logits. \(\mathbf e_y\): the one-hot vector of the generated token \(y\). Left, one term of the common loss (hard-target matching). Right, the forward KL of ASCENT (full-vocabulary distribution matching). Position indices are suppressed.
5Potential turn efficiency
Under a turn budget, an unnecessary turn uses up part of the budget, so a response that starts a detour can be less likely to reach verified success. Verifier selection gives such responses less weight, and the hindsight teacher, which sees how the accepted episode reached success, can potentially lower them at earlier student prefixes.
Empirically, ASCENT has the fewest mean turns in the main and held-out tables and uses fewer turns than the base along the stream, including on tasks that both the un-evolved base model and ASCENT solve. Without turn-level credit, an accepted trajectory can still contain unnecessary turns, and one attempt cannot show what an alternative action would have achieved. A theoretical understanding and complete proof of this mechanism are left for future work.
What the teacher sees
Results
ASCENT consistently improves long-horizon task success
At both scales, ASCENT improves exact success over the base by more than 22 points on ALFWorld and WebShop and exceeds the strongest online comparator. It also has the highest WebShop mean score and outperforms every offline in-context adaptation method on ALFWorld. ASCENT outperforms TT-OPSD, our test-time adaptation of OPSD, which uses the same self-distillation with a teacher that sees only the current turn. This shows the benefit of hindsight from the complete verified trajectory.
These consistent gains show that cross-episode parameter updates can extract reusable behavior from the same sparse environment verification available to offline and online in-context adaptation methods.
Setup. Qwen3.5-4B and Qwen3.5-9B. The base, all compared methods, and ASCENT use the same ReAct template and greedy decoding. Online methods follow the OaTTT protocol on the same task order: one pass over the stream, a single scored attempt per task with a 50-turn budget, and each task scored before its trajectory can update the method. They receive the same environment verification signal. Offline baselines instead build memories, skills, or prompts from the same fixed-size training split before evaluation. They fall outside OaTTT and serve as reference points. The direct imitation and policy-gradient baselines are left out of these tables because they severely degrade under OaTTT (Figure 1). Main results are mean ± std over 3 independent runs.
ALFWorld seen
| Method | Qwen3.5-4B | Qwen3.5-9B | ||
|---|---|---|---|---|
| SR (%) | Turns | SR (%) | Turns | |
| Base | 46.4 | 35.1 | 55.0 | 32.6 |
| Offline | ||||
| MemP | 61.4±6.5 | 28.6 | 69.3±2.4 | 26.3 |
| ACE | 60.7±9.2 | 31.1 | 61.4±5.5 | 27.9 |
| EvoSkill | 49.8±4.9 | 34.4 | 65.7±3.5 | 29.0 |
| GEPA | 58.6±3.4 | 34.5 | 65.4±4.6 | 30.1 |
| TextGrad | 42.1±4.9 | 37.5 | 59.3±4.5 | 32.6 |
| Trace2Skill | 51.4±6.1 | 33.0 | 67.9±6.4 | 28.4 |
| Online | ||||
| MemP | 63.6±5.9 | 29.7 | 67.9±2.9 | 27.4 |
| ReasoningBank | 49.0±1.5 | 35.6 | 66.0±2.4 | 29.0 |
| ACE | 51.2±13.8 | 34.6 | 69.3±6.3 | 28.3 |
| A-Mem | 46.7±5.4 | 34.5 | 60.4±0.4 | 30.7 |
| MemRL | 45.0±5.2 | 35.4 | 63.6±2.5 | 29.5 |
| Dynamic Cheatsheet | 48.6±4.1 | 36.9 | 57.1±4.8 | 32.2 |
| TT-OPSD | 51.0±3.7 | 32.9 | 61.2±1.8 | 29.3 |
| ASCENT | 69.5±5.6 | 26.3 | 77.4±1.8 | 21.0 |
Table 1. ALFWorld seen stream (140 tasks, six task families): success rate and mean turns per episode. Bold marks the highest SR and fewest mean turns within each backbone. The paper's Table 1 also reports SR per task family and runtime.
WebShop
| Method | Qwen3.5-4B | Qwen3.5-9B | ||||
|---|---|---|---|---|---|---|
| Strict-SR | Mean-Score | Turns | Strict-SR | Mean-Score | Turns | |
| Base | 17.8 | 25.9 | 39.5 | 17.2 | 27.8 | 37.5 |
| Offline | ||||||
| MemP | 33.2±1.3 | 43.7±2.0 | 30.0 | 26.6±1.3 | 40.6±2.8 | 31.9 |
| ACE | 5.8±0.5 | 21.4±1.1 | 42.0 | 13.2±0.6 | 22.4±1.4 | 41.6 |
| EvoSkill | 20.0±1.6 | 29.3±1.9 | 38.0 | 20.2±2.1 | 33.4±1.4 | 38.2 |
| GEPA | 23.4±1.2 | 37.5±2.4 | 34.0 | 35.4±3.2 | 48.6±1.6 | 30.4 |
| TextGrad | 10.4±2.1 | 30.5±2.5 | 35.9 | 18.4±2.5 | 34.7±2.8 | 34.5 |
| Trace2Skill | 14.2±0.6 | 21.7±1.2 | 41.6 | 10.0±4.2 | 21.0±3.6 | 40.4 |
| Online | ||||||
| MemP | 36.8±3.4 | 47.7±4.8 | 28.7 | 22.4±2.5 | 38.9±4.1 | 32.9 |
| ReasoningBank | 10.6±0.7 | 21.6±1.0 | 42.1 | 12.6±2.1 | 21.8±1.5 | 41.5 |
| ACE | 10.6±3.2 | 23.6±4.3 | 40.6 | 16.0±1.4 | 32.0±2.6 | 36.1 |
| A-Mem | 23.2±3.1 | 36.4±2.9 | 33.9 | 15.4±0.6 | 32.5±4.3 | 35.6 |
| MemRL | 21.8±1.2 | 32.3±2.5 | 37.1 | 14.0±4.2 | 27.4±1.7 | 38.9 |
| Dynamic Cheatsheet | 16.2±1.5 | 28.4±2.1 | 39.4 | 13.4±2.9 | 27.4±3.0 | 39.1 |
| TT-OPSD | 25.6±0.5 | 37.1±1.7 | 34.6 | 24.2±2.4 | 35.9±2.3 | 34.8 |
| ASCENT | 41.9±0.3 | 62.4±1.0 | 17.8 | 41.5±1.8 | 56.8±1.7 | 26.7 |
Table 2. WebShop, 500-task stream per scale. Strict-SR (%) counts episodes with task score = 1, and Mean-Score (%) averages the task score over the stream. ASCENT's adaptation gate accepts a score of at least 0.9. Bold marks the highest Strict-SR and Mean-Score and the fewest turns within each backbone. The paper's Table 2 also reports runtime.
Cross-episode gains extend to difficult tasks
The base rarely solves Heat or Cool tasks (SR 6.2 and 8.0 at 4B, 0.0 and 8.0 at 9B), whereas ASCENT improves both task families at each scale (20.8 and 54.7 at 4B, 37.5 and 48.0 at 9B). Its gain over the base is also much larger on later 4B tasks, while the base remains flat. ASCENT derives this improvement from its own verified experience during deployment, without a stronger teacher or a pre-deployment adaptation stage.
Tables 1 and 2 show that ASCENT reduces capped mean turns per episode at both scales on both benchmarks. On ALFWorld tasks that both the un-evolved base model and ASCENT solve, its trajectories are shorter too. Episode time, which includes the attempt and the LoRA update after verified successes, is modestly higher than the base on ALFWorld (4B 97 → 115 s, 9B 115 → 133 s) and lower on WebShop (4B 87 → 76 s, 9B 166 → 164 s).
Transfer
The evolved policy transfers and continues learning under scene shift
On ALFWorld unseen (134 tasks in held-out rooms and layouts), ASCENT leads the online comparators in both frozen transfer and continued adaptation at both scales. At 4B, continued OaTTT, which resumes online updates, substantially exceeds frozen transfer (+15.9 points). At 9B, continuation can preserve the already strong transfer result. These results show that seen-stream experience is retained in the evolved policy and that the self-evolution protocol remains effective.
Analysis
Experience internalized in the weights completes a long-horizon task
On the task “put a cool egg in microwave”, the base and five in-context adaptation variants exhaust the 50-turn budget, whereas ASCENT completes it in 20 turns. The in-context variants fail either by retrieving irrelevant or contradictory experience or by not following relevant guidance. ASCENT retrieves no memory. It carries LoRA updates from two earlier cool-then-place tasks, internalizing that same-family experience in its weights, and transfers it across objects and destinations. Across whole streams, retrieving available same-type prior experience does not reliably lead to task completion (paper Appendix C.3).
Scroll sideways or tap to open the figure at full size.
Interaction efficiency
In Figure 5a, ASCENT's turn savings over the base grow later in the stream. They also persist on shared successes, the tasks that both the un-evolved base model and ASCENT solve. ASCENT uses fewer turns than the base on 64% (4B) and 67% (9B) of them (paper Table 8), which supports a gain in interaction efficiency beyond the change in the success/failure mix.
In-context and parametric co-evolution Preliminary
ASCENT stores experience in fast weights, whereas in-context adaptation stores it in the context. We combine ASCENT with ACE and MemP on ALFWorld seen and find that the two are complementary: adding ASCENT improves both in-context methods at both scales. Whether retrieved memory also helps on top of ASCENT depends on the model. On the smaller 4B model, ASCENT alone performs best, so the extra retrieved context may distract a weaker model, in line with its preference for shorter privileged information (Table 3 below). On the relatively stronger 9B model, both combinations exceed ASCENT alone, maybe because the model can use the retrieved memory and its evolved weights together. Co-evolving context and weights is thus promising for stronger agents.
Ablations
Generally reasoning is the key component of the privileged information, and gains hold across divergences
| Privileged information \(z_i\) | Filter | Qwen3.5-4B | Qwen3.5-9B | ||||
|---|---|---|---|---|---|---|---|
| SR (%) | Turns | Valid (%) | SR (%) | Turns | Valid (%) | ||
| Base | – | 46.4 | 35.1 | 70.1 | 55.0 | 32.6 | 78.7 |
| Action only: \(a_{i,t}\) | filter off | 48.8 | 32.9 | 64.4 | 42.9 | 34.1 | 66.2 |
| filter on | 48.3 | 33.5 | 61.2 | 48.6 | 32.3 | 71.4 | |
| Reasoning + action + observation: \((w_{i,t},o_{i,t+1})\) | filter off | 62.9 | 28.0 | 74.6 | 77.9 | 21.1 | 86.6 |
| filter on | 64.5 | 28.3 | 79.5 | 77.1 | 23.8 | 88.2 | |
| Reasoning + action: \(w_{i,t}\) | filter off | 60.0 | 30.6 | 74.1 | 75.3 | 21.6 | 81.8 |
| filter on, ASCENT's default | 69.5 | 26.3 | 78.0 | 77.4 | 21.0 | 85.9 | |
Table 3. Privileged-information ablation on ALFWorld seen, with the action-validity filter off (✗) or on (✓). Valid (%) is the valid-action rate, and Turns is the mean turns per episode. The shaded row is ASCENT's default (reasoning + action with the filter). Bold marks the best value in each column.
Action-only hindsight (\(a_{i,t}\)) performs near or below the base, while adding reasoning (\(w_{i,t}\)) consistently improves success (by more than 11 points) and turn efficiency.
This may be because reasoning explains why each action was taken, such as which subgoal it serves. It also makes up most of the tokens the student generates, while an action is only a short command, so action-only hindsight informs few of the positions the student is trained on.
Smaller models benefit more from shorter, filtered privileged information. At 4B, reasoning + action with the filter gives the highest success and fewest turns, and with reasoning in \(z_i\) the filter helps most at this scale.
On the relatively stronger 9B model, the full trajectory (\(w_{i,t}\), \(o_{i,t+1}\)) performs well, maybe due to the model's stronger general language capability. Here the filter changes success slightly while still raising the valid-action rate, possibly because the stronger model is less affected by invalid-action turns. Without the filter, observations may help by showing which actions the environment flagged as invalid, so they add little once the filter removes those turns. The default setting (\(w_{i,t}\) with the filter) at 9B also performs comparably, with slightly better turn efficiency.
Distillation divergence
ASCENT still improves on the base at both scales when reverse KL or Jensen–Shannon (JS) divergence replaces forward KL. Reverse KL trails forward KL at both scales, most at 9B. JS matches reverse KL at 4B but exceeds forward KL at 9B. Forward KL favors coverage of the teacher distribution, whereas reverse KL is mode-seeking.
Citation
BibTeX
@misc{lu2026ascent,
title = {{ASCENT}: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience},
author = {Lu, Haodong and Gong, Dong},
year = {2026},
eprint = {2610.05303},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2610.05303}
}