Make the state explicit.
Before replacement, the policy writes a bounded continuation file. After replacement, it receives the path - not a hidden summary - and reads the state back through a normal file action.
Research project / arXiv:2609.34422
Unlocking the memory potential of pre-trained file operations for long-horizon tasks via reinforcement learning.
JD.com
Research thesis Files are not just artifacts. They are memory.
CAMG-RL trains a task-acting policy to create, revise, search, and reuse ordinary workspace files across context boundaries - using the same file operations that code-pretrained agents already know.
A long task does not fail only because the model forgets. It fails because the useful state has nowhere durable to live.
CAMG gives the policy a persistent workspace and lets reinforcement learning discover when to write a plan, preserve evidence, or read a decision back. The interface is familiar before training begins.
Before replacement, the policy writes a bounded continuation file. After replacement, it receives the path - not a hidden summary - and reads the state back through a normal file action.
Shell commands, text files, search, and revision are native coding-agent behavior. CAMG-RL turns that prior into deliberate long-horizon memory with task reward.
There is no separate memory reward. The only question is whether the saved state helps the downstream task succeed.
Each environment keeps its native task semantics while sharing executable shell access and an episode-persistent workspace.
Search, compare, and purchase under a budget while keeping product evidence and constraints available over a long interaction.
.agent_memory/preferences.md
The policy does not emit a separate hidden reasoning channel. It writes the reasoning state it needs into files that can be inspected, revised, and reused.
Use the native tool and ordinary shell actions in the same response space.
Record the objective, evidence, open questions, and next step in editable files.
After context replacement, retrieve the state with a normal file read and keep going.
The trained 4B policy moves a little, but its behavior changes where it counts: toward complete memory chains and better long-horizon task outcomes.
Greedy decoding on the frozen 128-task panel per environment. Success rates are the paper's reader-facing values.



The final CAMG-RL checkpoint is evaluated beyond its training distribution on software engineering and machine-learning engineering tasks.
See the evaluation protocol| Model / policy | SWE-bench Verified | MLE-bench Lite | Average |
|---|---|---|---|
| Qwen3.5-4B | 7.6% | 0.0% | 3.8% |
| Qwen3.5-35B-A3B | 15.6% | 4.5% | 10.1% |
| CAMG-RL-4B | 15.8% | 4.5% | 10.2% |
| Qwen3.5-122B-A10B | 22.0% | 9.1% | 15.6% |
| CAMG-RL-9B | 27.6% | 9.1% | 18.4% |
| Qwen3.5-397B-A17B | 34.4% | 13.6% | 24.0% |
Matched task assignments, decoding budgets, environments, and graders.
Project release
Unlocking the memory potential of pre-trained file operations for long-horizon tasks via reinforcement learning.