CAMG.RL arXiv:2609.34422

Research project / arXiv:2609.34422

Coding Agent
Memory Post-training

Unlocking the memory potential of pre-trained file operations for long-horizon tasks via reinforcement learning.

Lirui LuoKelong MaoHeming XiaRongqing LiXinwei YangLuyu ChenKieran WongYudong GuoXinrui WangJiayin ZhuSimiu GuSulong XuCong Fang

JD.com

Research thesis Files are not just artifacts. They are memory.

CAMG-RL trains a task-acting policy to create, revise, search, and reuse ordinary workspace files across context boundaries - using the same file operations that code-pretrained agents already know.

4task environments
2model scales
54.4%average success
memory loop / active
CAMG-RL compares context-only work with file-backed memory across a context boundary.
FIG. 01 A continuation file keeps the work alive when context does not.
01ordinary shell actions+persistent workspace
CAMG/CAMG-RL/ARXIV:2609.34422/JD.COM/2026
01 / The idea

When the prompt
runs out of room.

A long task does not fail only because the model forgets. It fails because the useful state has nowhere durable to live.

CAMG gives the policy a persistent workspace and lets reinforcement learning discover when to write a plan, preserve evidence, or read a decision back. The interface is familiar before training begins.

01context boundary

Make the state explicit.

Before replacement, the policy writes a bounded continuation file. After replacement, it receives the path - not a hidden summary - and reads the state back through a normal file action.

02pre-training prior
“

Use what the model already knows.

Shell commands, text files, search, and revision are native coding-agent behavior. CAMG-RL turns that prior into deliberate long-horizon memory with task reward.

$ cat .agent_memory/CONTINUATION.md
03learning signal

Memory earns its keep.

There is no separate memory reward. The only question is whether the saved state helps the downstream task succeed.

task rewardpolicy update
02 / The gym

One interface.
Four kinds of work.

Each environment keeps its native task semantics while sharing executable shell access and an episode-persistent workspace.

native task + file memory01 / 04

Shop

Search, compare, and purchase under a budget while keeping product evidence and constraints available over a long interaction.

Native taskbrowse + buy
Memory rolepreferences, evidence
Shared actionshell_command
episode workspace.agent_memory/preferences.md
CAMG-RL architecture: rollout GPUs, task environments, memory files, an episode queue, and PPO learner GPUs.
FIG. 02 A single policy samples native task actions and filesystem actions into one asynchronous PPO loop.
03 / The method

Memory as
reasoning.

The policy does not emit a separate hidden reasoning channel. It writes the reasoning state it needs into files that can be inspected, revised, and reused.

A
act

Work on the task

Use the native tool and ordinary shell actions in the same response space.

B
write

Save what matters

Record the objective, evidence, open questions, and next step in editable files.

C
return

Read it back

After context replacement, retrieve the state with a normal file read and keep going.

.agent_memory/CONTINUATION.md persisted
# checkpoint before context replacementobjective: fix response logging after recent changesevidence: added _mqpush in scrapy/core/scheduler.pytested: no _mqpush errornext: run tests, confirm logging, submit$ cat .agent_memory/CONTINUATION.md
04 / The evidence

Small moves.
Longer memory.

The trained 4B policy moves a little, but its behavior changes where it counts: toward complete memory chains and better long-horizon task outcomes.

Average CAMG success54.4%+6.4 pts vs. CompactionRL
Complete memory chains20.9%last training quarter
External transfer15.8%SWE-bench Verified
Model size4BQwen3.5 base
native held-out panel

Files beat dedicated memory tools.

Greedy decoding on the frozen 128-task panel per environment. Success rates are the paper's reader-facing values.

Four-stage coding episode showing a reproducer, patch, continuation file, and verification.
CASE STUDY A coding episode carries the task state through a context replacement.
Analysis of filesystem memory actions versus dedicated memory tools.
INTERFACE PRIOR The base policy already finds file operations familiar.
CAMG-RL ablation results across environments.
ABLATION Memory-action gradients and general files both matter.
05 / Beyond the gym

Train on memory.
Transfer the behavior.

The final CAMG-RL checkpoint is evaluated beyond its training distribution on software engineering and machine-learning engineering tasks.

See the evaluation protocol
Model / policySWE-bench
Verified
MLE-bench
Lite
Average
Qwen3.5-4B7.6%0.0%3.8%
Qwen3.5-35B-A3B15.6%4.5%10.1%
CAMG-RL-4B15.8%4.5%10.2%
Qwen3.5-122B-A10B22.0%9.1%15.6%
CAMG-RL-9B27.6%9.1%18.4%
Qwen3.5-397B-A17B34.4%13.6%24.0%

Matched task assignments, decoding budgets, environments, and graders.

Project release

Coding Agent Memory Post-training

Unlocking the memory potential of pre-trained file operations for long-horizon tasks via reinforcement learning.

CAMG / CAMG-RLQwen3.5-4B / 9BAgentic RL
Read PDF Code release: preparing the public snapshot.