Photo: Tara Winstead / Pexels
MITās minimalist JAZ agent beats Letta and ACE on memory tasks
A new MIT CSAIL paper finds that letting an agent treat its own history as code variables can outperform dedicated memory systems at lower cost
Researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) have introduced JAZ, a stripped-down agent framework. In benchmark tests it beat Letta, formerly known as MemGPT, on recall-heavy tasks. It also topped ACE, a specialized self-improvement harness, while spending less money to do it.
Instead of giving a language model a filing cabinet, give it the ability to write code that reaches into its own past.
One primitive to rule them all
The paper is titled “Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity.” It was posted to arXiv under the identifier arXiv:2609.26891.
At its core, JAZ relies on a single LLM-based primitive called invoke. That’s the whole toolkit, more or less.
JAZ takes a different route. The agent’s history and its prompt are exposed as variables inside a code environment. The model can then write executable code to inspect, slice, or manipulate that history directly.
The benchmark numbers
The researchers tested JAZ on two fronts: long-term recall and self-improvement.
For recall, they used the StuLife benchmark and ran JAZ on the GPT-5.4 nano model. JAZ scored 70%. Letta scored 62%.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
That eight-point gap is notable on its own. The cost side makes it more interesting: JAZ reportedly got there at approximately half the cost of Letta.
Letta’s design is built around a stateful memory hierarchy for managing context. In plainer terms, it organizes what the model remembers into tiers, deciding what stays close at hand and what gets archived.
The second test focused on self-improvement, using the AppWorld benchmark. Here JAZ went up against ACE, a harness designed specifically for iterative self-improvement in CodeAct environments.
JAZ scored 74%, finishing 4 percentage points ahead of ACE. It did so at lower cost.
What this means for agent builders
For developers building AI agents, the most practical takeaway is about cost. Getting a higher score at roughly half the price, as JAZ reportedly did against Letta, is the kind of result that gets attention in engineering budget meetings. When agents run thousands of tasks, per-task savings compound quickly.
There are reasons for caution. The results cover two benchmarks, StuLife and AppWorld, and the StuLife comparison was run on a single small model, GPT-5.4 nano.
Giving a model the ability to write and run arbitrary code against its own history also raises its own engineering questions. Sandboxing, error handling, and predictability matter more when the agent is effectively programming its own memory access.
The team has published the framework and its evaluation code on GitHub. The repositories are jaz-lang/jaz and jaz-lang/jaz-evals, so other researchers can poke at the results themselves.