MIT’s minimalist JAZ agent beats Letta and ACE on memory tasks

Photo: Tara Winstead / Pexels

MIT’s minimalist JAZ agent beats Letta and ACE on memory tasks

A new MIT CSAIL paper finds that letting an agent treat its own history as code variables can outperform dedicated memory systems at lower cost

Researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) have introduced JAZ, a stripped-down agent framework. In benchmark tests it beat Letta, formerly known as MemGPT, on recall-heavy tasks. It also topped ACE, a specialized self-improvement harness, while spending less money to do it.

Instead of giving a language model a filing cabinet, give it the ability to write code that reaches into its own past.

One primitive to rule them all

The paper is titled “Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity.” It was posted to arXiv under the identifier arXiv:2609.26891.

At its core, JAZ relies on a single LLM-based primitive called invoke. That’s the whole toolkit, more or less.

JAZ takes a different route. The agent’s history and its prompt are exposed as variables inside a code environment. The model can then write executable code to inspect, slice, or manipulate that history directly.

Advertisement

The benchmark numbers

The researchers tested JAZ on two fronts: long-term recall and self-improvement.

For recall, they used the StuLife benchmark and ran JAZ on the GPT-5.4 nano model. JAZ scored 70%. Letta scored 62%.

That eight-point gap is notable on its own. The cost side makes it more interesting: JAZ reportedly got there at approximately half the cost of Letta.

Letta’s design is built around a stateful memory hierarchy for managing context. In plainer terms, it organizes what the model remembers into tiers, deciding what stays close at hand and what gets archived.

The second test focused on self-improvement, using the AppWorld benchmark. Here JAZ went up against ACE, a harness designed specifically for iterative self-improvement in CodeAct environments.

JAZ scored 74%, finishing 4 percentage points ahead of ACE. It did so at lower cost.

What this means for agent builders

For developers building AI agents, the most practical takeaway is about cost. Getting a higher score at roughly half the price, as JAZ reportedly did against Letta, is the kind of result that gets attention in engineering budget meetings. When agents run thousands of tasks, per-task savings compound quickly.

There are reasons for caution. The results cover two benchmarks, StuLife and AppWorld, and the StuLife comparison was run on a single small model, GPT-5.4 nano.

Giving a model the ability to write and run arbitrary code against its own history also raises its own engineering questions. Sandboxing, error handling, and predictability matter more when the agent is effectively programming its own memory access.

The team has published the framework and its evaluation code on GitHub. The repositories are jaz-lang/jaz and jaz-lang/jaz-evals, so other researchers can poke at the results themselves.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
MIT’s minimalist JAZ agent beats Letta and ACE on memory tasks
MIT’s minimalist JAZ agent beats Letta and ACE on memory tasks

A new MIT CSAIL paper finds that letting an agent treat its own history as code variables can outperform dedicated memory systems at lower cost

Photo: Tara Winstead / Pexels

Researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) have introduced JAZ, a stripped-down agent framework. In benchmark tests it beat Letta, formerly known as MemGPT, on recall-heavy tasks. It also topped ACE, a specialized self-improvement harness, while spending less money to do it.

Instead of giving a language model a filing cabinet, give it the ability to write code that reaches into its own past.

One primitive to rule them all

The paper is titled “Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity.” It was posted to arXiv under the identifier arXiv:2609.26891.

At its core, JAZ relies on a single LLM-based primitive called invoke. That’s the whole toolkit, more or less.

JAZ takes a different route. The agent’s history and its prompt are exposed as variables inside a code environment. The model can then write executable code to inspect, slice, or manipulate that history directly.

Advertisement

The benchmark numbers

The researchers tested JAZ on two fronts: long-term recall and self-improvement.

For recall, they used the StuLife benchmark and ran JAZ on the GPT-5.4 nano model. JAZ scored 70%. Letta scored 62%.

That eight-point gap is notable on its own. The cost side makes it more interesting: JAZ reportedly got there at approximately half the cost of Letta.

Letta’s design is built around a stateful memory hierarchy for managing context. In plainer terms, it organizes what the model remembers into tiers, deciding what stays close at hand and what gets archived.

The second test focused on self-improvement, using the AppWorld benchmark. Here JAZ went up against ACE, a harness designed specifically for iterative self-improvement in CodeAct environments.

JAZ scored 74%, finishing 4 percentage points ahead of ACE. It did so at lower cost.

What this means for agent builders

For developers building AI agents, the most practical takeaway is about cost. Getting a higher score at roughly half the price, as JAZ reportedly did against Letta, is the kind of result that gets attention in engineering budget meetings. When agents run thousands of tasks, per-task savings compound quickly.

There are reasons for caution. The results cover two benchmarks, StuLife and AppWorld, and the StuLife comparison was run on a single small model, GPT-5.4 nano.

Giving a model the ability to write and run arbitrary code against its own history also raises its own engineering questions. Sandboxing, error handling, and predictability matter more when the agent is effectively programming its own memory access.

The team has published the framework and its evaluation code on GitHub. The repositories are jaz-lang/jaz and jaz-lang/jaz-evals, so other researchers can poke at the results themselves.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.