MIT CSAIL paper says minimalist JAZ agent beat Letta and ACE on two benchmarks

Editorial illustration: Two parallel teal tracks contain cream-colored steps and rectangular archways. Compact wheeled blocks sit farther along each track than intricate wheeled machines, under warm side lighting.

In brief

  • JAZ, a stripped-down agent framework from MIT CSAIL, is built on one LLM primitive called invoke.
  • StuLife: JAZ scored 70% versus Letta's 62% using GPT-5.4 nano, the paper reports.
  • AppWorld: JAZ scored 74%, 4 points ahead of ACE, at lower cost, the researchers said.
  • Results cover only two benchmarks; the StuLife test used one small model.
  • The framework and evaluation code are published on GitHub.

How JAZ works

JAZ is built around one idea. Per Crypto Briefing's report, the framework relies on a single LLM-based primitive called invoke, and it exposes the agent's history and prompt as variables inside a code environment. The model can then write executable code to inspect, slice or manipulate that history directly.

That's a different approach from Letta. Its design is built around a stateful memory hierarchy for managing context.

The paper is titled "Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity," and it's posted to arXiv under the identifier arXiv:2609.26891.

The benchmark numbers

The researchers tested JAZ on two fronts: long-term recall and self-improvement. For recall, they used the StuLife benchmark and ran JAZ on the GPT-5.4 nano model. JAZ scored 70% and Letta scored 62%, according to the paper as reported by Crypto Briefing, and JAZ reportedly got there at approximately half the cost of Letta.

The second test used the AppWorld benchmark. There, JAZ went up against ACE, a harness built for iterative self-improvement in CodeAct environments, and the researchers said JAZ scored 74% (4 percentage points ahead of ACE) at lower cost.

Those are the researchers' own numbers.

The caveats

The scope is narrow. The results cover just two benchmarks, StuLife and AppWorld, and the StuLife comparison was run on a single small model, so the findings don't extend beyond the setups that were tested.

Crypto Briefing also flagged a design trade-off that's easy to miss in the headline scores. Letting a model run arbitrary code against its own history raises sandboxing, error-handling and predictability concerns, and those matter more when the agent is effectively programming its own memory access.

The team has published the framework and its evaluation code on GitHub, in the jaz-lang/jaz and jaz-lang/jaz-evals repositories. For now, the benchmark claims rest on the paper itself.

Frequently asked questions

How does MIT's JAZ agent framework handle memory?

JAZ relies on a single LLM-based primitive called invoke. The agent's history and prompt are exposed as variables inside a code environment, and the model can write executable code to inspect, slice or manipulate that history directly.

How did JAZ score against Letta and ACE?

According to the MIT CSAIL paper as reported by Crypto Briefing, JAZ scored 70% versus Letta's 62% on StuLife using GPT-5.4 nano. On AppWorld it scored 74%, 4 percentage points ahead of ACE, at lower cost.

What are the limits of the JAZ benchmark results?

The results cover only two benchmarks, StuLife and AppWorld, and the StuLife comparison was run on a single small model, GPT-5.4 nano. Crypto Briefing also noted that letting a model run arbitrary code against its own history raises sandboxing, error-handling and predictability concerns.