Harvey and EngramLab Release 100M-Token Synthetic Legal Dataset

An individual viewing glowing numbers on a screen, symbolizing technology and data.

In brief

  • Harvey and EngramLab released 100+ million token synthetic dataset covering 250+ client matters and 46 clients.
  • Dataset trains AI agents to replicate legal professional knowledge without confidentiality or privacy risks.
  • EngramLab's memory-layer tech compresses token usage by up to 100x, reducing computational costs.

Solving the confidentiality problem

Synthetic data offers a workaround to the confidentiality barriers that typically constrain legal AI training. The dataset provides all the structural complexity of real legal work without exposing sensitive client information. This approach lets law firms and AI developers build better models while respecting privacy obligations that govern the legal profession.

The partnership between Harvey and EngramLab predates this dataset release. Their collaboration has centered on encoding law-firm processes through AI, creating tools that let firms integrate AI into workflows without surrendering control of proprietary knowledge.

EngramLab's compression advantage

EngramLab's memory-layer technology can compress organizational context enough to cut token usage by up to 100x. This efficiency gain matters because training and running legal AI agents is computationally expensive. The compression layer lets developers build scalable systems without ballooning inference costs.

EngramLab raised $98 million in funding on June 23, 2026, at a valuation of approximately $600 million. The company's focus on memory compression and institutional knowledge encoding reflects a broader shift in legal AI toward practical, cost-effective deployment.

Harvey, currently valued at $11 billion, has established itself as a platform for law firms looking to integrate AI into their workflows without surrendering control of their proprietary knowledge. Earlier in 2026, the company open-sourced the Legal Agent Benchmark, known as LAB, which established standardized ways to measure how well AI agents perform on legal tasks.

The open-source synthetic dataset builds on that foundation. By releasing training data alongside measurement benchmarks, Harvey and EngramLab are creating infrastructure that lets the broader legal AI ecosystem develop faster, cheaper, and with fewer confidentiality trade-offs.