Coinbase benchmark finds newer AI models caught less payment fraud in replay

Editorial illustration: Two conveyor belts carry cream and red tiles through metal screens. A fine mesh screen holds back several red tiles; a wider grid catches fewer, with larger red tiles visible beyond it.

In brief

  • Coinbase replayed 16,140 Onramp transactions, 813 of them confirmed fraud, through paired AI model versions.
  • Every newer version scored lower on recall, F1 and dollar-weighted recall, Coinbase said.
  • Sonnet's recall fell 22.2 points; GPT's precision rose 11.5 points while its recall fell 20.7.
  • Coinbase said it did not establish the cause and cannot release the dataset.
  • A post-trained Qwen3.5-9B model beat Opus 4.5, a separate Oct. 8 disclosure said.

What the replay measured

The test cohort covered nine weeks before Coinbase's risk agent rolled out, keeping all matured fraud cases while sampling legitimate traffic. Each model reviewed recent transaction behavior under fixed guidance and the same decision policy. That's the important design choice (it isolates the model's own behavior rather than comparing redesigned screening systems).

Coinbase compared Opus 4.5 against Opus 5 and Sonnet 4.6 against Sonnet 5, plus GPT-5.4 against GPT-5.6 (sol). Recall is the share of fraud cases a model catches. Dollar-weighted recall is how much of the total fraud value it catches.

Every newer version came in lower on recall and F1, and on dollar-weighted recall too.

Where the drops landed

Sonnet's were the steepest, according to Coinbase's Oct. 7 disclosure as reported by CryptoSlate: recall fell 22.2 percentage points and dollar-weighted recall fell 22.9 points. Opus's recall slipped just 0.8 points, though both newer Sonnet and Opus models also had lower precision. GPT moved the other way on one metric. Its precision rose 11.5 points, but recall fell 20.7 points and dollar-weighted recall fell 21.8.

There are real limits here. The replay didn't establish customer losses from deploying the newer versions, and Coinbase said it could identify the regressions without establishing their cause. The SR-Fraud researchers also said the proprietary dataset can't be released, which restricts independent replication and generalization. (Their related study first appeared Sept. 23 and was revised Sept. 30.)

An earlier online experiment found the agent-enabled flow with selective LLM review recorded 30% fewer fraudulent transactions and 22% less fraud value than existing models and rules alone. It didn't compare newer model versions.

A smaller model trained on fraud outcomes

In an Oct. 8 disclosure, Coinbase said a post-trained Qwen3.5-9B model beat Opus 4.5 across four fraud-detection metrics, with F1 up 9.6 points and dollar-weighted recall up 35.4 points. The company specialized it using historical fraud outcomes and deterministic rewards balancing fraudulent and legitimate examples. Production measurements from a separate evaluation put median end-to-end request latency at 0.683 seconds versus 1.515 seconds for Opus 4.5 (a 55% relative reduction).

Coinbase's advice is practical. Test a candidate model in the existing configuration first, then evaluate changed prompts or thresholds separately, weighing latency and reliability (and cost) alongside detection quality.

Frequently asked questions

What did Coinbase's AI fraud benchmark test?

Coinbase replayed 16,140 Onramp transactions across 7,293 users, including 813 confirmed fraud cases, from the nine weeks before its risk agent rolled out. Each model worked under fixed guidance and the same decision policy, so the test isolated the model's behavior rather than comparing redesigned systems.

Does the benchmark show newer AI models caused fraud losses?

No. The replay did not establish customer losses from deploying the newer versions. Coinbase said it could identify the regressions without establishing their cause, and the proprietary dataset can't be released for independent replication.

What is dollar-weighted recall?

Recall measures the share of fraud cases a model catches. Dollar-weighted recall measures how much of the total fraud value it catches. In Coinbase's replay, every newer model version scored lower on both.