Live Prices
Markets

Coinbase AI Benchmark Reveals Upgraded Models Missed More Payment Fraud

TheCryptoDesk Editorial · 2m read
Coinbase AI Benchmark Reveals Upgraded Models Missed More Payment Fraud

On Oct. 7, Coinbase disclosed benchmark results showing that updated versions of major artificial intelligence models caught fewer fraudulent transactions and missed a higher proportion of dollar-weighted fraud during historical payment screening tests for its Onramp service.

The evaluation replayed 16,140 transactions across 7,293 users, which included 813 confirmed fraudulent transactions spanning nine weeks prior to the company's risk agent deployment. Despite applying an unchanged decision policy and identical screening parameters, newer model iterations consistently underperformed their predecessors in detecting fraudulent activity.

Model Upgrades Lead to Detection Regressions

Coinbase evaluated three major model families: comparing Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Across every test, the newer iterations registered lower recall, reduced F1 scores, and lower dollar-weighted recall, which measures the total value of caught fraud.

  • Sonnet 5: Recall dropped 22.2 percentage points and dollar-weighted recall dropped 22.9 points compared to Sonnet 4.6.
  • Opus 5: Recall declined 0.8 points against Opus 4.5, alongside a drop in precision.
  • GPT-5.6 (sol): Precision increased by 11.5 percentage points, but recall dropped 20.7 points and dollar-weighted recall fell 21.8 points.

While Coinbase noted that the historical replay does not establish actual customer losses, it highlighted that upgrading to newer frontier models can introduce unexpected performance regressions in dedicated financial compliance setups. These model behavioral shifts highlight ongoing industry discussions around model capabilities, such as when the Ledger CTO rejected claims that AI will break Bitcoin ECDSA cryptography.

Custom Models Outperform Generalist Systems

Following the initial findings, Coinbase released a separate disclosure on Oct. 8 showing that a specialized, post-trained Qwen3.5-9B model outperformed Opus 4.5 across four key fraud-detection metrics. The custom model achieved a 9.6 percentage point improvement in F1 score and a 35.4 point jump in dollar-weighted recall by utilizing deterministic rewards tied to historical fraud outcomes.

In production measurements, Qwen3.5-9B recorded a median end-to-end request latency of 0.683 seconds, representing a 55% relative reduction compared to 1.515 seconds for Opus 4.5.

The benchmark limitations noted by SR-Fraud researchers indicate that the proprietary dataset cannot be publicly released, preventing third-party replication. The original underlying study appeared on Sept. 23 and was revised on Sept. 30. This technical development follows other recent corporate updates from the exchange, including when Coinbase secured a shareholder lawsuit dismissal.

Why It Matters

These findings counter the standard assumption in crypto compliance that upgrading to the newest general-purpose large language model automatically improves domain-specific security. For exchanges processing high-volume fiat and crypto transactions, off-the-shelf LLMs present operational risks if deployed without backtesting against fixed historical datasets. Coinbase's success with custom-trained open models suggests that smaller, domain-tailored neural networks may provide faster, safer, and more cost-effective fraud infrastructure than proprietary frontier generalist models.

Read next