Coinbase reported Oct. 7 that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service, despite an unchanged decision policy. The findings challenge the assumption that upgrading a model improves an existing payment screener.
The company’s evaluation replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions. The cohort covered nine weeks before its risk agent rolled out, retaining all matured fraud cases while sampling legitimate traffic.
Each candidate reviewed recent transaction behavior under fixed guidance and the same policy for turning risk classifications into decisions. This isolated the decision model’s behavior within that setup, rather than comparing redesigned screening systems.
Results from a fixed historical replay
Coinbase compared Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Every newer version had lower recall, a lower combined precision-and-recall score called F1, and lower dollar-weighted recall. Recall measures the share of fraud cases a model catches; dollar-weighted recall measures how much of the total fraud value it catches.
Sonnet’s recall fell 22.2 percentage points and its dollar-weighted recall dropped 22.9 points. Opus’s recall declined 0.8 points. Both newer models also had lower precision, meaning a smaller share of transactions they classified as fraud were actually fraudulent.
GPT showed why one improving score can be misleading. Its precision rose 11.5 percentage points, but recall fell 20.7 points and dollar-weighted recall fell 21.8 points. Its fraud flags were more accurate, while more fraud cases and value escaped detection in the replay.


The replay does not establish customer losses from deploying those versions. Coinbase also said it could identify the regressions without establishing their cause.
Coinbase’s earlier online experiment compared adding selective LLM review with the existing models and rules alone. That agent-enabled flow recorded 30% fewer fraudulent transactions and 22% less fraud value; it did not compare newer model versions.
In their limitations, the SR-Fraud researchers say the proprietary dataset cannot be released, restricting independent replication and generalization. Their related payment-fraud study first appeared Sept. 23 and was revised Sept. 30, before the October blogs.
A separate case for a custom model
In its Oct. 8 disclosure, Coinbase reported that a post-trained Qwen3.5-9B model exceeded Opus 4.5 across four fraud-detection metrics. F1 improved 9.6 percentage points and dollar-weighted recall rose 35.4 points. The company specialized it using historical fraud outcomes and deterministic rewards balancing fraudulent and legitimate examples.
Separately, production measurements put median end-to-end LLM-request latency at 0.683 seconds versus 1.515 seconds for Opus 4.5, a 55% relative reduction. Faster inference and stronger benchmark detection came from different evaluations.
For payment providers, the upgrade question is whether a candidate improves fraud coverage under their actual decision setup. Coinbase recommends testing that configuration first, then evaluating changed prompts or thresholds separately, with latency, reliability and cost alongside detection quality.
AI,Featured,Payments,Research,Anthropic,Coinbase,OpenAI,payments,SolanaAI,Anthropic,Coinbase,OpenAI,payments,Solana#Payment #fraud #detection #fell #newer #Coinbase #test1791657649
