Coinbase fine-tunes Qwen3.5-9B model for fraud prevention, beating frontier models

Coinbase

Coinbase fine-tunes Qwen3.5-9B model for fraud prevention, beating frontier models

A small open model trained for under $100 outperformed Opus 4.5 on Coinbase's Onramp fraud benchmark while answering in less than half the time

Coinbase has spent the past stretch teaching a relatively small AI model to catch fraudsters. According to the exchange, the small model now does that job better than some of the most expensive AI systems on the market.

In its “Owning Intelligence” blog series, published October 7-8, 2026, Coinbase said a fine-tuned version of Alibaba’s Qwen3.5-9B model outperformed frontier models on fraud detection for its Onramp service. It was faster and far cheaper to train.

What Coinbase actually built

The project centers on an LLM-based fraud risk agent built specifically for Onramp, Coinbase’s service for buying crypto. The agent sits on top of Coinbase’s existing machine-learning models and looks at transactions that have already cleared those systems.

The agent’s output is deliberately narrow. It returns a risk level and nothing more, so upstream systems keep running as they were.

The results from A/B testing were meaningful. Coinbase said the agent cut fraudulent transactions by 30% and lowered the dollar value of fraud by 22%. The savings were estimated at three times the cost of running the model.

Advertisement

The benchmark showdown

To compare models, Coinbase built a proprietary Onramp fraud benchmark. It contains 16,140 transactions, of which 813 were confirmed fraud cases.

Coinbase fine-tuned Qwen3.5-9B using reinforcement learning. The tuned model then went head-to-head with frontier systems, including Anthropic’s Opus 4.5.

The 9-billion-parameter model came out ahead across four metrics: precision, recall, F1 score, and dollar-weighted recall. The margins ranged from 7.5 to 35.4 percentage points.

Speed also tilted toward the smaller model. Qwen posted a median latency of 0.683 seconds, compared with 1.515 seconds for Opus 4.5. That’s a 55% reduction.

Then there’s the price tag. Coinbase said training the Qwen model cost under $100 on a single NVIDIA RTX PRO 6000 96GB GPU.

Newer isn’t always better

The most interesting finding may be what happened with the latest frontier releases. Coinbase reported that Opus 5, Sonnet 5, and GPT-5.6 underperformed their own predecessors on the 16,140-transaction benchmark.

For Coinbase, the takeaway appears to be strategic. The company is leaning toward post-trained models tailored to narrow problems rather than relying on whichever general-purpose model topped the latest leaderboard.

Why this matters beyond Coinbase

A 30% drop in fraudulent transactions means fewer stolen cards slipping through and fewer chargebacks. A 22% cut in fraud value means the losses that do get through are smaller.

The caveats matter. The benchmark is proprietary, so outsiders can’t independently verify the results. The findings also apply to one specific task on one specific product, and fraud patterns shift constantly as attackers adapt.

The regression in newer frontier models also deserves scrutiny over time. It could reflect something specific to how those models handle this kind of structured risk judgment, or it could change with future updates.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Coinbase fine-tunes Qwen3.5-9B model for fraud prevention, beating frontier models
Coinbase fine-tunes Qwen3.5-9B model for fraud prevention, beating frontier models

A small open model trained for under $100 outperformed Opus 4.5 on Coinbase's Onramp fraud benchmark while answering in less than half the time

Coinbase

Coinbase has spent the past stretch teaching a relatively small AI model to catch fraudsters. According to the exchange, the small model now does that job better than some of the most expensive AI systems on the market.

In its “Owning Intelligence” blog series, published October 7-8, 2026, Coinbase said a fine-tuned version of Alibaba’s Qwen3.5-9B model outperformed frontier models on fraud detection for its Onramp service. It was faster and far cheaper to train.

What Coinbase actually built

The project centers on an LLM-based fraud risk agent built specifically for Onramp, Coinbase’s service for buying crypto. The agent sits on top of Coinbase’s existing machine-learning models and looks at transactions that have already cleared those systems.

The agent’s output is deliberately narrow. It returns a risk level and nothing more, so upstream systems keep running as they were.

The results from A/B testing were meaningful. Coinbase said the agent cut fraudulent transactions by 30% and lowered the dollar value of fraud by 22%. The savings were estimated at three times the cost of running the model.

Advertisement

The benchmark showdown

To compare models, Coinbase built a proprietary Onramp fraud benchmark. It contains 16,140 transactions, of which 813 were confirmed fraud cases.

Coinbase fine-tuned Qwen3.5-9B using reinforcement learning. The tuned model then went head-to-head with frontier systems, including Anthropic’s Opus 4.5.

The 9-billion-parameter model came out ahead across four metrics: precision, recall, F1 score, and dollar-weighted recall. The margins ranged from 7.5 to 35.4 percentage points.

Speed also tilted toward the smaller model. Qwen posted a median latency of 0.683 seconds, compared with 1.515 seconds for Opus 4.5. That’s a 55% reduction.

Then there’s the price tag. Coinbase said training the Qwen model cost under $100 on a single NVIDIA RTX PRO 6000 96GB GPU.

Newer isn’t always better

The most interesting finding may be what happened with the latest frontier releases. Coinbase reported that Opus 5, Sonnet 5, and GPT-5.6 underperformed their own predecessors on the 16,140-transaction benchmark.

For Coinbase, the takeaway appears to be strategic. The company is leaning toward post-trained models tailored to narrow problems rather than relying on whichever general-purpose model topped the latest leaderboard.

Why this matters beyond Coinbase

A 30% drop in fraudulent transactions means fewer stolen cards slipping through and fewer chargebacks. A 22% cut in fraud value means the losses that do get through are smaller.

The caveats matter. The benchmark is proprietary, so outsiders can’t independently verify the results. The findings also apply to one specific task on one specific product, and fraud patterns shift constantly as attackers adapt.

The regression in newer frontier models also deserves scrutiny over time. It could reflect something specific to how those models handle this kind of structured risk judgment, or it could change with future updates.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.