Chutes AI trains an 8B model on distributed gaming GPUs, runs it on a phone CPU

Photo: KyoRa Kee / Pexels

Chutes AI trains an 8B model on distributed gaming GPUs, runs it on a phone CPU

Jon Durbin showed off Parallax at the Exploit Summit, a decentralized training system built on 240 consumer graphics cards across 13 countries

Training an AI model usually conjures images of warehouse-sized data centers humming with specialized chips. Chutes AI just did it with the same graphics cards people buy to play video games.

At the Exploit Summit in Montreal, held September 28-29, 2026, Jon Durbin presented an 8 billion parameter model trained on distributed consumer GPUs. The finished model runs on a phone’s CPU.

What Chutes AI actually built

The system is called Parallax, and it is Chutes AI’s approach to decentralized model training. Instead of renting one giant cluster, Parallax stitches together hardware scattered across the globe.

For this run, the setup used 240 RTX 5090 GPUs spread across 30 hosts in 13 countries.

Advertisement

The price tag is the headline number. Chutes AI put the training cost at approximately $6,500, or about $11 per billion tokens processed.

The 8B model reportedly hits approximately 59.6 tokens per second on mobile CPUs, generating text on a phone processor with no cloud connection required.

Chutes AI also showed a larger variant with 40.75 billion parameters. That version ran at 26.9 tokens per second while using 12GB of peak memory.

The privacy and resilience pitch

Local inference was a central selling point of the demo. Because the models run directly on the device, no prompts were sent off the user’s hardware.

Chutes AI says it tackles distributed training challenges with architectures like libp2p, a peer-to-peer networking framework used for efficient syncing between machines.

Where Bittensor fits in

Chutes AI operates as Bittensor subnet SN64, focused on serverless decentralized compute for open-source AI models. SN64’s job is providing compute, and Parallax extends that mission from running models to training them.

Chutes AI previously ran 20B-scale experiments across a range of GPUs, including H100s and RTX 6000 models. The current cluster was built entirely on RTX 5090 consumer cards.

The company plans to publish a comprehensive tech report and the full model.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Chutes AI trains an 8B model on distributed gaming GPUs, runs it on a phone CPU
Chutes AI trains an 8B model on distributed gaming GPUs, runs it on a phone CPU

Jon Durbin showed off Parallax at the Exploit Summit, a decentralized training system built on 240 consumer graphics cards across 13 countries

Photo: KyoRa Kee / Pexels

Training an AI model usually conjures images of warehouse-sized data centers humming with specialized chips. Chutes AI just did it with the same graphics cards people buy to play video games.

At the Exploit Summit in Montreal, held September 28-29, 2026, Jon Durbin presented an 8 billion parameter model trained on distributed consumer GPUs. The finished model runs on a phone’s CPU.

What Chutes AI actually built

The system is called Parallax, and it is Chutes AI’s approach to decentralized model training. Instead of renting one giant cluster, Parallax stitches together hardware scattered across the globe.

For this run, the setup used 240 RTX 5090 GPUs spread across 30 hosts in 13 countries.

Advertisement

The price tag is the headline number. Chutes AI put the training cost at approximately $6,500, or about $11 per billion tokens processed.

The 8B model reportedly hits approximately 59.6 tokens per second on mobile CPUs, generating text on a phone processor with no cloud connection required.

Chutes AI also showed a larger variant with 40.75 billion parameters. That version ran at 26.9 tokens per second while using 12GB of peak memory.

The privacy and resilience pitch

Local inference was a central selling point of the demo. Because the models run directly on the device, no prompts were sent off the user’s hardware.

Chutes AI says it tackles distributed training challenges with architectures like libp2p, a peer-to-peer networking framework used for efficient syncing between machines.

Where Bittensor fits in

Chutes AI operates as Bittensor subnet SN64, focused on serverless decentralized compute for open-source AI models. SN64’s job is providing compute, and Parallax extends that mission from running models to training them.

Chutes AI previously ran 20B-scale experiments across a range of GPUs, including H100s and RTX 6000 models. The current cluster was built entirely on RTX 5090 consumer cards.

The company plans to publish a comprehensive tech report and the full model.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.