Photo: Tima Miroshnichenko / Pexels
Chutes AI and Harvard release public dataset of 6.12 billion LLM requests
The year-long dataset spanning 9,174 models reveals that 99% of requests are repeats within 15 minutes, a finding with major implications for how AI infrastructure gets built.
If you’ve ever wondered what billions of AI requests actually look like under the hood, now you can find out. Chutes, a decentralized AI inference platform running on Bittensor’s Subnet 64, has teamed up with Harvard University researchers to release what may be the largest public dataset of LLM serving metadata ever assembled.
The numbers are staggering: 6,122,413,756 requests across 9,174 models, generated by 314,970 anonymized users over a full year of production traffic. The dataset spans from April 11, 2025, to April 12, 2026, and is now freely available through GitHub and a Harvard S3 bucket.
What’s actually in this thing
The dataset captures metadata, not the actual conversations people had with AI models. It includes request timing, token counts, latency measurements, and time-to-first-token (TTFT) metrics, but zero prompts or responses. Over the year-long period, Chutes processed roughly 35.8 trillion input tokens and 2.52 trillion output tokens.
User identifiers in the dataset rotate every three months, adding another layer of anonymization. The accompanying documentation, co-authored by researchers from Harvard, the University of Chicago, and Chutes, provides the kind of detailed methodology notes that make the data actually usable for academic work.
The 99% repeat problem
Perhaps the most consequential finding buried in this dataset: 99% of repeat requests occur within a 15-minute window. The research team found that prefix-aware routing strategies can achieve nearly optimal cache-hit rates with only slight load imbalances across servers.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
Output lengths have been declining over the dataset’s timespan, dropping from hundreds of tokens per response to fewer than 100. The researchers interpret this as evidence of a shift toward agentic or machine-driven queries, where automated systems are making quick, targeted requests rather than humans asking for lengthy explanations.
The Bittensor connection
Chutes operates as part of Bittensor’s decentralized network, specifically Subnet 64, running open-source LLMs on a distributed GPU infrastructure. Payments on the platform flow through TAO, Bittensor’s native token.
The collaboration between Chutes and Harvard included an opt-in period from March to July 2026 where researchers received a 25% discount for contributing their usage data to the dataset.