Elon Musk plans 660K GB300 GPUs for Colossus 2 by year-end

Elon Musk plans 660K GB300 GPUs for Colossus 2 by year-end

The xAI supercomputer in Memphis is already the world's largest AI data center by power consumption, and it's about to nearly double its GPU count

Elon Musk wants to pack 660,000 NVIDIA GB300 GPUs into his Colossus 2 AI training cluster by late December. If that number sounds absurd, consider that the facility already houses roughly 330,000 B300 GPUs and 110,000 B200 GPUs, making it the largest AI data center on Earth by IT power consumption.

The planned expansion would bring the Memphis complex’s total GPU count well past the half-million mark for Blackwell-series chips alone. In raw compute terms, the facility is projected to reach approximately 1.825 million H100-equivalent units.

What Colossus 2 looks like today

The current installation is already staggering. With its existing 440,000 GPUs (combining B300 and B200 variants), Colossus 2 generates about 1.112 million H100-equivalent compute units.

Advertisement

Powering all of this requires 946 megawatts of IT power. The capital cost so far sits at an estimated $35.8 billion, according to projections from Epoch AI. The full buildout, once the additional GPUs are installed and operational, is modeled at approximately $58 billion.

xAI, operating under the designation SpaceXAI for this project, isn’t keeping the facility to itself either. Leasing agreements reportedly include Anthropic running workloads on the adjacent Colossus 1 cluster at roughly $1.25 billion per month, and Google utilizing approximately 110,000 GPUs.

The power problem

Right now, Colossus 2 relies on a combination of temporary gas turbines as a stopgap measure. These interim power solutions have drawn regulatory scrutiny over environmental emissions.

The longer-term answer is a 1.2 GW permanent power plant in Southaven, Mississippi, expected to reach full operational capacity by mid-2027.

The broader Memphis complex, which encompasses Colossus 1, Colossus 2, and future buildings internally referred to as “Minihard,” aims to eventually exceed 1 million total GPUs with an aggregate power capacity of roughly 2 GW.

Why NVIDIA’s Blackwell generation matters here

The choice of GB300 and B300 GPUs isn’t incidental. NVIDIA’s Blackwell architecture represents a generational leap in AI training efficiency compared to the H100 chips that defined the previous era of large-scale compute buildouts. Each Blackwell GPU delivers substantially more floating-point operations per watt, which is why 612,000 B300 GPUs can deliver the equivalent of 1.8 million H100s worth of compute.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Elon Musk plans 660K GB300 GPUs for Colossus 2 by year-end
Elon Musk plans 660K GB300 GPUs for Colossus 2 by year-end

The xAI supercomputer in Memphis is already the world's largest AI data center by power consumption, and it's about to nearly double its GPU count

Elon Musk wants to pack 660,000 NVIDIA GB300 GPUs into his Colossus 2 AI training cluster by late December. If that number sounds absurd, consider that the facility already houses roughly 330,000 B300 GPUs and 110,000 B200 GPUs, making it the largest AI data center on Earth by IT power consumption.

The planned expansion would bring the Memphis complex’s total GPU count well past the half-million mark for Blackwell-series chips alone. In raw compute terms, the facility is projected to reach approximately 1.825 million H100-equivalent units.

What Colossus 2 looks like today

The current installation is already staggering. With its existing 440,000 GPUs (combining B300 and B200 variants), Colossus 2 generates about 1.112 million H100-equivalent compute units.

Advertisement

Powering all of this requires 946 megawatts of IT power. The capital cost so far sits at an estimated $35.8 billion, according to projections from Epoch AI. The full buildout, once the additional GPUs are installed and operational, is modeled at approximately $58 billion.

xAI, operating under the designation SpaceXAI for this project, isn’t keeping the facility to itself either. Leasing agreements reportedly include Anthropic running workloads on the adjacent Colossus 1 cluster at roughly $1.25 billion per month, and Google utilizing approximately 110,000 GPUs.

The power problem

Right now, Colossus 2 relies on a combination of temporary gas turbines as a stopgap measure. These interim power solutions have drawn regulatory scrutiny over environmental emissions.

The longer-term answer is a 1.2 GW permanent power plant in Southaven, Mississippi, expected to reach full operational capacity by mid-2027.

The broader Memphis complex, which encompasses Colossus 1, Colossus 2, and future buildings internally referred to as “Minihard,” aims to eventually exceed 1 million total GPUs with an aggregate power capacity of roughly 2 GW.

Why NVIDIA’s Blackwell generation matters here

The choice of GB300 and B300 GPUs isn’t incidental. NVIDIA’s Blackwell architecture represents a generational leap in AI training efficiency compared to the H100 chips that defined the previous era of large-scale compute buildouts. Each Blackwell GPU delivers substantially more floating-point operations per watt, which is why 612,000 B300 GPUs can deliver the equivalent of 1.8 million H100s worth of compute.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.