Skip to content

AI's Real Bottleneck Is a Gas Turbine, Not a GPU

#ai-infrastructure #power-generation #risc-v #gpu #memory-chips #data-centers

The compute cart is ahead of the power horses

At a G20 innovation ministers meeting on September 1, Elon Musk made a prediction that should worry anyone building for AI: a clear global electricity shortage next year. His math, as reported via 36Kr: AI chip output grows 40-50% per year, while power supply growth outside China runs 10-20% per year. The slower curve wins.

That gap is reshaping AI infrastructure faster than any model release. Four stories landed in the same news cycle and read like unrelated events: a 50 MW Chinese gas turbine loaded onto a truck in Sichuan, a RISC-V accelerator startup closing a 2 billion RMB round, AMD's rumored next flagship claiming a 15-25% edge over the RTX 5090, and China's CXMT pushing a new memory platform into mass production.

They are one story. AI's constraint chain has moved from algorithms to chips, and now to the physical inputs: power, memory, and manufacturing.

The power wall

Let me put the numbers where they hurt. Microsoft, Meta, and Google are planning clusters of 100,000 GPUs. One cluster, including switches, liquid cooling, and storage, draws 300-500 MW. That's a small city's electricity consumption, for a single building complex. They aren't planning one each. They're planning dozens.

The rack math is just as brutal. A pre-AI rack drew 5-10 kW. A GB200 liquid-cooled rack runs past 100 kW. You don't plug that into existing wiring; you build a power plant next door, a model the industry calls "behind-the-meter" generation.

But you can't buy the plant. The large heavy-duty gas turbine market is a three-company monopoly: GE, Siemens, and Mitsubishi. All three have full order books. A turbine ordered today delivers around 2030, and that's the real reason Musk said he'd build turbines himself. Not for fun. Because nobody will sell him one.

The frustrating part is that turbines are genuinely hard to make. A fighter jet engine runs an hour or two and lands. A heavy-duty turbine runs 8,000+ hours continuously in conditions that melt steel. F-class combustors sit around 1,400°C; H-class approaches 1,600°C, and steel melts between roughly 1,300°C and 1,500°C. The remedy is single-crystal blades with internal cooling channels machined by EDM, plus thousands of laser-drilled film-cooling holes. Cold air bled from the compressor flows through the blade core and exits through those holes, forming a protective film between metal and flame.

That's why three companies have owned this market for decades, and why the shortage won't fix itself.

Key numbers: a 100K-GPU cluster draws 300-500 MW, about one small city's worth. Rack power went from 5-10 kW to 100+ kW with GB200 liquid cooling. New Western large-turbine orders deliver around 2030, and the G50 produces 50 MW, so a single cluster needs six to ten of them.

The G50 and why China can deliver

On August 28, Dongfang Electric loaded its first 50 MW heavy-duty gas turbine, the G50, onto a truck in Deyang, Sichuan, for export. SASAC, China's state asset regulator, featured the shipment on its website days later. That's the tell: this is a national capability statement.

The G50 took 13 years from project start to grid connection. Dongfang launched in 2009, hit full-load testing in 2020, and first fed the grid in late 2022. No drawings, no material formulas, thousands of alloy recipe experiments, hundreds of blade design iterations. The 36Kr report sums it up: there were no shortcuts, only starting from zero and advancing step by step.

Once that 0-to-1 threshold cleared, the rest followed. Shanghai Electric, Harbin Electric, and the national gas turbine program delivered F-class, H-class, and 300 MW machines. Blade casting, film-hole drilling, and control systems that the West treated as state secrets got replicated across institutes and private companies.

The structural advantage is supply chain density. Around Dongfang's Deyang plant, within a few hundred kilometers, sit specialty steel mills, giant forging presses, precision CNC shops, and assembly-and-test lines. GE's equivalent supply chain is fragmented across four countries: blades in the UK, machining in Germany, controls in France, assembly in the US. One broken link delays the whole machine. China's suppliers are close enough to truck parts between them, and once a design is proven, scaling is a matter of will and capacity, not coordination.

That timing lines up almost too neatly with global AI demand. The Middle East has already figured it out. OpenAI's Stargate UAE is a 5 GW compute buildout. Musk and Saudi sovereign funds launched a 500 MW Humain center. Google put $10 billion into a Saudi AI hub. The Gulf has gas it can't use up, desert land at bargain prices, and none of the US interconnection queues. Chinese contractors already won the Taiba2 and Qassim2 gas combined-cycle plants in Saudi Arabia, 3.6 GW combined.

The emerging division of labor: America sells chips and algorithms, the Middle East sells land and gas, China sells turbines and power infrastructure.

Quick Take: China's turbine matters because of lead time, not geopolitics. A Western turbine ordered today lands in 2030. A Chinese turbine deployment at a Gulf site can generate within about a year.

GPU competition makes a comeback

While infrastructure people wrestle with turbines, the consumer front is heating up. GameGPU reports AMD's Radeon RX 10800 XT could beat the RTX 5090 by 15-25% in 4K gaming and local AI workloads. If it holds, that's the first time in two generations AMD has a credible claim on the top performance slot, not just the value position.

That 15-25% means visibly more tokens per second in local inference, assuming VRAM and the software stack measure up. The rumor says nothing about memory capacity, and for local LLMs, VRAM decides which models you can run at all. A card that's faster at 16 GB interests me less than one with 32 GB and a working ROCm path for the framework I actually use.

Reading the thread on this one, my takeaway matched the consensus: more competition in the GPU market is long overdue. When I tested the previous AMD generation for local inference, the hardware was fine. The software path took a full weekend to sort out. That friction, not FLOPS, is what keeps people on NVIDIA.

Memory becomes the quiet constraint

The same week, Reuters reported that CXMT, China's largest memory maker, has pushed a new memory-chip platform into mass production. That sentence deserves more attention than it gets.

Inference is bandwidth-bound. Every generated token requires streaming the model's full weight set through the memory system, and HBM allocation determines whether next-gen GPUs ship at list price or at scalper rates. For two generations, the HBM supply map was simple: SK hynix, Samsung, Micron, and everyone else queued behind them. CXMT's ramp adds a supplier in a category where incumbents can't build fast enough, and it does so under US export restrictions, which is exactly why the ramp matters for Chinese supply chains.

Practical effect: memory pricing stops being a one-way bet, and Chinese server makers get a supply line that sanctions can't easily cut. For anyone buying servers at volume, that's the difference between watching a cost line track one region's fabs or having a genuine second source.

RISC-V goes to the data center

The most interesting hardware story this month is EVAS Intelligence, the RISC-V cloud AI chip unicorn. It closed a nearly 2 billion RMB B+ round, following the 1.5 billion RMB B round and a China Mobile strategic investment earlier this year. Post-money valuation sits near 15 billion RMB, with more than 20 funds in the round.

What did the money buy? A production track record. The first-generation Epoch chips are in volume production, with floating-point performance, interconnect bandwidth, and token throughput for large models rated at the top tier of domestic production silicon. The 64-node supernode cluster, built with liquid-cooled E200-L OAM modules and an orthogonal backplane-free design, is live in a carrier data center. It links 64 chips with sub-100-nanosecond latency and routes MoE experts across all 64 chips cooperatively. That's what lifts inference throughput on mixture-of-experts models: you're no longer pinning all the experts to one chip or a small group.

This is the first RISC-V supernode in production anywhere. Industry chatter has long treated RISC-V as an embedded-core technology, good for microcontrollers and little else. A 64-node liquid-cooled cluster in a real operator facility changes that conversation.

The second generation is already taped out, with several times the compute and interconnect performance of the first. Combined with the second-gen supernode, EVAS claims a 10x gain on prefill-decode separated inference. Ten times on PD separation is the difference between a token factory that pays back in months and one that pays back in years.

The software play matters more than the silicon. EVAS is open-sourcing its VISA virtual instruction set and building the EVACA software stack to integrate with mainstream AI frameworks and domestic stacks like FlagOS. The stated goal is to make VISA the de facto standard connecting RISC-V to the AI software world, bypassing CUDA entirely. Whether that works is the open question, but the direction is clear: China's cloud builders want an open ISA they control, and they're paying unicorn valuations to get one.

One supply chain, four fronts

Strip away the company names and the shape is simple. The AI industry has split into four competitions that are really one contest over physical supply:

Each front has its own actors and constraints:

FrontThe constraintPlayer to watchWhat changed
PowerA cluster needs 300-500 MWDongfang Electric, GE, Siemens, MitsubishiChina's G50 export opens a new turbine supply line; Western order books are full until 2030
MemoryInference is bandwidth-bound; HBM is a chokepointCXMT, SK hynix, Samsung, MicronCXMT's new platform in mass production adds a supply channel outside US control
GPULocal AI and gaming performanceAMD, NVIDIARumored RX 10800 XT claims 15-25% over the RTX 5090
AcceleratorsNVIDIA lock-in and MoE serving costsEVAS, the RISC-V ecosystemFirst production RISC-V supernode; nearly 2 billion RMB round at a ~15 billion RMB valuation

What trips people up

Sizing data centers by GPU count instead of power. I've seen cluster plans obsess over interconnect topology and treat grid interconnection as a formality. In the US, interconnection queues run for years. A 100K-GPU cluster without a signed power agreement is a warehouse full of silicon.

Assuming a gas turbine is a commodity gas engine. Converting a retired jet engine into a peaking generator works; the 36Kr piece notes a Yantai company did exactly that and sold the result for $1.465 billion. But peaking is not baseload. Single-crystal blades and film cooling are the entire moat, and they're why fast-follower vertical integration can't shortcut the incumbents.

Buying on leaked GPU specs. The 15-25% RX 10800 XT claim is a rumor, and local AI performance is a function of VRAM, quantization support, and the software stack, not headline FLOPS. Wait for measurements on real model workloads before planning purchases.

Treating RISC-V accelerators as drop-in CUDA replacements. The chip is the easy part; the software stack is the real cost. EVAS's VISA and EVACA work exists precisely because model migration is the adoption barrier. Budget for the migration, not just the hardware.

Ignoring memory bandwidth. Doubling FLOPS gets you nothing if the HBM or DDR5 feed can't keep up. CXMT's ramp is a supply story, but it's also a reminder that the memory system is the silent co-processor in every inference budget.

One thing to remember

Every AI company is now an energy company and a hardware company whether it wants to be or not. The model is only as good as the power source behind it, the memory feeding it, and the supply chain that manufactured both. Teams that plan around physical constraints will ship on schedule. Teams that don't will watch their clusters sit dark.

The Bottom Line

If you're planning a large training or inference cluster, put power procurement on the critical path before the GPU orders. A turbine ordered from GE, Siemens, or Mitsubishi today arrives around 2030, while a Middle East deployment with Chinese turbines can close within about a year.

If you're doing local AI on a budget, buy on VRAM and verified software support, not leaked performance numbers. A 15-25% edge in rasterization is irrelevant if the model you need doesn't fit on the card or the framework path costs you a weekend.

One thing to watch: the RISC-V supernode story is moving on a one-year generation cadence, with second-gen silicon already taped out. Expect the first credible RISC-V versus CUDA cluster benchmarks within 6-12 months, probably at Chinese carrier data centers first. If those numbers hold, the accelerator map rewrites itself.