Analytics24
3 September
My list

The Long Read · 24 Aug

The GPU is sold. The stack is not.

The chip can be sold while the rest of the stack is still a promise. Power, water, and the hall decide.

An unmarked wafer boat and two unmarked memory stacks on a bright clean bench.
Wafer boat on the clean bench.

NVIDIA sold the GPU. The stack is the other queue.

3

Qualified HBM4 suppliers on Huang’s June 2026 telling, Bloomberg and Reuters. Cantor and Wedbush have been running sold-out through 2026 and into 2027. Their claims.

Wafers and stacks.

Model shoptoken list priceWorkflow ownerseats · usagePowerqueue
A GPU that has wafers and a GPU that has stacks are different animals. Schematic, not a fitted beta.

The public argument is demand. Tokens, training runs, inference at the edge of a product that did not exist three years ago. That argument is loud, and it is mostly priced. The private argument is a slab of memory that has to be stacked, tested, and qualified before the GPU you already paid for can ship.

SK Hynix, at Hot Chips in August 2026, Herald Business among them, said HBM4 in a twelve-layer stack was in mass production and a sixteen-layer stack was in qualification. Jensen Huang said, that all three qualified suppliers were in production and racing for Vera Rubin. Cantor and Wedbush have been running a sold-out-through-2026 and into-2027 line. Those are their claims. They are. They are the weather the tape is already using.

Allocation is. A GPU maker who has stacks is a different animal from a GPU maker who has wafers. It is the rent. It is a short list. Short lists do not mean infinite price.

What people are ignoring is the shape of the bottleneck. Last cycle the story was silicon. Then it was memory in the ordinary sense, DRAM, a node, a bit. It is a stack, a through-silicon via, a yield on a height that did not exist when the last cycle’s models were trained. You can add wafer starts and still miss the stack. You can have a GPU design win and still sit in a warehouse because the high-bandwidth memory that makes the win real is allocated to someone else’s rack. It is the spread between the companies that can qualify a taller stack and the companies that can only announce one.

The visionary version is that the next decade of models is rented at the stack. It is a reservation. That is a quieter market than a keynote. It lives in qualification letters, in who got the taller sample, in which cloud already booked 2027. If that market is the real one, share prices that treat every GPU print as a clean demand print are reading the wrong line.