Etched Unveils Low-Voltage Inference Chips With Cluster-Scale Memory
Patrick O'Shaughnessy's video and quoted post describe Etched's claimed inference innovations: low-voltage operation for more FLOPs per watt and cluster-scale memory bandwidth. The company says these enable over 80% MFU on trillion-parameter models, versus 20 to 50% on GPUs.
Original post · 1 min read
The combination runs trillion-parameter models at over 80% MFU, where today's GPUs deliver 20 to 50%.
"Everyone said you can't run at voltages lower than GPUs. That was dissatisfying, because plenty of other chips already do.
Bitcoin miners run at under a quarter of the voltage of GPUs. We found a new mechanism to run at much lower voltages, and we think all AI chips in the future will be low voltage chips.
People ask how much memory bandwidth is on your chip. You should be asking how much is on your full scale-up cluster.
We added far more bandwidth at much lower latency from chip to chip."
Patrick OShaughnessy @patrick_oshagThree years ago, two Harvard dropouts set out to build a better AI chip than the largest companies in the world.
Almost everyone I called at the time said it was impossible.
Today, Etched (@Etched) comes out of stealth with $800M total raised, $1B in signed customer contracts, and a working next-gen AI chip.
This was my excuse to ask the two founders, @UbertiGavin and @robertwachen, every question I have about compute and inference.
We discuss:
- Why they built an entire rack and not just a chip
- The two technical bets behind their architecture no one else has tried
- How two founders in …




