Analyst Argues Hardware-Algorithm Codesign Will Decide AI Hardware Winners
Bubble Boi argues that new accelerators offer inference features with no GPU or TPU equivalent, making hardware-algorithm codesign the decisive competitive frontier, and quotes Gavin Baker on diverging scale-up architectures eroding model portability across chips.
Original post · 1 min read
I know of several features that new accelerators are adding that have no analogous operation on GPU, TPU, etc. most of these “special tricks” are on the inference domain and not only lower the TCO but also increase the model quality & capabilities.
Hardware-algo codesign is the last frontier left and it’s going to be the area that picks the winner in the end.
Gavin Baker @GavinSBakerMuch of Dwarkesh's argument hinges on this statment which *was* accurate but will be increasingly inaccurate on a go forward basis imo:
“American labs port across accelerators constantly. Anthropic's models are run on GPUs, they're run on Trainium, they're run on TPUs. There are so many things you can do, from distilling to a model that's well fit for your chips.”
As system level architectures diverge (torus vs. switched scale-up topologies, memory hierarchies, networking primitives), true portability is eroding. The Mi300 and Mi325 had roughly the same scale-up domain size as Hopper whil…



