The latest generation of AI models achieves breakthrough performance by spending vastly more tokens to solve complex problems. Demand is now compounding faster than the infrastructure beneath it can scale. The constraint sits in how tokens are made. Producing a token is not one operation but many, each placing different demands on hardware, yet every AI system in production runs them all on the same general purpose chip. Moving work between chips has always cost too much energy and time for anyt...See more