Paul Kedrosky
Paul Kedrosky
@paul@paulkedrosky.com · Mar 3

One of the dynamics playing out as model improvement slows, and LLMs move toward being inference-centric, is that new silicon is emerging that runs cooler, with more caching, and much faster speeds.

For example, check Taalas's demo here to see 70x token production speedup, often approach 16,000 tok/s. https://chatjimmy.ai/