Xeons have a much longer shelf life and diverse workloads. If you order hardware specifically for LLM inference and then some new hardware/model combination is much better at that (which it will be, because a lot of people are working on that), you might be in trouble.
It's like setting up a warehouse of GPUs to mine bitcoin while others are switching to ASICs.
No I mean inference. The idea is that inference demand will be massive and a race to the bottom with razor thin margins.
Training costs can be amortized over the entire lifetime of the model, but if you lose money on inference or can't offer competitive usage limits for subscribers, there's no amortizing that.
No it's all about having the top model first and training time is what's crucial. OpenAI has already shown willingness to bleed money for the sake of brand and we can expect that to continue.
It's like setting up a warehouse of GPUs to mine bitcoin while others are switching to ASICs.