Literally 3 days before openai announced the navier-stokes stuff, (because there were rumors about it) I was commenting to someone "if someone actually uses ai to prove navier-stokes, I will consider LLMs one of the more substantial technological developments ever."
(granted they found a counter-example, which is perhaps less impressive, since counter-examples seem to be the more common success story for LLMs looking at big problems, since they are presumably easier to "brute-force." + some of the drama w.r.t. the external math researchers uploading their work into chatgpt, and it possibly being used to train the new model...)
quintoshsome analysts predict that local AI models can be as good as AI that ran in a datacenter half a year ago, it's not just the hardware improving but several other factors too like compression techniques and algorithmic efficiencies. A GTX 3090 can run qwen 3.8-27b which gets a better AI intelligence score (doesn't just include googling information) than Claude Opus 4.5, GPT-5.2 and Gemini 3 7 months ago
IIRC, qwen has to use a ton of tokens to get that quality on the benchmark. So it's small/fast model, but it needs to spit out a ton of tokens to work that well, which may cancel out some of the performance and memory benefits.
Something interesting: nvidia is potentially struggling to make tensor cores faster. part of why they've added 4- and 8-bit precisions is that they were unable to really make 16-bit tensor cores faster (according to someone at nvidia)