Comment Re: We are going so fast we need to slow down! (Score 1) 84
One of the neat things about Deep Seek is they actually do something OAI and Anthropic do not;- They actually publish their methods, training code and all that.
Theres actually a lot more going on with these models than "Trained it on ChatGPT" (And its not like the chinese are alone in that. When claude first came out, it kept randomly claiming to be ChatGPT, a smoking-gun level sign of distillation, and OIA has *heavily* implied that they've caught XAI out multiple times trying to distill CGPT. And to be clear , amongst ML researchers, distillation is considered an entirely legitimate and normal part of AI training.). What is clever about deep seek is they've made pretty significant progress on getting a very big model and running it cheap and fast on more commodity software. And importantly , they've published how they do it in a reproducable way. They've also made *significant* progress on training methods that drastically lower the cost of training up a model the hard way.
So yeah, its more than just distilling western models. The chinese are actually making real progress. And..... Deepseek and Kimi and the like are actually really good models. You can do 90% of what the big american frontier models are doing with them, and on a budget, 90% is good enough.