60 points krackers 57 minutes ago 21 comments
wolttam 24 minutes ago | parent
krm01 22 minutes ago | parent
kibae 12 minutes ago | parent
Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.
speedgoose 19 minutes ago | parent
thehamkercat 15 minutes ago | parent
rozab 14 minutes ago | parent
bayindirh 7 minutes ago | parent
Keeping the garage door open, or at least making the door translucent. It's always cool.
liuliu 14 minutes ago | parent
lucrbvi 11 minutes ago | parent
jampekka 6 minutes ago | parent
I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.
https://en.wikipedia.org/wiki/Training,_validation,_and_test...
liuliu 2 minutes ago | parent
SwellJoe 4 minutes ago | parent
ProfessorLayton 13 minutes ago | parent
For some reason I thought training took much, much longer than what the progress bar suggests.
This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.
joelwallis 12 minutes ago | parent
The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.
-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
james2doyle 8 minutes ago | parent
I always found that those Mimo models to be really good at tool calling and following instructions