88 points seelos 1 hour ago 37 comments
mydreamof 1 hour ago | parent
harmonic18374 55 minutes ago | parent
Also the submitter's account is very new which makes me suspicious of self-promotion.
_doctor_love 1 hour ago | parent
Tsarp 1 hour ago | parent
airstrafer 44 minutes ago | parent
Maybe still worth it if their "64% cheaper" figure holds.
Bolwin 20 minutes ago | parent
xlbuttplug2 32 minutes ago | parent
I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models.
scronkfinkle 59 minutes ago | parent
samyok 43 minutes ago | parent
:)
Disclaimer: I work at Cognition, although was not involved in SWE-2
scronkfinkle 27 minutes ago | parent
CamperBob2 7 minutes ago | parent
monkeydust 54 minutes ago | parent
postalcoder 45 minutes ago | parent
Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
enraged_camel 30 minutes ago | parent
nullbio 27 minutes ago | parent
thereitgoes456 20 minutes ago | parent
While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others.
Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.
mediaman 27 minutes ago | parent
Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
felixgallo 18 minutes ago | parent
general_reveal 5 minutes ago | parent
letmevoteplease 4 minutes ago | parent
Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases.
> does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities)
And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves.
general_reveal 12 minutes ago | parent
You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world.
Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase. That means smart money would have to know the valuation of model makers needs to drop, but not necessarily hardware makers. There will be massive marketing pushback from Anthropic and OpenAI to keep the illusion up that models havent been commodified, but the acceleration of lying is exposing them. Invest wisely :)
Cheers.
iLoveOncall 6 minutes ago | parent
Yes? Just like every single model from every single AI lab.
eranation 10 minutes ago | parent
TheJCDenton 44 minutes ago | parent
On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.
nullbio 40 minutes ago | parent
I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).
The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.
eru 21 minutes ago | parent
Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.
ltsSmitty 37 minutes ago | parent
llmslave 37 minutes ago | parent
I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.
Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job
hnedeotes 26 minutes ago | parent
eyeris 34 minutes ago | parent
The write-up from yesterday was by somebody from cognition using Devin to translate existing cpu sieving methods to gpu and to optimize the gpu sieve.
bobtheborg 31 minutes ago | parent
Looking forward to 2 -- maybe it'll be usable
bluelightning2k 23 minutes ago | parent
I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex.
pkilgore 16 minutes ago | parent
gruez 6 minutes ago | parent
https://www.youtube.com/watch?v=tNmgmwEtoWE
As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.