312 points crorella 1 hour ago 244 comments
hlynurd 59 minutes ago | parent
tedsanders 58 minutes ago | parent
thejazzman 56 minutes ago | parent
https://amphetamem.es/meme?id=the-simpsons_06_12_71&text=We%...
pkulak 57 minutes ago | parent
This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.
prodigycorp 58 minutes ago | parent
These moves all make sense when you take into account the enterprise market.
aaronbrethorst 57 minutes ago | parent
t-sauer 57 minutes ago | parent
nsingh2 54 minutes ago | parent
SirMaster 54 minutes ago | parent
algoth1 50 minutes ago | parent
oh_no 18 minutes ago | parent
gradus_ad 57 minutes ago | parent
nojito 55 minutes ago | parent
I remember when bandwidth was super expensive and now it’s dirt cheap.
vanviegen 39 minutes ago | parent
iAMkenough 12 minutes ago | parent
Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975
mixdup 43 minutes ago | parent
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
semiquaver 40 minutes ago | parent
Edit: removed a comment that was uncharitable and rude, for which I apologize.
ActionHank 38 minutes ago | parent
We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.
CuriouslyC 27 minutes ago | parent
mixdup 38 minutes ago | parent
phoghed 36 minutes ago | parent
arctic-true 26 minutes ago | parent
famouswaffles 17 minutes ago | parent
semiquaver 9 minutes ago | parent
CamperBob2 13 minutes ago | parent
serf 38 minutes ago | parent
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
mixdup 35 minutes ago | parent
password54321 29 minutes ago | parent
LPisGood 38 minutes ago | parent
theturtletalks 36 minutes ago | parent
colechristensen 33 minutes ago | parent
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
CuriouslyC 30 minutes ago | parent
xienze 30 minutes ago | parent
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
redanddead 21 minutes ago | parent
azan_ 18 minutes ago | parent
People were talking about plateau for years already.
sebzim4500 10 minutes ago | parent
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
luma 3 minutes ago | parent
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
jorblumesea 40 minutes ago | parent
it's also why there have been so many calls for regulation and slowdowns.
jimbob45 27 minutes ago | parent
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
simianwords 7 minutes ago | parent
amelius 57 minutes ago | parent
mholm 44 minutes ago | parent
condour75 44 minutes ago | parent
SkyBelow 33 minutes ago | parent
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
phpnode 57 minutes ago | parent
jesse_dot_id 56 minutes ago | parent
sharpshadow 51 minutes ago | parent
wg0 48 minutes ago | parent
Wheen 6 minutes ago | parent
Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
LPisGood 35 minutes ago | parent
mckirk 51 minutes ago | parent
system2 50 minutes ago | parent
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
wg0 35 minutes ago | parent
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
copperx 9 minutes ago | parent
jonatron 49 minutes ago | parent
SwabbyNat74 49 minutes ago | parent
tjwebbnorfolk 48 minutes ago | parent
colpabar 47 minutes ago | parent
infamouscow 36 minutes ago | parent
Aboutplants 46 minutes ago | parent
scrollop 38 minutes ago | parent
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
sockaddr 35 minutes ago | parent
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
geeky4qwerty 19 minutes ago | parent
mattnewton 45 minutes ago | parent
motoboi 43 minutes ago | parent
mynameisjonny_ 42 minutes ago | parent
agluszak 41 minutes ago | parent
orbital-decay 40 minutes ago | parent
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring for some.
jchw 38 minutes ago | parent
esafak 31 minutes ago | parent
toasty228 31 minutes ago | parent
az226 10 minutes ago | parent
cmrdporcupine 57 minutes ago | parent
Lapalux 57 minutes ago | parent
fraywing 56 minutes ago | parent
Astra is a pretty impressive model. Excited to try this.
Nevin1901 56 minutes ago | parent
glimshe 55 minutes ago | parent
IshKebab 55 minutes ago | parent
rs_rs_rs_rs_rs 50 minutes ago | parent
Edit: for context, just Steam alone has ~200million monthly active users.
paulryanrogers 40 minutes ago | parent
How many DCs are devoted solely to gaming?
lp92 37 minutes ago | parent
HelloMcFly 33 minutes ago | parent
rs_rs_rs_rs_rs 23 minutes ago | parent
Yes but it adds up when you consider that just on Steam alone there are 200 million monthly active users.
empthought 9 minutes ago | parent
rs_rs_rs_rs_rs 24 minutes ago | parent
An entire planet. Just Steam alone has one or two hundres million monthly active users.
lbrito 31 minutes ago | parent
rs_rs_rs_rs_rs 20 minutes ago | parent
Yeah? Show me the big movements against computer gaming.
JDups 27 minutes ago | parent
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
otterley 47 minutes ago | parent
Bolwin 35 minutes ago | parent
sergiotapia 24 minutes ago | parent
barrenko 55 minutes ago | parent
Starlevel004 55 minutes ago | parent
dcchambers 54 minutes ago | parent
iamdelirium 54 minutes ago | parent
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
A_D_E_P_T 54 minutes ago | parent
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
ekun 50 minutes ago | parent
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
godwinson__4-8 41 minutes ago | parent
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.
lukan 35 minutes ago | parent
CuriouslyC 25 minutes ago | parent
therealdrag0 16 minutes ago | parent
minimaxir 52 minutes ago | parent
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
TuxSH 47 minutes ago | parent
bigwheels 38 minutes ago | parent
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
Infinity315 35 minutes ago | parent
toasty228 34 minutes ago | parent
I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous
copperx 27 minutes ago | parent
AndrewKemendo 33 minutes ago | parent
squidbeak 22 minutes ago | parent
edgyquant 19 minutes ago | parent
colinhb 22 minutes ago | parent
rspeele 12 minutes ago | parent
Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours.
The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.
The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.
Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.
Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.
phoghed 3 minutes ago | parent
They form these super strong opinions after a few prompts, then face reality over time.
People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.
mmis1000 29 minutes ago | parent
However it's less willing to obey your instruction so it's less usable for general runtine flows.
jauntywundrkind 26 minutes ago | parent
(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)
dotancohen 16 minutes ago | parent
> Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?peterbell_nyc 4 minutes ago | parent
There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
TuxSH 7 minutes ago | parent
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
dom96 10 minutes ago | parent
joshstrange 29 minutes ago | parent
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
throwitaway222 52 minutes ago | parent
Guess not?
murbard2 49 minutes ago | parent
minimaxir 48 minutes ago | parent
nimonian 47 minutes ago | parent
alvis 51 minutes ago | parent
slopinthebag 51 minutes ago | parent
sehw 48 minutes ago | parent
mekpro 47 minutes ago | parent
jdw64 47 minutes ago | parent
godwinson__4-8 44 minutes ago | parent
When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.
When do we get the next jump in capability? When is 6.1 Astra released?
ColonelPhantom 36 minutes ago | parent
godwinson__4-8 32 minutes ago | parent
The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.
Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.
Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.
wren6991 26 minutes ago | parent
After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.
vb-8448 44 minutes ago | parent
tultra 44 minutes ago | parent
gavin_gee 41 minutes ago | parent
nicce 41 minutes ago | parent
thefounder 39 minutes ago | parent
The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable/Claude just cannot get/fix even when you point it.
the_duke 38 minutes ago | parent
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
nxc18 34 minutes ago | parent
jstummbillig 29 minutes ago | parent
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents)
the_duke 21 minutes ago | parent
nicce 18 minutes ago | parent
copperx 17 minutes ago | parent
Marha01 5 minutes ago | parent
Eridrus 5 minutes ago | parent
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
epolanski 37 minutes ago | parent
modeless 35 minutes ago | parent
mkaic 35 minutes ago | parent
slekker 28 minutes ago | parent
lynx97 34 minutes ago | parent
Aboutplants 34 minutes ago | parent
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...
Yikes
surgical_fire 21 minutes ago | parent
The only way is for prices to go up. Way up.
mrtesthah 4 minutes ago | parent
moregrist 20 minutes ago | parent
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
5555watch 6 minutes ago | parent
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
TomGarden 20 minutes ago | parent
Our VC-backed subscription days are numbered
m3kw9 7 minutes ago | parent
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
glaslong 7 minutes ago | parent
honkycat 19 minutes ago | parent
I can justify $200/mo but more than double is not appealing to me.
WinstonSmith84 14 minutes ago | parent
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
cmrdporcupine 3 minutes ago | parent
Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.
I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.
Whether that survives contact with Chinese open weight models is hard to say.
enraged_camel 1 minute ago | parent
MCArth 1 minute ago | parent
spiderice 4 minutes ago | parent
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
torginus 9 minutes ago | parent
LeBit 6 minutes ago | parent
intenex 33 minutes ago | parent
zarzavat 26 minutes ago | parent
toasty228 22 minutes ago | parent
nater5000 18 minutes ago | parent
They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.
sergiotapia 33 minutes ago | parent
{"type":"item.completed","item":{"id":"item_0","type":"error","message":"Model metadata for `gpt-6.1-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues."}}
sampton 33 minutes ago | parent
TomGarden 32 minutes ago | parent
Alifatisk 31 minutes ago | parent
One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.
In other news:
> In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.
arctic-true 29 minutes ago | parent
MetaverseClub 20 minutes ago | parent
neosat 20 minutes ago | parent
The last time a model announcement felt like a leap in capability beyond other things out there was Fable - which was promptly taken away. Sol and recently Opus 5.5 were strong because they approach that capability with a lot more efficiency and don't blabber incoherently (looking at you Opus 5.1).
Deepseek is a workhorse for those who prefer open and API usage. Other than that the model announcements all just seem like a blur and quite interchangeable but I wonder if that's just me tuning out or do others feel the same way?
dom96 18 minutes ago | parent
MisterMunchkin 13 minutes ago | parent
joduplessis 10 minutes ago | parent
itzikkatz 7 minutes ago | parent
moinism 6 minutes ago | parent
prometheus1992 6 minutes ago | parent
dangoodmanUT 3 minutes ago | parent
ghm2180 2 minutes ago | parent
Fuck altruism, ammi right? lets make money, gobs of it by screwing the middle users as much as we can to push them into just two tiers: Ones that use it for recreation and others that pay through their noses.
ChaseRensberger 1 minute ago | parent
If OpenAI cuts alternative harness support it will be a weird day trying to figure out what to do next, it's been so clearly the best bang for your buck (imo) for a while. maybe id finally have to give smaller models a try.
anything to avoid using the dogwater codex & claude code tuis.
anyways this seems like a nice cost improvement over GPT 6 Sol and I expect this will be my new daily driver.