109 points bryan0 1 hour ago 62 comments
gigatexal 1 hour ago | parent
aabhay 1 hour ago | parent
refulgentis 57 minutes ago | parent
(which has led me to believe that's a good approximation for hedonic adaptation, I've seen tons of attempts at demonstrating nerfing via benches, none persist)
bbg2401 9 minutes ago | parent
It’s frustrating to observe communities made up of smart, professional individuals as they behave like spoiled children on the day after Christmas when new toy novelty has begun to wane.
I understand it’s relatively harmless but for goodness sake, take a step back and appreciate what you have instead of immediately wanting the thrill of a newer model. Slow down and do deliberate work to get the most out of these amazing tools. Don’t just live off the temporary thrill of finding something marginally better than what you have.
aloukissas 56 minutes ago | parent
judge2020 54 minutes ago | parent
solenoid0937 51 minutes ago | parent
solfox 49 minutes ago | parent
nba456_ 48 minutes ago | parent
AnimalMuppet 45 minutes ago | parent
voiceeh 45 minutes ago | parent
jyoung8607 34 minutes ago | parent
If so, please share. This should be measurable, and I'm glad this project is measuring it.
Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.
dude250711 47 minutes ago | parent
empath75 45 minutes ago | parent
wccrawford 44 minutes ago | parent
raincole 44 minutes ago | parent
jascha_eng 33 minutes ago | parent
It would be economical suicide from anthropic and OpenAI to actually need models intentionally.
But hey I guess it's hard with technology that truly seems like magic. People say if you'd bring electricity to the middle ages you'd be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn't ready for yet.
solfox 44 minutes ago | parent
After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.
I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?
whs 42 minutes ago | parent
madeofpalk 37 minutes ago | parent
ENGNR 31 minutes ago | parent
Computer0 25 minutes ago | parent
gr_norm 42 minutes ago | parent
The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.
nico 40 minutes ago | parent
The quality of the output/work seems the same, but the speed at which it gets stuff done is a lot slower, because it's asking for permission so much more
I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output
LeoPanthera 37 minutes ago | parent
colordrops 37 minutes ago | parent
Razengan 37 minutes ago | parent
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
jug 36 minutes ago | parent
https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Grimblewald 17 minutes ago | parent
Rapzid 3 minutes ago | parent
Of course it's almost entirely unsubstantiated BS.
octoberfranklin 36 minutes ago | parent
Open models are the endgame.
paradox460 32 minutes ago | parent
octoberfranklin 17 minutes ago | parent
The classifier is a model; it examines the actual prompt.
They already do this for the safety "guardrails".
johnfn 34 minutes ago | parent
I made a graphic to explain why people feel like the models get nerfed:
https://x.com/thesilenceturns/status/2103551351825543610
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
hbn 31 minutes ago | parent
My work at my job has stayed the same. But the model quality has varied.
They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.
Computer0 24 minutes ago | parent
johnfn 23 minutes ago | parent
> It’s not a crazy conspiracy that the same model can be stupider
Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.
frde_me 22 minutes ago | parent
This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)
With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.
Grimblewald 21 minutes ago | parent
My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.
How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?
gobdovan 18 minutes ago | parent
Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.
prodigycorp 17 minutes ago | parent
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Your chart is wrong.
swader999 14 minutes ago | parent
johnfn 12 minutes ago | parent
I am more skeptical about the compute provider claim - do you have any evidence of that?
HawtAds 11 minutes ago | parent
bitexploder 7 minutes ago | parent
physicallyIllfr 8 minutes ago | parent
By the way, dont for a second think LLM hourly limits are all about revenue, they're playing into this psychology. They hire literal gambling UX designers, they want to turn you all into addicts. They want to make you reliant.
Want to run your llm like a slot machine? They'll let you do that spin the generation on a multiple, get 6x results, pick your favorite. Feel that high.
Just know you can get that same hit of dopamine by fostering your own intelligence and creating something with it. Token dealers are just selling you the shortcut, straight to the reward, short circuiting the the natural process.
Bad times ahead for many. This shit isnt good for your brain. And you all know the truth, you just wont admit it. Its doing damage, making you lazier, less intelligent.. Making you an addict.
xlayn 27 minutes ago | parent
Yesterday I fought claude fable to not just jump to make changes like a dog following a treat, that we were researching... at some point I introduced the word HAWAI... and only if I say HAWAI the thing can start making changes..
I was going to post here in HN just to have a "I knew this was the reason" when they release fable > 5.1
I had the exact same feeling every time they have a new big release
bethekidyouwant 17 minutes ago | parent
zeroonetwothree 7 minutes ago | parent