225 points D2OQZG8l5BI1S06 1 hour ago 153 comments
ramish94 1 hour ago | parent
Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
bigyabai 1 hour ago | parent
OpenAI and Anthropic's lead is vanishingly small at this point.
SubiculumCode 1 hour ago | parent
TuxSH 58 minutes ago | parent
Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.
Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.
bbor 50 minutes ago | parent
level87 59 minutes ago | parent
bbor 51 minutes ago | parent
johnmlussier 1 hour ago | parent
This is bollocks. Their safeguards are shit.
AshamedBadger56 1 hour ago | parent
sebzim4500 58 minutes ago | parent
In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.
AshamedBadger56 53 minutes ago | parent
polski-g 23 minutes ago | parent
bbor 48 minutes ago | parent
The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!
icedchai 55 minutes ago | parent
tom1337 54 minutes ago | parent
jchw 52 minutes ago | parent
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
film42 48 minutes ago | parent
jauntywundrkind 23 minutes ago | parent
newspaper1 46 minutes ago | parent
solenoid0937 46 minutes ago | parent
https://support.claude.com/en/articles/14604842-real-time-cy...
> This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models
You obviously should not expect the CVP to cover this model either.
machomaster 40 minutes ago | parent
solenoid0937 40 minutes ago | parent
gowld 33 minutes ago | parent
solenoid0937 10 minutes ago | parent
giancarlostoro 43 minutes ago | parent
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
rfgplk 40 minutes ago | parent
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.
nightpool 32 minutes ago | parent
AIorNot 29 minutes ago | parent
dom96 5 minutes ago | parent
pookieinc 1 hour ago | parent
radial_symmetry 59 minutes ago | parent
lanthissa 59 minutes ago | parent
WinstonSmith84 59 minutes ago | parent
bpodgursky 50 minutes ago | parent
jrflo 40 minutes ago | parent
iagocc 1 hour ago | parent
s3p 1 hour ago | parent
takerofnaps 1 hour ago | parent
wongarsu 1 hour ago | parent
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
ttul 59 minutes ago | parent
taurath 59 minutes ago | parent
nicoburns 47 minutes ago | parent
Sol- 58 minutes ago | parent
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
phainopepla2 49 minutes ago | parent
Imustaskforhelp 33 minutes ago | parent
In short, seems to describe vibe-coding to me? What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.
> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
[0]: https://news.ycombinator.com/item?id=49808422
[1]: https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-no...
RGS1811 20 minutes ago | parent
For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.
schwarzrules 48 minutes ago | parent
doctoboggan 48 minutes ago | parent
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
bbor 46 minutes ago | parent
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding.
Yes. Our career is over, as is our economy. Soooo... FYI :(jobs_throwaway 18 minutes ago | parent
> the economy is over
Hackernews' neuroticism remains undefeated
JMKH42 46 minutes ago | parent
Imustaskforhelp 45 minutes ago | parent
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
> And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
gregwebs 44 minutes ago | parent
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
alansaber 39 minutes ago | parent
chrismustcode 22 minutes ago | parent
Not quite sure where this fits well. Maybe small one one off requests like using Claude desktop/web?
maherbeg 21 minutes ago | parent
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
afro88 19 minutes ago | parent
rtuin 58 minutes ago | parent
avree 58 minutes ago | parent
onlyrealcuzzo 57 minutes ago | parent
> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).
For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.
This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.
Hopefully they release a Haiku that actually has a reason for existing.
criemen 54 minutes ago | parent
enraged_camel 34 minutes ago | parent
>> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
jchw 54 minutes ago | parent
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
eli 51 minutes ago | parent
jchw 46 minutes ago | parent
MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.
I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.
I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.
Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.
eli 37 minutes ago | parent
GLM 5.3 Flash is also very good. I think a little smarter and a little more expensive.
velcrovan 49 minutes ago | parent
cbg0 21 minutes ago | parent
It's been out for an hour and you've already concluded this?
wkcheng 57 minutes ago | parent
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
SubiculumCode 46 minutes ago | parent
ricardobeat 42 minutes ago | parent
wkcheng 36 minutes ago | parent
oh_no 15 minutes ago | parent
i think they see what openai charges for luna and just don't want to try and compete
solenoid0937 41 minutes ago | parent
RussianCow 40 minutes ago | parent
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
delillos 40 minutes ago | parent
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
usaar333 35 minutes ago | parent
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
Jcampuzano2 31 minutes ago | parent
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
dominotw 18 minutes ago | parent
_fw 55 minutes ago | parent
I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.
Give me the frontier, or give me the cheapest form of good enough.
calumcl 37 minutes ago | parent
Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.
EMM_386 35 minutes ago | parent
Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.
abejora 54 minutes ago | parent
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
eli 53 minutes ago | parent
abejora 46 minutes ago | parent
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
joeyhage 3 minutes ago | parent
radlad 45 minutes ago | parent
> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...
I cannot find a Sonnet 5.5 system card.
abejora 41 minutes ago | parent
Leary 45 minutes ago | parent
manojlds 26 minutes ago | parent
oh_no 22 minutes ago | parent
AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k
Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.
bayesianbot 52 minutes ago | parent
ChickeNES 51 minutes ago | parent
ahriad 50 minutes ago | parent
square_usual 46 minutes ago | parent
simianwords 34 minutes ago | parent
Some tasks are reasoning shaped by nature and you can't just throw a big model at it.
Jcampuzano2 46 minutes ago | parent
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
ghoshbishakh 45 minutes ago | parent
yapfrog 42 minutes ago | parent
solenoid0937 41 minutes ago | parent
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
ricardobeat 40 minutes ago | parent
solenoid0937 37 minutes ago | parent
The difference in perception for Opus 5.5 on HN vs the real world is what convinced me HN is totally detached from reality.
alansaber 40 minutes ago | parent
limsungkee 39 minutes ago | parent
ghoshbishakh 38 minutes ago | parent
In xhigh effort it is a lot cheaper and possibly lot less impressive?
dack 38 minutes ago | parent
s314 35 minutes ago | parent
zozbot234 28 minutes ago | parent
lol, MiMo 2.6 Pro basically matches Sonnet 5.5 high (mind you, not xhigh or max) at a far lower price point.
tombert 33 minutes ago | parent
croemer 31 minutes ago | parent
alasano 30 minutes ago | parent
enraged_camel 28 minutes ago | parent
If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.
OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.
It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...
Alifatisk 27 minutes ago | parent
> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
simonw 27 minutes ago | parent
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Here's how the thinking effort levels compare:
low
27 input, 1,623 output, thinking_tokens: 0
1.6284
Duration: 10138ms (10s)
medium
27 input, 1,796 output, thinking_tokens: 0
1.7914 cents
Duration: 11266ms (11s)
high
27 input, 2,334 output, thinking_tokens: 745
2.3394 cents
Duration: 17376ms (17s)
xhigh
27 input, 5,730 output, thinking_tokens: 2535
5.7354 cents
Duration: 41882ms (41s)
max (failed to return response)
27 input, 128,000 output, thinking_tokens: 128000
$1.28
Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.croemer 19 minutes ago | parent
gumby271 17 minutes ago | parent
TomGarden 16 minutes ago | parent
heyjstn 8 minutes ago | parent
thefourthchime 6 minutes ago | parent
2nd only to Opus 5.5 https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/
aimaxxed 4 minutes ago | parent
SeriousM 25 minutes ago | parent
heyjstn 22 minutes ago | parent
- Fable 5.1 for planning/adversarial reviewer
- Opus 5.5 for well-scoped tasks break down
- Sonnet 5.5 for these well-scoped tasks implementation
I think the blocker might be how efficient the context is compacted and sending around between these agents
chrismustcode 21 minutes ago | parent
Changing model would be cache busting spiking usage for no good reason when Opus can do it all.
Haiku 5.5 might fit well though depending on pricing.
SirMadam 14 minutes ago | parent
enraged_camel 4 minutes ago | parent
afro88 16 minutes ago | parent
AM1010101 21 minutes ago | parent
On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.
If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.
gregwebs 21 minutes ago | parent
From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.
OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.
jtrn 19 minutes ago | parent
If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.
BUT
It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.
Some of the more interesting things I found from scanning the system card:
- It is the only model tested that shows no preference for rude or polite style.
- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.
- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).
- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.
- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.
- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.
- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).
Clinical behaviour:
Suicide and self-harm handling is reported as weaker in the API because it
It sometimes called a wish to die understandable.
It sometimes validated self-harm as functional.
It sometimes suggested harmful substitute behaviours.
As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."
Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be
dude250711 16 minutes ago | parent
dom96 7 minutes ago | parent
https://bench.killswitch-lang.org/
Claude Sonnet 5 17.8%
Claude Sonnet 5.5 7.4%swingboy 4 minutes ago | parent
laurenz-bauer 1 minute ago | parent