407 points sfkgtbor 2 hours ago 197 comments
TheAmazingRace 2 hours ago | parent
himata4113 2 hours ago | parent
serious bit: if you think about how these smaller models work, at the end of the day it seems that they are now capable of forgetting useless information because they're able to derive it in reasoning allowing models to become smaller at the cost of requiring more reasoning tokens to solve a task.
qeternity 2 hours ago | parent
dyauspitr 2 hours ago | parent
bravetraveler 2 hours ago | parent
ChaseRensberger 2 hours ago | parent
onlyrealcuzzo 2 hours ago | parent
You'll know when the trend stops -> when the intelligence differential between smaller models like 7B starts to grow instead of shrink from 32B models -> that means 7B is getting about as smart as it can get. Then, 32B will follow next, then 70B, etc etc.
We haven't yet seen that at any size AFAIK.
thefourthchime 2 hours ago | parent
onlyrealcuzzo 1 hour ago | parent
I wouldn't be surprised if less than 1B param equivalent of our brain deals with solving math and writing computer programs and physics and all the things we tend to associate with "intelligence" - especially if you ultra optimized for that, I doubt our brain works like that.
Dealing with the real world, I highly highly doubt it.
jstummbillig 1 hour ago | parent
Given that humans learn to talk while having encountered a measly number of word instances, and, given enough time, we should always be able to improve on the lottery that is biology, it does seems fairly likely.
istjohn 2 hours ago | parent
> The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. [0]
0. https://epoch.ai/publications/the-plunging-price-of-thought
FooBarWidget 1 hour ago | parent
jstummbillig 1 hour ago | parent
teaearlgraycold 1 hour ago | parent
adgjlsfhk1 1 hour ago | parent
jrflo 1 hour ago | parent
srdjanr 37 minutes ago | parent
minimaxir 2 hours ago | parent
Input
$0.10 / MTok for prompts up to 100,000 tokens
$0.50 / MTok for prompts over 100,000 tokens
Output
$0.50 / MTok for prompts up to 100,000 tokens
$2.50 / MTok for prompts over 100,000 tokens
100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])
giancarlostoro 2 hours ago | parent
I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.
tr4656 2 hours ago | parent
From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.
minimaxir 2 hours ago | parent
Fixed.
tripleee 1 hour ago | parent
usef- 10 minutes ago | parent
The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though.
tripleee 5 minutes ago | parent
j45 2 hours ago | parent
Eridrus 2 hours ago | parent
Neither encode nor decode are linear in compute, so providers need to price for average expected length.
This is just getting closer to the true cost of generating tokens.
insanitybit 2 hours ago | parent
dannyw 2 hours ago | parent
For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.
These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.
In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.
WinstonSmith84 1 hour ago | parent
system2 2 hours ago | parent
user43928 2 hours ago | parent
wyrdcurt 1 hour ago | parent
mrngld 1 hour ago | parent
aesthesia 1 hour ago | parent
skeledrew 1 hour ago | parent
pkulak 6 minutes ago | parent
usef- 6 minutes ago | parent
esafak 2 hours ago | parent
AustinDev 2 hours ago | parent
There are plenty of workflows like translations where you'd easily be under the cap.
enraged_camel 2 hours ago | parent
Your vibes don't appear to be supported by facts. From the announcement:
>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.
Philpax 2 hours ago | parent
StilesCrisis 1 hour ago | parent
enraged_camel 39 minutes ago | parent
Tiberium 2 hours ago | parent
You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc
AtNightWeCode 1 hour ago | parent
So, it is might be even worse.
Tiberium 1 hour ago | parent
HarHarVeryFunny 2 hours ago | parent
For this application 100K token input is plenty.
Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.
martianvoid 2 hours ago | parent
HarHarVeryFunny 1 hour ago | parent
For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.
I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.
port3000 2 hours ago | parent
mnicky 1 hour ago | parent
alexchamberlain 1 hour ago | parent
sixtyj 10 minutes ago | parent
maz1b 2 hours ago | parent
GDPval-AA v2.1 as of now: 1620
GDPval-AA v2.1 for Haiku 4.5: 735
The 100k tokens pricing makes sense, looks to be a hedge against OpenAI's decisions API and Jev or its open source alternatives that are springing up.
Nice release, congrats to Anthropic.
sroussey 2 hours ago | parent
iagocc 2 hours ago | parent
TomGarden 2 hours ago | parent
margorczynski 2 hours ago | parent
onlyrealcuzzo 2 hours ago | parent
This is more expensive, but it also looks like it's better enough that it's far more useful.
I also won't be surprised if you look at cost per completed task + wall clock time that it comes out ahead for the majority of what you'd want to actually use it for.
Luna will still be a great option for doing non-engineering tasks super cheaply.
TomGarden 2 hours ago | parent
For prompts over 100k tokens it's 5 times more expensive - $0.50 in, $2.50 out.
simianwords 2 hours ago | parent
She said she was using Haiku 4.5 because she was advised to be careful with the spending.
I hate that model so much lol.
seaal 2 hours ago | parent
Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.
0gs 2 hours ago | parent
thepasch 2 hours ago | parent
https://support.claude.com/en/articles/15036540-use-the-clau...
InsideOutSanta 1 hour ago | parent
skeledrew 1 hour ago | parent
Wait what? This has gotten their blessing?
patrickwdaly 2 hours ago | parent
svachalek 2 hours ago | parent
steve_adams_86 2 hours ago | parent
mariocesar 2 hours ago | parent
I also have an "ask" script that I use daily to ask simple stuff, it can access websearch and webfetch, it's more than enough to parse logs, ask for commands, quick research on the internet, small stuff. https://github.com/mariocesar/dotfiles/blob/main/common/.loc...
I use haiku for things that needs to be quick, have really clear instructions.
gghootch 2 hours ago | parent
Planning on doing flash analyses of PRs that impact evals in some way, and then post comments on GitHub whenever there’s flaws in them
swalsh 2 hours ago | parent
dannyw 2 hours ago | parent
Of course it’s not as big, and hence falls-off quicker. I’d consider the 100k a “promotional price” to match Luna’s token pricing while delivering noticeably more intelligence.
Plutoberth 2 hours ago | parent
I've been using Luna, but I'll probably switch to Haiku.
notatoad 1 hour ago | parent
used <10% of my 5hr limit on a $100 codex plan.
hector_vasquez 1 hour ago | parent
tpoacher 2 hours ago | parent
afrnswrth 2 hours ago | parent
caaqil 2 hours ago | parent
If you block pentest or "other techniques more likely to be used by attackers", then what does "permit a wider range of defensive tasks" even mean?
Any defensive task that's meaningful is almost indistinguishable from legitimate red-teaming that then falls under 'likely to be used by attackers". If only they would just stop nerfing these models, that'd be great. No APT is waiting around for Anthropic's permission, so might as well let us have some cool stuff.
TuxSH 1 hour ago | parent
(sadly Mistral Large 4 isn't up to par - but Mistral serves GLM at 130 tps!)
simianwords 2 hours ago | parent
Did anyone read this? We get free API credits on some plans now
AtNightWeCode 2 hours ago | parent
AtNightWeCode 53 minutes ago | parent
charlesabarnes 2 hours ago | parent
This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes
geek_at 2 hours ago | parent
enraged_camel 1 hour ago | parent
charcircuit 1 hour ago | parent
enraged_camel 1 hour ago | parent
That's 5-15 minutes of work at most. Not exactly the type of lock-in the parent is implying.
charcircuit 1 hour ago | parent
tomjen3 1 hour ago | parent
losvedir 1 hour ago | parent
This is more to encourage people to try out adding AI into their product, which is a totally different flow and experience from using AI to build the product.
tech234a 2 hours ago | parent
alasano 1 hour ago | parent
Anthropic isn't even close to being this useful.
Iolaum 1 hour ago | parent
Biggest loss is that Ant models look like they are genuinely better.
matsz 1 hour ago | parent
This changes on a weekly basis, I ended up with subscriptions to most of the providers (except for X.ai).
thepasch 1 hour ago | parent
https://support.claude.com/en/articles/15036540-use-the-clau...
sambaumann 1 hour ago | parent
sanex 1 hour ago | parent
martinald 1 hour ago | parent
thepasch 1 hour ago | parent
https://code.claude.com/docs/en/headless
So, as written, yes.
ziga 32 minutes ago | parent
tekacs 58 minutes ago | parent
Being forced through the non-OSS Claude Code with all of its quirks and issues is... such an exhausting use of force by Anthropic.
To the extent that you _can_ choose to disable telemetry and training on your traces in CC, it's not all that obvious what they gain by crippling your ability to use the subscription with other – better – tools.
It's also remarkable that it's coincident with OpenAI adding "Sign in with OpenAI", so that you can use your tokens with other tools.
eli 28 minutes ago | parent
You might be right and they will change this in the future, but that's speculative
thepasch 27 minutes ago | parent
This text has replaced the entirety of the page called "Use the Claude Agent SDK with your Claude plan."
eli 17 minutes ago | parent
thepasch 14 minutes ago | parent
What more do you need?
eli 7 minutes ago | parent
Topfi 1 hour ago | parent
Of course, they don't do this out of pure kindness, but I really struggle to see a negative for subscribers already using a Claude Max subsection, especially given changing to another model is essentially frictionless.
Compared with "Sign in via OpenAI" which they just announced, this is far less lock-in for anyone hosting services but less interesting for users of said services. With Anthropics approach, you can just use the allowance on your users however you see fit along with any other models and once it's used up, you can still just decide not to use their models for the remainder. With users bringing their tokens meanwhile, there is less flexibility in terms of switching for you, though might be cheaper for users.
Both interesting, each approaching this from a very different direction, each having their own trade-offs. On the OpenAI front, will be interesting whether developers can set specific temp, reasoning budgets, etc. for such "provided tokens" or whether OpenAI exposes that only via the actual API.
crooked-v 2 hours ago | parent
garo-pro 2 hours ago | parent
Opus 5.5 runs 117 tps average on Openrouter, so it must be at least 10-20 tps slower for them to mention. IDK why they mention this as it does not help for marketing though. https://openrouter.ai/anthropic/claude-opus-5.5
wyrdcurt 2 hours ago | parent
jjcm 2 hours ago | parent
Haiku 5.5: https://html.non.io/lcars-haiku-5.5/
Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5
Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...
One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.
BrokenCogs 1 hour ago | parent
FranzFerdiNaN 1 hour ago | parent
twostorytower 1 hour ago | parent
BrokenCogs 1 hour ago | parent
jjcm 1 hour ago | parent
It's meant to be a good test, not a good design.
thefourthchime 1 hour ago | parent
Considering the price, no model comes close to being as good as this. However, it did take an extremely long time.
TIME 19m COST $0.16 https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5
All results: https://jonclegg.github.io/pacman-bakeoff/
myzie 1 hour ago | parent
rpcope1 50 minutes ago | parent
saretup 1 hour ago | parent
jjcm 1 hour ago | parent
sparklingmango 56 minutes ago | parent
Hasn't this always been the case with Haiku?
bouk 1 hour ago | parent
gizmodo59 1 hour ago | parent
bouk 1 hour ago | parent
- https://blog.chrislewis.au/using-coding-agents-to-decompile-nintendo-64-games/
- https://blog.chrislewis.au/the-long-tail-of-llm-assisted-decompilation/
And to setup a harness that will decompile the game and start doing a matching decompilation of every function. It set up a bunch of tooling and started a service in the background to do this actual decompilation campaign. I put some instructions into the main opus chat now and then to e.g. add automatic git pushing including a nice svg chart of progress and to switch model strategies here and there i.e. to do a first pass with a cheap model and then switch to opus/sol if the small model can't solve it.I could now one-shot a new game, yeah.
supersour 1 hour ago | parent
anthonypasq 1 hour ago | parent
bouk 1 hour ago | parent
MisterMunchkin 35 minutes ago | parent
Like could total war become a browser game?
kro 25 minutes ago | parent
skeledrew 1 hour ago | parent
satvikpendem 1 hour ago | parent
areoform 1 hour ago | parent
> but they still block penetration testing and other techniques more likely to be used by attackers.
>
> Haiku 5.5’s biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm. Organizations working on wider-ranging biology and cyber activities can apply to our Life Sciences Verification Program and Cyber Verification Program.
I would like to take a moment of your time to tell you about some of the "bioweapons" Anthropic has blocked that involved Haiku!These are the examples from "Detecting and countering misuse of AI: September 2026" - https://news.ycombinator.com/item?id=49647300
> Importantly, because our biological safety classifiers robustly block content involving high-risk biological research (in this case, the construction of enhanced pandemic potential pathogens), all of these exchanges occurred on models in our weakest class of models (specifically, the models were Claude Sonnet 4 and Haiku 4.5, the latter of which the user began using after Sonnet 4 was deprecated).
>
> Upon a detailed examination of the exchanges, we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design. This is consistent with our understanding of the capabilities of Sonnet 4 and Haiku 4.5, which are not able to perform expert-level biology research tasks; we estimate that the uplift provided to the researcher was limited and substantially lower than it would have been from one of our more capable models.
Anthropic then says for the above, "we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design"While doing my best to avoid comment, please note, they're talking about a domain expert in a state research institution using Claude to do paperwork.
What did they save us from? What bioweapons did these filters prevent? From the front matter report,
> The above LLM platform is not the only route via which researchers engaged in viral gain-of-function research have used our platform. In May 2026, we discovered a researcher outside the US using Claude in their research on highly-pathogenic avian influenza (“bird flu”). The research focused on viruses’ adaptation to mammals, and the mechanism by which it causes severe disease beyond the respiratory tract.
OK. Sounds serious. "Gain of function research..." but who and why? > The researcher pursued this work in a credible institutional context, and interacted with Claude over the course of several weeks, exchanging thousands of messages. In these exchanges, the researcher leveraged Claude’s knowledge of the scientific literature to assist the researcher in study planning and design, data analysis, and the interpretation and prioritization of experiments. The researcher also used Claude for editorial assistance in writing up the research.
So this was a researcher inside of some country's national lab ("credible institutional context") doing research on dangerous viruses using Claude for "for editorial assistance in writing up the research."What "uplift" are you providing to scientists working at specialized global BSL-4 labs that already have – and I quote their report - "physical access to such isolates." (as in samples of viruses)? Are we uplifting their grammar?
These "safeguards" are being expanded. The scientists I know can't use Claude for grammar checks or anything serious. You can try it for yourself.
Topfi 1 hour ago | parent
Was a big fan of Haiku 4.5, though understand why for most Sonnet was the far better option back then.
johnisom2001 1 hour ago | parent
I ask:
> how many r's in diminished
It answers:
> Diminished has 1 r.
mattz56 1 hour ago | parent
axthauvin 1 hour ago | parent
d1l 1 hour ago | parent
saretup 1 hour ago | parent
3371 1 hour ago | parent
dcchambers 1 hour ago | parent
dangoodmanUT 1 hour ago | parent
This is kind of nuts
djeastm 34 minutes ago | parent
declanjackson 1 hour ago | parent
oh_no 1 hour ago | parent
MisterMunchkin 1 hour ago | parent
They’re definitely planning to make the subscriptions API based so they can charge you full price.
simonw 1 hour ago | parent
> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens
Haiku and Luna now have the exact same price up to 100,000 tokens. Luna is now cheaper for anything after 100,000 tokens, even after Luna's own price increases at 270,000 it's still less than Haiku.
So it sounds like they've directly addressed that problem. Their self-reported benchmarks are all higher than Luna too.
mrbungie 49 minutes ago | parent
Good to know that is going back to being an actual option from perf/price perspective.
justmaris 59 minutes ago | parent
simonw 58 minutes ago | parent
Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.
The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.
The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:
llm install llm-anthropic -U
llm anthropic refresh
llm -m claude-haiku-5.5 'prompt goes here'vinni2 44 minutes ago | parent
abirch 42 minutes ago | parent
1ucky 42 minutes ago | parent
jonshariat 42 minutes ago | parent
fakedang 42 minutes ago | parent
ghoshbishakh 41 minutes ago | parent
simonw 39 minutes ago | parent
Here's a Blender model I had Claude Opus 5.5 create: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
And here's some animated pixel art by Opus 5.5: https://tools.simonwillison.net/kakapo-party
And some Monkey Island style music (Opus can compose music too): https://tools.simonwillison.net/scrimshaw-jukebox
Anthropic's models do all of this by outputting code. GPT-6 Astra has similar capabilities - I got this Blender model using that: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
jansan 28 minutes ago | parent
They are still not great at SVG. I just asked Opus and Fable to add a background to an SVG and the results were, well, not great.
sixtyj 11 minutes ago | parent
vunderba 10 minutes ago | parent
I put this KQ Style Builder together which creates Sierra AGI-style adventure game scenes painted live from *basic drawing instructions* so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style.
sixothree 6 minutes ago | parent
Please don't judge me too harshly for this poop video. But here is an example of something 100% generated with claude prompts only.
ozgung 11 minutes ago | parent
the__alchemist 50 minutes ago | parent
djoldman 32 minutes ago | parent
This section makes the reader think: why would I not pick Sonnet 5.5 instead of Haiku 5.5?
waximabbax 20 minutes ago | parent
hidelooktropic 19 minutes ago | parent