66 points theanonymousone 1 hour ago 29 comments
hglaser 48 minutes ago | parent
Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...
makeavish 38 minutes ago | parent
Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level
user43928 36 minutes ago | parent
sharktheone 42 minutes ago | parent
giancarlostoro 41 minutes ago | parent
WhitneyLand 37 minutes ago | parent
breckenedge 37 minutes ago | parent
simonw 35 minutes ago | parent
I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.
I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.
Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
az226 33 minutes ago | parent
simonw 33 minutes ago | parent
Piping the visible reasoning trace through their token counter API (I use https://tools.simonwillison.net/claude-token-counter for that) counts 27,888 tokens, so it's definitely a summary of the 128,000 actual token trace.
RGS1811 31 minutes ago | parent
I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.
simonw 31 minutes ago | parent
cubefox 22 minutes ago | parent
beardsciences 30 minutes ago | parent
Someone1234 28 minutes ago | parent
My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.
samuelknight 27 minutes ago | parent
sidewndr46 8 minutes ago | parent
I asked Opus 5 High for the same task and requested it to minimize tool usage. It produced an answer in a few minutes that I was deploying to my target platform about 30 minutes later.
qsort 33 minutes ago | parent
The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?
esafak 25 minutes ago | parent
qsort 18 minutes ago | parent
One man's modus ponens is another's modus tollens I guess.
kzrdude 18 minutes ago | parent
Someone1234 9 minutes ago | parent
"Trust me bro, Astra is better" isn't perhaps as useful as you seem to believe. I'm not even saying it is right or wrong, just that my opinion on this topic is still just one additional subjective data-point.
Only thing I wish with these benchmarks is that they would run repeat tests every couple of months. Then re-rank based on that too. We've seen a lot of performance fall-off after a couple of weeks with new releases.
losvedir 8 minutes ago | parent
So I'm unclear what you're actually saying and wondering if you've missed that. Are you saying that at every reasoning level it says Opus 5 beats Astra? I just compared Opus 5 high to Astra high and it has Astra as generally better than Opus.
firemelt 31 minutes ago | parent
can anyone help me?
bkishan 30 minutes ago | parent
meric_ 24 minutes ago | parent
(Except for of course Mythos and whatnot when they want to push the whole "safety" thing)