88 points talhof8 3 hours ago 23 comments
fwip 1 hour ago | parent
tyingq 59 minutes ago | parent
> . The accepted runs cost $4.65. Failed attempts and replacement runs increased the complete cost to $5.14.
fwip 51 minutes ago | parent
Aldipower 58 minutes ago | parent
TuxSH 55 minutes ago | parent
I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.
Perhaps DS works better where targets have low-hanging fruits than can be found fast?
simlevesque 49 minutes ago | parent
d5lt5 30 minutes ago | parent
TuxSH 9 minutes ago | parent
But, well, number of subagents doesn't make a difference if model is dumb (GPT 5.4 High, in May,in Chat mode outperformed what I see with DS4.1-F).
That being said, pricing model makes a huge difference for "find at least one" tasks: with API/PAYG if you have a chance to save 90%, you go for it, whereas with subscriptions it is optimal to burn all your remaining allowance right before reset
mariopt 34 minutes ago | parent
Given how fast and cheap DS is, it's just an ideal model with enough "IQ" to let it loose. Another thing they left out of the article, DS becomes really good with if provide custom tools for the task, on it's own it's mediocre.
seemaze 12 minutes ago | parent
allie1 33 minutes ago | parent
surgical_fire 26 minutes ago | parent
I love DS flash, it is an amazing workhorse to implement plans created by more robust models (such as GLM). But a more fair comparison would be of DS Flash with GLM Flash.
cmrdporcupine 24 minutes ago | parent
I have good results with DS4.1 flash because I can iterate faster. I either provide it with correction, or it discovers its failures via the harness. And seems to respond well to empirical evidence rather than go in circles.
So it might need some prodding, but it's likely in this case it was able to brute force after several runs and collecting some evidence.
severino 13 minutes ago | parent
nickysielicki 43 minutes ago | parent
EGreg 37 minutes ago | parent
“If we ban CFCs now the Chinese will win!”
“If we ban chemical weapons, nuclear weapons, etc etc our enemies will triumph! They won’t stop!”
“If we switch to biodegradeable plastic then our rivals will have an advantage.”
“If we dont externalize the costs to our population, then they will, and then will win!”
I think workflows can do the job agents do, 20x cheaper and more predictably and safely. They can completely displace agents, just as HFCs displaced CFCs and then we were able to ban CFCs and phase them out through international COOPERATION. The language of COOPERATION is what saves us vs COMPETITION is all about cutting corners and externalizing costs. Google the Montreal Protocol, Geneva Conventions, Nuclear Non Proliferation Treaty, Unleaded Gasoline etc etc.
Agents have got to be marginalized. They are just popular because the labs need to make a ton of money for their investors and recoup their massive spending on training models.
airstrike 36 minutes ago | parent
You can't compare banning football to banning genetic experiments and say "they are both bans and therefore directly comparable"
palmotea 8 minutes ago | parent
Come on. Lecturing about hubris when your message is damn the consequences, full speed ahead?
If it's a race to build the torment nexus, or a race with a nonzero chance of building the torment nexus by accident, I don't care about winning.
jrflo 39 minutes ago | parent
wg0 22 minutes ago | parent
The 2 trillion dollar ROI on anthropic alone?
Good luck with that.