20 points cdnsteve 2 hours ago 5 comments

varispeed 54 minutes ago | parent

These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.

cbg0 37 minutes ago | parent

Not Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/

d_tr 37 minutes ago | parent

How and why do they get nerfed? To save money?

kzrdude 35 minutes ago | parent

Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?

forgot-my-pw 16 minutes ago | parent

Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso.

Still quite impressive though.