20 points cdnsteve 2 hours ago 5 comments
varispeed 54 minutes ago | parent
These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.
cbg0 37 minutes ago | parent
Not Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/
d_tr 37 minutes ago | parent
How and why do they get nerfed? To save money?
kzrdude 35 minutes ago | parent
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
forgot-my-pw 16 minutes ago | parent
Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso.
Still quite impressive though.