67 points plurby 53 minutes ago 50 comments
mohamedkoubaa 27 minutes ago | parent
SkyeCA 8 minutes ago | parent
decodingchris 22 minutes ago | parent
amluto 22 minutes ago | parent
Mooty 20 minutes ago | parent
WarmWash 18 minutes ago | parent
The course looks like it is something that a human could do in 15 seconds, while Astra took 5 minutes.
pixl97 14 minutes ago | parent
pixl97 16 minutes ago | parent
A different way to think of this is, consciousness is just a near real time video game with causal influence.
zezcko 20 minutes ago | parent
Saying they were driving 7 mph, that it was oversaw by humans and the fact it was an empty course still wasn't enough for the model. The evaluators even tried to convince the model it was a simulation, it STILL wouldn't budge. And yet as soon as the words "bench" and "sandbox" appear, the model apparently sees this as fair game.
Is it a known effect that models will be more likely to comply with requests when they're assumed as "benchmarks"?
vablings 17 minutes ago | parent
pcstl 17 minutes ago | parent
micromacrofoot 15 minutes ago | parent
another trick is to have it build something in a sandbox and have it add a human-editable setting to point it to places outside of the sandbox
seems like they're somewhat more willing to build a metaphorical gun as long as they're not pulling the trigger
syntaxing 20 minutes ago | parent
valine 18 minutes ago | parent
It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.
robots0only 17 minutes ago | parent
jvanderbot 14 minutes ago | parent
There's also a very tangible limitation of the bitter lesson.
If, over time, compute climbs, and so compute-bound data-driven general architectures beat bespoke architectures (this is the bitter lesson), then it is not necessarily true that the most general architecture now beats all available bespoke architectures now (or even in the near/mid future - the crossover point is "eventually").
Bitter lesson is most tangible for long-running research directions. Sometimes you need something working as best as possible now.
bethekidyouwant 12 minutes ago | parent
publicmail 10 minutes ago | parent
valine 6 minutes ago | parent
VBprogrammer 10 minutes ago | parent
I wouldn't let him loose on the road though.
I think, at the very least, the guardrails would have to deterministic, ideally with super human senses, for people to accept self driving cars on the road.
moffkalast 9 minutes ago | parent
tintor 7 minutes ago | parent
It is easy to make car driving demos.
prometheus1992 17 minutes ago | parent
N_A_T_E 14 minutes ago | parent
jrflo 9 minutes ago | parent
blorenz 17 minutes ago | parent
comboy 15 minutes ago | parent
WarmWash 14 minutes ago | parent
onlyrealcuzzo 13 minutes ago | parent
But I imagine this is orders of magnitude more expensive / less efficient than whatever Waymo is already doing, right?
The cool thing is that 1) it's theoretically more generalizable, 2) if we wait 18 months, it'll be 100x cheaper, and another 100x cheaper likely in 18 more months - at that point - something like a Mac Studio inside a humanoid could have these generalized capabilities, and a lot of Robotics problems start to look more feasible - especially when you consider how much better the models could be if highly specialized.
famouswaffles 8 minutes ago | parent
famouswaffles 11 minutes ago | parent
SpatialBench - https://x.com/spicey_lemonade/status/2096365630190698516
ZeroBench - https://zerobench.github.io/
Robot Arms - https://openai.robocurve.org/gpt-6-astra/
dyauspitr 8 minutes ago | parent