47 points Michelangelo11 1 hour ago 37 comments
prodigycorp 41 minutes ago | parent
There's so many "x generated this in one shot, this is agi" stuff that gives you the impression that you can vibe operate modern models the same way you operated last year's models. There's so much more to it than that. It requires you to put a faith in the leap in the capability of models, one that would've surely been a waste of time in previous models.
Not sure where i'm going with this other than I think most can relate that it's exhausting keeping up with. I cant imagine what it'd be like parenting a kid that went from toddler to puberty in the span of a year and planning for them to go to college the next year. This industry is moving so fast that it's becoming fact that it's the user that's "holding it wrong" every six months.
XenophileJKO 30 minutes ago | parent
The step function change on Opus 5.5 for visual work shocked me.. and I haven't been surprised like this in a long time with LLMs.
EDIT: When I first saw the "P(DOOM)" video and some of the other animations I was VERY skeptical that Opus 5.5 without a lot of tools could make something like that.. until I tried it for myself. It can.. 100%.
prodigycorp 27 minutes ago | parent
But it other cases, like the music videos, much of the magic is done by access to elevenlabs and suno apis.
Edit: just saw your edit about the pdoom video. Can you share how you prompted it? Would be helpful to know.
collabs 23 minutes ago | parent
I'm still more worried about the malice and any malicious acts by the people at these frontier labs than the models at the frontier labs.
alansaber 15 minutes ago | parent
prodigycorp 5 minutes ago | parent
mathisfun123 39 minutes ago | parent
ulfw 34 minutes ago | parent
ddosmax556 32 minutes ago | parent
mathisfun123 29 minutes ago | parent
Not sure how you missed it but that's exactly what I'm calling out as asinine.
> We're in a race right now
Again: consider the analogy about cars... which are literally used for racing (occasionally).
UrineSqueegee 21 minutes ago | parent
i dont understand how you're framing this. how is this a bad thing exactly? How is it asinine?
mathisfun123 16 minutes ago | parent
bronlund 31 minutes ago | parent
Edit: Someone commented that this is insulting to autists, and I guess it kind of is - sorry. What I ment was an intellectual; one that can be an absolute retard, but have read an aweful lot.
anotha_one 29 minutes ago | parent
kalleboo 13 minutes ago | parent
Cars were just like that during their early years, with tillers and knobs. See the video where Top Gear finds the first car with controls we recognize https://www.youtube.com/watch?v=fkwGJzU5B-I
Bishonen88 36 minutes ago | parent
> Asked for frontend work without design direction, Claude Opus 5.5 falls back on a few default styles, and a general instruction such as "avoid a generic AI look" mostly swaps one default for another. It responds well to instructions that name specific patterns to avoid, as in the following example. Work iteratively: check which styles the first result used instead, and extend the list if needed.
I hardly ever read tips for prompting etc. because things change too quickly, the writeups are kindof big. Glad I read this one, because I often did exactly what they assume users would do. I write "don't make it look like generic ai slop" and that seemed to work nicely. Now I know why there was still a chance of seeing similar styles across apps. I reckon doing some manual work in terms of scouting dribbble/behance for nice layouts will yield better results.
emsixteen 17 minutes ago | parent
skeledrew 13 minutes ago | parent
bronlund 35 minutes ago | parent
prodigycorp 32 minutes ago | parent
dhanushnehru 35 minutes ago | parent
Just saying "continue" when it gets stuck usually makes it repeat the same error. A better way is to save its last action and result, then make it try a new approach. If it tries the exact same thing twice, it should stop and ask the user for help instead of wasting money on a loop
skeledrew 25 minutes ago | parent
shaan7 21 minutes ago | parent
OtomotO 19 minutes ago | parent
Which was and is true to some extent.
And don't get me wrong, China is a dictatorship, and a tyranny for some.
But then again, the west is a tyranny for some.
muzani 15 minutes ago | parent
bob1029 22 minutes ago | parent
I'm not sure I understand this complexity. In all harnesses I've ever used, tool calls themselves are surfaced to the user as an indication of progress. When the UI/UX around this is engineered well, the user should be able to infer roughly what is going on. Different tools have different ideal presentations. You can't reduce everything to plaintext blobs.
If I absolutely needed intra-turn progress updates, I'd accumulate a separate per-turn transcript and feed it into a cheaper model at deterministic intervals.
pookieinc 20 minutes ago | parent
I say the above because I'm seeing entire worlds and games being one-shotted built on X and I just have no idea how they do it. I tried building a large prompt for Fable when it was first released and it didn't have anything close to resembling some of the stuff I'm seeing today.
derencius 15 minutes ago | parent
user43928 19 minutes ago | parent
Claude Code has an output style setting that I set to "Concise", with no apparent effect.
I am told this is merely something in the system prompt that the model tends not to pay attention to with large contexts.
Opus 5.5 writes whole essays at the end of the turn, with the important actionable steps somewhere at the bottom.
When prompted to give a concise summary, it usually overshoots into a super short summary and then you have to dig into the details again anyway.
In general I find Opus 5.5's writing to still have more "ticks" or "Claudisms" than the OpenAI models.
Its explanations often appear overcomplicated for simple concepts.
Sure, it's leagues above the ridiculous writing of Opus 5, but Anthropic still has a long way to go here.
derencius 17 minutes ago | parent
user43928 14 minutes ago | parent
I guess I could take some lengthy example explanation, and have it try various instructions and test what results in output that I find preferable.
Maybe I'll give that a try, thanks!
wren6991 7 minutes ago | parent
IIRC it's a system reminder injected after every single turn.
It must be pretty ingrained to be so resilient against prompting. I think heavy RL on relatively short-horizon programming tasks has given the model a tendency to write down absolutely everything, to beat the compaction. Longer-term tasks where this crap piles up are in the evolutionary shadow, so to speak.
kimseungyong 11 minutes ago | parent
preommr 8 minutes ago | parent
In which case we've royally fucked ourselves that the level of engineering we've reached is... prompts. Because there is a deadline where we have to show productivity to justify all the investment spending.
People need to build with tools in a reliable, constructive way. Not vodoo magic based off vibes. We need better structured output, better transparency on what these models can do, better controls overla, maybe new ideas on loops graphs, and ways to use the models. Like, at least people were trying new things with jev.
KellyCriterion 6 minutes ago | parent
I used Opus 5.5 for some simpler tests and was quite angry when I saw that each of my question was above 10USd
simonw 5 minutes ago | parent
Summarize the main complaints in this thread.
<pasted_content id="ab12">
...text the user pasted...
</pasted_content id="ab12">
Where those IDs are randomly generated and unknown to the user, and the model is told to use that markup to help avoid it suffering prompt injection attacks.In the past I've been very skeptical of this kind of protection. Anthropic have clearly trained their models for this though, so maybe Opus 5.5 is smart enough for this to work?
Will be interesting to see if minds more devious than mine can break it.