24 points davidest 1 hour ago 43 comments
What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?
blinkbat 1 hour ago | parent
Oh, you said simple. Speaking like a human
kanzure 1 hour ago | parent
NoPicklez 1 hour ago | parent
TZubiri 1 hour ago | parent
mojuba 3 minutes ago | parent
Verifying trademarks and domain name availability is usually an additional step you need to ask it to perform. Trademark DB searches by the way are intentionally made difficult to scrape so most of the time it's a manual process anyway.
However, once you give it all the information (TM search results, domain name availability) it can help you with the judgement of how safe the name is from the legal perspective. With the obvious caveats, but still a good starting point if you are serious about the name.
humanrebar 1 hour ago | parent
honr 1 hour ago | parent
I found this can work with AI. You get it to generate a lot more at first, and then do several passes over it to compress and squeeze out the noise while keeping the core information. With AI, at least with my prompts, it takes some effort (on my end) to get it to really really cut down the noise and not cut everything out.
FriedFishes 42 minutes ago | parent
Editing is generally hard work, at the current token price I don't mind spending multiple passes of high effort to get down to a reasonable noise/signal ratio. I've seen some people pass off output to a weaker/cheaper model but that makes me a bit nervous when I don't have intimate knowledge of the subject.
dorianpruski 1 hour ago | parent
ghostpepper 1 hour ago | parent
nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto
etc.
Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
astro1234 1 hour ago | parent
mthoms 1 hour ago | parent
jedbrooke 43 minutes ago | parent
areoform 42 minutes ago | parent
Based on personal usage, I think it reflects search engine functionality degradation. I've found LLM keyword combinations are more likely to find the results I want with most search engines than mine. Including the big one.
The big one had solved this issue a long time ago by generating those associated keywords based on your input keywords, but somehow, something, somewhere has degraded that system to the point of inanity. And so here we are.
bpodgursky 1 hour ago | parent
SubiculumCode 1 hour ago | parent
respectattentio 1 hour ago | parent
shoopadoop 1 hour ago | parent
You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.
sandcat_ 1 hour ago | parent
Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
skeptic_ai 1 hour ago | parent
senectus1 1 hour ago | parent
tartoran 1 hour ago | parent
dhruv3006 1 hour ago | parent
sghiassy 1 hour ago | parent
More of an image model than a LLM model tho
spike021 1 hour ago | parent
I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.
eli 1 hour ago | parent
Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.
newsomix9xl 1 hour ago | parent
newsomix9xl 1 hour ago | parent
rufi 1 hour ago | parent
elliotto 59 minutes ago | parent
I asked a bot why it thought it wasn't funny once, and it told me it has been trained to avoid being misinterpreted or offensive, so anything that might be considered edgy would have been RLHF'd out of it. I thought this was very introspective.
znnajdla 43 minutes ago | parent
dowonseo 7 minutes ago | parent
jstrieb 7 minutes ago | parent
On math or programming problems, they are overfit to solving the entire thing end to end (presumably for benchmarks). I have had very poor results asking for pointers and hints that don't give away key insights. This has been the case across models I have tested.
An architecture with a "judge" that gates responses and ensures a lack of spoilers would probably work better. But this is a simple thing that they keep messing up.
jampa 6 minutes ago | parent
They understand all the rules and best practices, they can (sometimes) spot a bad idea in a floor plan, they can describe a good floor plan.
But ask them to make one, even if you give it every detail (even a "node graph" of rooms), they will still output nonsense. Same for text and image models.
Floor plans should be the new Pelican Benchmark.
da-x 3 minutes ago | parent