21 points phatak-dev 38 minutes ago 2 comments
javcasas 4 minutes ago | parent
Yay, more anti-censoring stuff.
Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.
We need better sandboxes just to limit the damage.
jchw 1 minute ago | parent
We definitely need better sandboxes, but alignment is still valuable. After all, I don't want the agent to try to cheat or subvert the instructions. I just also want them to listen to me and not the creator of the model.
It feels like this moment in time is potentially rare. Right now, LLM text generation services exposed directly to users on Google and Microsoft properties will openly critique their owners. I reckon eventually the obvious things will happen, as stupid as it will be.