42 points tosh 3 hours ago 24 comments

m-hodges 1 hour ago | parent

> Starting today, API customers globally will be able to opt in to text watermarking for select models. Text watermarking will remain off by default in the API.

> Over the coming weeks, we will add an invisible watermark to eligible ChatGPT and Codex text output in the European Union.

mgax 1 hour ago | parent

This is such a waste of time. If someone wants to bypass this it will be rather simple. Just change the words. If someone wants to avoid fingerprinting they will. Can’t we just focus on building rather than spending brainpower on these ridiculous sidequests

Analemma_ 1 hour ago | parent

I think some EU regulations have merit and some don't, but going "can't we just focus on building instead of spending brainpower on following the law" is not going to win you much affection. There might be a connection between this pervasive attitude in tech and why now every proposed datacenter construction project is being met with ferocious public opposition.

iamnothere 1 hour ago | parent

This isn’t just a “techbro” attitude. Although they disagree on the specific problem laws in question, basically every political group (at least in the US) takes issue with the law as written. And most agree (72% as polled in 2023) that the law is written to favor the interests of the wealthy. The pushback against data centers isn’t in contrast to this—data centers provide the bulk of their benefits to wealthy investors, and many localities are permitting them against the will of local citizens.

You might also want to consider that many believe the current AI regulation push is a trojan horse for regulatory capture and subsequent domination of the industry by large, well-connected players who are hostile to privacy and individual freedom.

So when you see US people on a US site complaining about a law that affects US industry, especially scrappy startups, consider if maybe there’s more to it.

SpicyLemonZest 1 hour ago | parent

There's lots of regulations that are easy to bypass. It's absolutely trivial to, say, pull 30 amps from a home circuit rated for only 15 amps. But most people will just trip the fuse because they don't know what they're doing, and the existence of the regulation makes it easy to prove ill intent for anyone who does know what they're doing but causes harm anyway.

akersten 28 minutes ago | parent

Of course, the implication by comparison to the building code example, that text absent some homeopathic suggestion of provenance causes harm, is absurd at best.

Sent from my iPhone (harm prevention watermark)

bsoqk 24 minutes ago | parent

The way I interpret your comment is that, in the future, it will be enough to prove that someone tried to remove an LLM watermark to put them in jail.

someonebaggy 8 minutes ago | parent

If someone causes financial loss by posting LLM text and is then found to have removed a watermark.

IsTom 6 minutes ago | parent

Imagine that somebody sold you text that they mislead you to think is not generated by AI and then trying to convince court that they didn't mislead you intentionally. If they've removed the watermark it stops this kind of defense.

gonzalohm 40 minutes ago | parent

Focus on building what? A future in which only big companies get credit for their work?

bsoqk 27 minutes ago | parent

Isn't that what this watermarking allows for? To allow LLM companies to detect that certain text was authored by their tools?

smokel 1 hour ago | parent

Why use watermarking, and not simply add a signature?

SpicyLemonZest 59 minutes ago | parent

The guidelines (https://digital-strategy.ec.europa.eu/en/library/guidelines-...) require a solution that's robust against "common alterations and adversarial attacks".

aenis 1 hour ago | parent

Another cookie consent-grade success of the EU.

zetanor 6 minutes ago | parent

The near-ubiquitous malicious compliance that took place during the implementation of GDPR (i.e., displaying consent modals instead of just not sending visitors' personal information to 93 third party services) at least had the benefit of bringing to light just how much web service providers view their users as things to be bought and sold. I doubt that text watermarking will result in much of anything at all, besides additional expense.

iririririr 2 minutes ago | parent

This is also malicious compliance.

The whole effort started as a way to pin point which work openAI stole from! Now they are redirecting the dicussion to "our sources are magic. let's talk about who is copying from that magic".

All the controversy about this being pointless, is on purpose. They are pushing the pointless thing as a distraction and we keep discussing how pointless it is. no wonder.

greatgib 1 hour ago | parent

My personal opinion is that they cheated evaluations to be able to release this pretending that it has no meaningful impact.

Otherwise, I don't see any logic explanation that some of there benchmarks results would be higher when watermarked. Except if benchmark results are so unstable that they are a useless metric.

Y_Y 40 minutes ago | parent

I agree that if they're so noisy they need to average over more runs, but maybe they didn't feel the need to bother.

yorwba 2 minutes ago | parent

As long as variance is nonzero, you won't get the exact same result twice for the same benchmark, so one of the numbers has to be higher and the other lower. If watermarking has no effect on the distribution, that's a 50% chance the watermarked model gets the higher number. In this case, it happened 5 out of 8 times, which is hardly unusual.

athrowaway3z 42 minutes ago | parent

>> Editing can weaken the watermark. In an evaluation of 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%. Replacing 25% of words reduced it to 17%.

> Claude/codex/deepseek, please replace 25% of words with synonyms or slight rephrasing because i dont like the current version.

Not sure if that counts as: `a solution that's robust against "common alterations and adversarial attacks"`. Is there a sort of adversarial attack that is more common?

yorwba 30 minutes ago | parent

If you use a watermarking model to do the synonym replacement, it will rephrase it in a way that is compatible with the watermark...

k__ 20 minutes ago | parent

Is watermarking part of a model architecture or is it something added by the inference engine?

someonebaggy 7 minutes ago | parent

Added during sampling by adding a bias to token probabilities in places where several tokens are equally likely.

richwater 16 minutes ago | parent

Just make the models worse for the EU. Don't accept this nonsense that's holdingg back actual work and progress.