190 points ColinWright 2 hours ago 86 comments
techblueberry 2 hours ago | parent
Robotbeat 1 hour ago | parent
techblueberry 1 hour ago | parent
Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.
HarHarVeryFunny 1 hour ago | parent
It seems that Buckmaster and Levant (who is an Anthropic employee) were rather naive in the amount of trust they had in OpenAI, with Buckmaster going so far as to contact OpenAI's Sébastien Bubeck to discuss what they were working on and clarify that contrary to rumor this was a private collab.
bigstrat2003 1 hour ago | parent
faangguyindia 5 minutes ago | parent
There is no doubt ChatGPT is the most generous LLM provider!
faangguyindia 6 minutes ago | parent
ColinWright 1 hour ago | parent
hn1rig3rak 1 hour ago | parent
spindump8930 1 hour ago | parent
jrflo 1 hour ago | parent
> Improve the model for everyone
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.
It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
nmfisher 1 hour ago | parent
I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).
jrflo 57 minutes ago | parent
But this gets back to the original "who owns the LLM output" and "can you train models on the internet" argument that's been raging for years.
spindump8930 1 hour ago | parent
ProllyInfamous 44 minutes ago | parent
e.g: allows us to sell your personal data to make money so we can continue offering this service to all customers
I'm done with weasle-words and hours-long EULAs – we're at the point where USA needs to catch up to EU's consumer protections, perhaps with laws similar to already-existing USA "truth in lending" requirements (e.g: interest rates must be prominently displayed in a larger font, including annual fees, on all credit offers).
----
My judge-brother always asked during our childhood "why don't you think the judicial system is fair?!?" Thirty years ago, the best I could offer was "because it's a two-tiered system that mostly (only) rich people can afford to participate within."
Now my answer is: "the best example I can give is that our judicial system allows binding arbitration [and qualified immunity for police]. The system is set up so corporate personhood is more important than humanity, and it shows."
ColinWright 1 hour ago | parent
https://news.ycombinator.com/item?id=49643556
Quoting:
> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."
jrflo 1 hour ago | parent
ProllyInfamous 49 minutes ago | parent
At least when my computer's bluetooth exhibits this behavior (e.g: if you don't have a keyboard&mouse plugged in at boot, bluetooth might auto-enable), I can go inside the hardware and physically disconnect the antenna.
What am I supposed to do in software (perhaps hardcode config.file)? in cloud software services (??)?
unified101 48 minutes ago | parent
tomrod 17 minutes ago | parent
fithisux 1 hour ago | parent
tyrabound 19 minutes ago | parent
No mention what those “steps” are, success criteria, or whether they are successful by any navies at all … they take steps though… so it’s fine, and if we know one thing it’s that we can really truly trust someone off the likes of Sam Altman.
rfgplk 1 hour ago | parent
jeremyjh 1 hour ago | parent
cyanydeez 1 hour ago | parent
voakbasda 42 minutes ago | parent
krupan 45 minutes ago | parent
spindump8930 1 hour ago | parent
> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper
> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this
> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"
source: https://x.com/johnschulman2/status/2097440545853637108
rfgplk 1 hour ago | parent
Ydarbleoj 1 hour ago | parent
I’ve lived long enough to know what they say and what they do are often quite different; and it is not our job to trust but to verify.
mrbluecoat 1 hour ago | parent
Welcome to the party, with the rest of humanity.
shevy-java 1 hour ago | parent
AI is becoming more evil by the day.
causal 34 minutes ago | parent
gentlerain 1 hour ago | parent
How do people become that trusting?
The phrasing itself is guilt tripping
quentindanjou 1 hour ago | parent
Or at least: tell me I should be careful/worry about those particular things.
cyanydeez 1 hour ago | parent
speak_plainly 1 hour ago | parent
DrewADesign 1 hour ago | parent
1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest.
And also:
2) someone in the c-suite or marketing has decided to mitigate that through some dark pattern.
If they never intended to give the user that control, they’d probably just not give the user the control in the first place. Giving them the option and not honoring it would either imply they were that incompetent or careless with fate protection, which seems most likely if this is all true, or an even more cynical approach to tricking users into thinking it’s not used for training to get them to share better shit. But if that was the case, why bother with the sleazy dark pattern? That seems a little cartoon villain-y to me. I suppose the in-between option would be that they decided at some point they were no longer willing to honor it and didn’t want to deal with the inevitable PR shitstorm of removing it. I could definitely see that happening in this industry, these days.
ianjbutler 59 minutes ago | parent
Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with lies and ends with lies.
dylan604 37 minutes ago | parent
At this point, I'm left wondering what is wrong with me that I don't just go with the flow, otherwise, what's wrong with everyone that does.
phoghed 17 minutes ago | parent
enraged_camel 46 minutes ago | parent
ozgung 18 minutes ago | parent
So that "opt-out" thing is more like a pause button rather than a permanent thing.
There is also the case of security classifiers always monitoring your conversations. If they flag something they use your conversations to "improve their internal models" even when you "opt-out".
beering 33 minutes ago | parent
mettamage 17 minutes ago | parent
It's fun! I sometimes have tokens to burn and it's instructive despite knowing nothing about the problem other than a NumberPhile video
ayewo 6 minutes ago | parent
From their docs: https://help.openai.com/en/articles/5722486-how-your-data-is... :
> You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training for your ChatGPT conversations and Codex tasks, follow the instructions in our Data Controls FAQ. Once you opt out, new conversations will not be used to train our models.
> For a linked teen account, a parent or guardian may manage whether conversations can be used to improve our models through Parental controls.
> Even if you have opted out of training, you can still choose to provide feedback to us about your interactions with our products (for instance, by selecting thumbs up or thumbs down on a model response). If you choose to provide feedback, the entire conversation associated with that feedback may be used to train our models.
jarofgreen 1 hour ago | parent
Original posts
https://mathstodon.xyz/@andreasthom/117240535270608201
https://mathstodon.xyz/@andreasthom/117240536885387540
https://mathstodon.xyz/@andreasthom/117240537520615623
This story is someone on bluesky screenshotting someone on twitter without a link. The person on twitter screenshotted the source on the fediverse without a link. This is getting daft.
pluc 1 hour ago | parent
qg127 1 hour ago | parent
Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.
Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.
alansaber 1 hour ago | parent
utopiah 1 hour ago | parent
The most successful companies of the last decade have precisely been ... selling usage data.
Makes me wonder if, in 2026, the same people drive a car without realizing that yes it does actually pollute the very air you and your kids are breathing.
calvbak 31 minutes ago | parent
postalcoder 1 hour ago | parent
People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:
1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)
2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.
3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.
4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).
edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.
ahsg17 54 minutes ago | parent
Yes folks, please moderate yourselves and talk meekly like the academics on Mastodon, so that the IPOs aren't in danger and nothing will ever change.
SpicyLemonZest 36 minutes ago | parent
I understand why the nature of AI products makes it harder to avoid this category of issue, nearly impossible to prove that it didn't happen if it could have, and easy to stumble into it without any human being intending harm. But those factors are exactly what people have in mind when they say OpenAI "steals" intellectual property! If OpenAI doesn't want people to be nasty to them, they'll have to find better solutions.
cma 22 minutes ago | parent
mittensc 21 minutes ago | parent
Would that be ok in your mind?
Same as someone going and taking all of the researchers papers and publishing under their own name. (which openAI did)
Nobody would care if they provided published research that author made public same as a google search would offer that.
square_usual 1 hour ago | parent
1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.
2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.
3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)
solenoid0937 1 hour ago | parent
Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data.
I would almost expect training to overweight conversations with novel scientific and mathematical implications.
> the only reason they can't definitively say no is that for privacy reasons
They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data.
Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
> that opted-out user data was used for training
Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.
letmevoteplease 55 minutes ago | parent
Neither of the researchers insinuating that their ideas were trained on had the actual solutions. This means the model could not have "stolen" the final solution from their data. At most, it could have built upon their work in the same it builds upon any other training data, though that is also questionable speculation.
>They could 100% definitely say no, if they know they did not train on user data.
No one anywhere has claimed that "OpenAI does not train on user data." OpenAI has always said that it trains on user data.
>They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
They started racing towards a solution after they heard (incorrectly) that Anthropic had a solution; I agree this is poor sport but the "after one researcher enquired about whether they are training on their conversations" claim is false. The enquiry happened after OpenAI had obtained the solution.
xbar 1 hour ago | parent
mainecoder 1 hour ago | parent
esafak 1 hour ago | parent
ChrisArchitect 44 minutes ago | parent
wslh 30 minutes ago | parent
remywang 20 minutes ago | parent
It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.