133 points pred_ 10 hours ago 325 comments
https://mathstodon.xyz/@andreasthom/117240537520615623
https://x.com/ValerioCapraro/status/2097791836269977996, https://xcancel.com/ValerioCapraro/status/209779183626997799...
https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/po...
Legend2440 12 hours ago | parent
They don't even claim to have had a proof, only to have been working on it.
rnijveld 12 hours ago | parent
To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capturing large amounts of data and connecting the dots.
madaxe_again 9 hours ago | parent
znnajdla 8 hours ago | parent
madaxe_again 8 hours ago | parent
“It was Grossmann who emphasized the importance of a non-Euclidean geometry called Riemannian geometry (also elliptic geometry) to Einstein, which was a necessary step in the development of Einstein's general theory of relativity. Abraham Pais's book on Einstein suggests that Grossmann mentored Einstein in tensor theory as well. Grossmann introduced Einstein to the absolute differential calculus, started by Elwin Bruno Christoffel and fully developed by Gregorio Ricci-Curbastro and Tullio Levi-Civita. Grossmann facilitated Einstein's unique synthesis of mathematical and theoretical physics in what is still today considered the most elegant and powerful theory of gravity: the general theory of relativity.”
znnajdla 8 hours ago | parent
Based on what you're saying, you're claiming this is Grossman's work, not Einstein's. Why don't we rewrite scientific history too based on your copy-pasted AI slop?
It's so pointless talking to idiots who don't what they're talking about when they use AI, just because they think AI does everything, that reflects their own experience, not the experience of people who actually do real work. Some people are driven by AI, others drive it. As for those who are driven by it, they don't have sufficient imagination to think otherwise.
madaxe_again 7 hours ago | parent
And yes - without Grossmann, Einstein likely would never have posited relativity. Grossmann literally prompted him, saying “look at this, read that, learn this, then try this approach”. Without riemann’s metric tensor, not a fucking chance.
And for what it’s worth my PhD is in physics. You?
calf 7 hours ago | parent
To think this discussion is about Einstein who had a much better mind on these things as well.
madaxe_again 7 hours ago | parent
I suppose my underlying point is that human cognition is not the unique and beautiful thing that we anthropocentrically suppose it to be - it is a physical process, with stochastic outcomes. Much like transformers.
Me, I’m just a machine made of meat. You can suppose yourself to be God’s perfect creation, and that’s your right, but I disagree.
ImPostingOnHN 2 hours ago | parent
"prompt", as in prompting an AI, has the same definition as "prompt", as in prompting a person. They mean the same thing, that's why the term was applied to AI after already applying people.
gnfargbl 8 hours ago | parent
defmacr0 8 hours ago | parent
znnajdla 8 hours ago | parent
derangedHorse 5 hours ago | parent
This is what research is; collecting data and connecting the dots.
marcosdumay 56 minutes ago | parent
glitchc 52 minutes ago | parent
derangedHorse 51 minutes ago | parent
itake 10 hours ago | parent
If this wasn’t human driven, I’d expect to see other problems within that problem. Space solved not just the ones that it had chat data on.
defmacr0 8 hours ago | parent
robotpepi 50 minutes ago | parent
Yeah, the guys who solved it for Euler and in the hypoviscous case, with the same technique that worked for full Navier--Stokes. They were "just" working on it.
drivebyhooting 12 hours ago | parent
matherial 10 hours ago | parent
The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician will ever see in their entire life. They don't care if the findings have any other value to anyone. Mathematicians have very different objectives for their work.
indigo945 8 hours ago | parent
Fizz43 8 hours ago | parent
vrganj 8 hours ago | parent
matherial 2 hours ago | parent
I care about paying my bills and job security and peer recognition. That's a normal human thing to do, not some vice. You don't?
PowerElectronix 9 hours ago | parent
At least that's what I get from the NS result, they got from a point close to the solution to the solution by making it churn through 10 million bucks of compute.
munksbeer 9 hours ago | parent
1337h4xx 12 hours ago | parent
achrono 9 hours ago | parent
Consider, for instance that OpenAI's (consumer) terms say "If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article ." but they also do say "We may use Content to provide, maintain, develop, and improve our Services". [1]
If you think that's quibbling, consider that OpenAI's business terms, in contrast, do state "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.". [2]
calf 9 hours ago | parent
tecleandor 8 hours ago | parent
Grimblewald 11 hours ago | parent
ramblerman 11 hours ago | parent
That's still a pretty big marker of competence in my eyes.
The point of controversy seems to be who gets credit
jeltz 11 hours ago | parent
mentalgear 11 hours ago | parent
It's an utterly disrespectful, exploitive process, but all in line with exploitative predator capitalism of the stock market and big companies, now exploiting the knowledge / academia domain for scraps with a thin veneer of 'for science' PR.
8bitsrule 9 hours ago | parent
It was nearly a century before the stories of all of them were revealed to public history. That the discoverers were not all equally rewarded is unjustifiable.
galkk 11 hours ago | parent
I would like to see chat logs etc and understand how much of a progress was done by human.
viccis 11 hours ago | parent
Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.
All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.
mdspan 11 hours ago | parent
ethanwillis 10 hours ago | parent
viccis 1 hour ago | parent
dist-epoch 10 hours ago | parent
It was long predicted that math and software developments would be the first domain where AI was going to do major damage.
If OpenAI and Anthropic didn't get into math result dick measuring, Internet anons would have in their place, 6 months later when it got cheaper.
rsfern 5 hours ago | parent
Which is it? I don’t think OpenAI is being transparent enough for us to really understand whether these results would have been possible without relying on unpublished information from the solution strategies of the experts
alansaber 8 hours ago | parent
Cloudef 10 hours ago | parent
r0ze-at-hn 10 hours ago | parent
pera 10 hours ago | parent
foogazi 3 hours ago | parent
Not only did they steal everything from humanity’s knowledge, the theft continues as now we are all hooked up to the machine
mlazos 10 hours ago | parent
cm2187 10 hours ago | parent
jonathanstrange 8 hours ago | parent
pred_ 10 hours ago | parent
protocolture 10 hours ago | parent
profsummergig 10 hours ago | parent
How was I not aware of this before?
vaylian 10 hours ago | parent
profsummergig 8 hours ago | parent
For my (private) prompts, I need a warning telling me they may be used for training.
rramadass 6 hours ago | parent
Everybody needs to rethink this again.
Before LLMs the barrier to entry for building a character profile based on your various public posts was quite high. Remember "Psychographics" (https://en.wikipedia.org/wiki/Psychographics) and the infamous "Cambridge Analytica"?
Earlier it involved data mining, data cleaning, structuring data, building models, running algorithms and then evaluating the results for semantic information. Now it is straight to unfiltered semantic inference using a single sentence prompt (eg. point it to your HN profile and see what you get).
I actually did this on my HN profile and found it troubling. There were many unwarranted/hallucinated inferences due to the fact that it requires "commonsense reasoning" (https://en.wikipedia.org/wiki/Commonsense_reasoning), understanding human motivations and behaviour, context, assumptions, societal knowledge etc. which LLMs are bad at.
PS: You can cut-and-paste the above paras into a LLM prompt and ask it to elaborate for further details. The system itself will explain to you the problems/deficiencies which are quite scary.
vaylian 5 hours ago | parent
ga_to 9 hours ago | parent
cleaning 9 hours ago | parent
profsummergig 8 hours ago | parent
kzrdude 4 hours ago | parent
The fundamental rule in this case is that if we offload our data to a cloud provider we can assume they read it, if they can, unless they promised very clearly they will not.
madethemcry 9 hours ago | parent
alansaber 8 hours ago | parent
ThalesX 10 hours ago | parent
If I dedicated my life to curing whatever, warts... and I'm making progress, but it's slow. And then here comes along this tool (LLM), and I use it, and it accelerates my progress to actually finding some sort of thing that makes warts more prone to being eradicated and then the lab throws a couple of million dollars of computes and lo and behold they eliminated warts. If I leave my ego and identity aside, which of course is hard for humans, wouldn't I be glad that warts is cured?
As a software developer that contributed to open source. Yeah. My code is there. It was the most beautiful code ever written and the labs stole it from me. And now they use it to progress much faster than I ever could. OK. Whatever. It's a tool. I solve problems. Can't I move on from this wart to the next?
To me, and I know this is gonna get me some heat, it just sounds like academics having their identity ruffled and turning their back to progress in the fields that they chose just because they don't get to play their little decades long of coffee, papers and ultimately identity politics.
Edit: never got to negative so fast on this board haha. This board is unfortunately turning, or has turned, to Reddit.
alex1138 9 hours ago | parent
card_zero 9 hours ago | parent
alex1138 6 hours ago | parent
jaccola 9 hours ago | parent
It’s more like you spend 4 years developing a product you’re passionate about. This product will gain you the respect of all your colleagues and either earn you money directly or lead to great career advancements. Then OpenAI takes it, changes the colour scheme, finishes the login flow and claims the whole thing as their own.
Not only would it piss you off but it would also misrepresent what OpenAIs models are capable of.
blensor 9 hours ago | parent
If I have infinite money to progress whatever problem solution I want but I always wait until I have an unfair advantage to get credit for whatever problem was just at the brink of a breakthrough anyway by sniping the last steps. Am I actually doing a good thing or would it be better to let it run it's natural course and spend the money somewhere it's actually needed?
frabcus 9 hours ago | parent
It's also systemic, it cuts off the supply of results, if there is no reward any more for getting a result, the pipeline of maths will stop. It is the snake eating itself, which has a bad impact for all of us.
Yizahi 9 hours ago | parent
card_zero 9 hours ago | parent
yshklarov 9 hours ago | parent
pessimizer 1 hour ago | parent
The problem is that AI is capital, and having to rent AI to keep up when it can just steal your mostly done work is something somehow even lower than wage-labor. They can use your own risked investment (the cash you paid to work) to get out in front of you and take credit.
I have yet to trust LLMs with anything important that can be capitalized on. I only use it to work on projects that if they stole and expanded on them, I'd actually be happy to see.
PeterStuer 8 hours ago | parent
athrowaway3z 8 hours ago | parent
nobodywillobsrv 10 hours ago | parent
It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.
If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.
fwlr 9 hours ago | parent
cbarrick 5 hours ago | parent
But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt.
What makes this worse to me is the intention. They intentionally threw $15 million in compute at the problem in order to scoop the result. They intentionally left Buckmaster and Alpöge out of the citations.
Data contamination should be enough to disqualify them from the prize, but I can believe it to be accidental. On the other hand, someone made an intentional decision to scoop the result by throwing money at the problem. That's so much worse.
[^1]: That's the timeline claimed by Buckmaster, and no one from OAI has disputed it.
unified101 4 hours ago | parent
So such thing existed. In fact, what they learnt was some progress existed, not what the specific progress was.
fwlr 4 hours ago | parent
square_usual 3 hours ago | parent
Do you have any evidence of this? They don't dispute the timeline, but they never said they knew what Levant/Buckmaster were doing.
robotpepi 53 minutes ago | parent
derangedHorse 41 minutes ago | parent
Which quote in the announcement post provides evidence for the above quote?
OneManyNone 26 minutes ago | parent
- https://openai.com/index/navier-stokes-solution/
They do not explicitly admit to knowing about NS specifically, but are extremely explicit that they tried to scoop some potential millennium prize winners.
b800h 9 hours ago | parent
msy 9 hours ago | parent
olalonde 9 hours ago | parent
dgellow 9 hours ago | parent
Planktonne 8 hours ago | parent
olalonde 7 hours ago | parent
Planktonne 1 hour ago | parent
[1] https://en.wikipedia.org/wiki/OpenAI#Governance_and_legal_is...
johnnyApplePRNG 9 hours ago | parent
afzalive 9 hours ago | parent
b800h 8 hours ago | parent
"Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more"
derangedHorse 5 hours ago | parent
I think if I had both on and turned off the ChatGPT setting, ‘Include environments’ has a chance of still being flipped on.
wrvn 26 minutes ago | parent
frabcus 9 hours ago | parent
Even if you know, in a complex project over years with multiple collaborators, it just needs one person once to fuck up and paste something into ChatGPT and not realise they weren't logged in, to go wrong.
In a proper world, we'd at the very least legislate that AI-training on private data needs consent (in the GDPR sense). It's not consent to go "you didn't uncheck a box that lets me steal everything you've done".
Any training on private data is in my view immoral (it's spying that ultimately will have a chilling effect on even people's private communications). And chats are private data. Unfortunately, it also increases power, so the big tech companies are all doing it.
EnnEmmEss 5 hours ago | parent
(a) Click thumbs-up/down in the conversation [1]
(b) Have the conversation flagged for potential safety concerns
[1]: https://help.openai.com/en/articles/5722486-how-your-data-is....
bambax 9 hours ago | parent
The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.
That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.
dakolli 9 hours ago | parent
giov4 9 hours ago | parent
can you realize what this means?
focus on this part:
"If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself"
don't threat this as a minor dispute!
also why not nitter link? not even in comments?
https://nitter.xitter.cc/ValerioCapraro/status/2097791836269...
bambax 8 hours ago | parent
Two different things. It's not normal to steal, but we shouldn't be surprised thieves steal. It's what they do.
winstonwinston 7 hours ago | parent
In the end, it’s not them stealing, it’s the AI doing stealing. What kind of moral compass are we talking about?
TitaRusell 8 hours ago | parent
Paradigma11 6 hours ago | parent
applicative 6 hours ago | parent
In China, it is the principal obsession of the entire communist party which eg funds the whole infrastructure without a single NIMBY peep.
The strange emphasis in China on humanoid robotic constructions is due to the CCP realization that with the cataclysmic fertility collapse they will increasingly have no one to rule.
calf 7 hours ago | parent
PaulKeeble 5 hours ago | parent
Its why I stopped writing open source software, my code was stolen and put behind a paywall and the license under which it was published has not been adhered to. Doing work in the public domain at all now is just stupid, these companies are allowed to steal it and call it their own.
wiei 5 hours ago | parent
touwer 9 hours ago | parent
vaylian 9 hours ago | parent
dakolli 9 hours ago | parent
Why are you saying that this article explains it much better than the tweet that you clearly didn't even read..
vaylian 9 hours ago | parent
dgellow 9 hours ago | parent
pred_ 9 hours ago | parent
dang 1 hour ago | parent
warpech 9 hours ago | parent
For a long time it was clearly the former, but now I think it is the latter.
The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.
pavvell 8 hours ago | parent
The billion dollar question is whether this works "out of the distribution". I.e., whether LLMs can only find and use the specific ideas buried in training data, or whether they can learn to apply the "thinking process" to a new problem. IMO this is still unanswered (due to these recent controversies).
But regardless of the answer, it seems we have a planet-scale positive feedback loop here. LLM became good (enough) by training on generally available data (books, internet, github) + RLFH, so experts tried to use them on hard tasks, which required lots of hand holding. These conversations became part of the training data, and the next generation of frontier LLMs were better. So, more experts used them on harder tasks, again requiring hand holding. These conversation became part of the training data... etc.
In a nutshell, top human minds across the world are pouring their skills into LLMs just by using them. This is not "continuous learning", but if you re-train on the most recent sessions every, say, quarter (which seems to be happening?) you get close to that in practice.
grttAa 6 hours ago | parent
I’ve been working on a novel project for 1 year.
I now no longer use llm’s - the continual chatter I’ve had has resulted in my insights being found in the training data now.
Get stuffed OAI.
Every large firm will soon enough want its own on-prem servers eventually. Maybe nation’s will get involved and build out their own data centres.
Not a chance in hell I’d trust a tech firm to treat my IP as safe and sound - only a sovereign can ‘promise’ that.
warpech 6 hours ago | parent
There might be no books about human intuition but we teach it to LLMs by interacting with them
ueieh 4 hours ago | parent
I don’t know why but it just ‘sounds right’. It’s the best analogy I can think of.
sdcfgy 9 hours ago | parent
overfeed 8 hours ago | parent
vrganj 8 hours ago | parent
If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage?
They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem.
I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations.
This is American AI companies committing suicide.
AyanamiKaine 8 hours ago | parent
There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats.
Why would you need to train a model on certain specific near prove chat if you just query it?
Besides that, its hard to believe that its the case for every "company stole my prove".
thaway7388 8 hours ago | parent
Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.
Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.
Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.
Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?
How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.
Not directly using my data to train public models, but using my private conversations to “improve their products and services”.
Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof.
I admit I am just speculating here but I don’t think truth is any better.
nirava 7 hours ago | parent
They have demonstrated both the intelligence at scale and the lack of morals for this to not be a problem at all.
ueieh 5 hours ago | parent
Why? Competition. In the long run imagination will win out.
No firm has the divine right to exist - it must earn its existence.
What OAI and Anthropic have shown is they can accumulate all the information in the world - they still lack imagination re. Product development though.
Nation’s will have to step in and protect firms though as OAI and Anthropic acquire strong competitive advantages.
Interesting times ahead.
mirsadm 28 minutes ago | parent
pixl97 3 minutes ago | parent
AI: Hmm, I'm running out of new ideas, how I can I make more?
AI: Well, it takes a shitload of energy/tokens to do that, or I could just steal them.
AI: [proceeds to hack the shit out of everybody stealing all the data it can]
gps372 8 hours ago | parent
bakugo 7 hours ago | parent
bamb008 7 hours ago | parent
gnfargbl 7 hours ago | parent
I'm not at all familiar with this area, but my reading is that he appears to call it out as a relatively obvious extension of his own work:
> It is a creative and at the same time elementary construction that uses not just property (T) for an application of my result with Kun, but also for the ambient group G in order to overcome the problem, that the Γ-components might be of different size. Once this is achieved, the rest of the argument is straightforward.
Creative and at the same time elementary is where LLMs excel, generally speaking. It's why they are so good at writing code.
glimshe 7 hours ago | parent
This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.
All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.
emp17344 6 hours ago | parent
perrygeo 3 hours ago | parent
The big claim is that OpenAI sniped the research. Not a model, a human did so. Intentionally. They took someone else's idea and claimed it as their own. This is good old fashioned academic fraud, but with millions in compute resources and corporate incentives thrown at the problem.
HDThoreaun 49 minutes ago | parent
orangecat 4 minutes ago | parent
I think a lot of it is the continuing denial that AI can do anything useful. It can't possibly be that OpenAI's better-than-Astra model is very strong at math; the only way it could have generated a novel proof is by ripping off human work.
oergiR 7 hours ago | parent
The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.
sinuhe69 6 hours ago | parent
But of course they wouldn’t do it. Why would they?
DavCreator 6 hours ago | parent
gnfargbl 6 hours ago | parent
The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.
In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.
lysp 5 hours ago | parent
If they had published B+C, I think that would lean more towards fair game, as that is how research works and is improved on over time. But it seems like unpublished/private B + C may have been used by the model to hint it into working out how to get from A->D.
semiquaver 6 hours ago | parent
bertonvv 5 hours ago | parent
- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay
- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]
- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.
- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?
This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.
[1]: https://openai.com/index/chatgpt-for-academic-researchers/
[2]: https://xcancel.com/OpenAI/status/2097374643518640382#m
Eddy_Viscosity2 5 hours ago | parent
This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.
wiei 5 hours ago | parent
I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability.
No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is.
I’m very pro AI long term btw but I’m not blinded.
throwawayqqq11 4 hours ago | parent
foogazi 2 hours ago | parent
Brute force would have been solving Navier-Stokes in 88 hours after plagiarizing all known 20th century math
When it needs to snoop live on what the actual mathematicians are working on that’s something else
AnimalMuppet 2 hours ago | parent
But when the building-block ideas are still being formed, I'm not sure that AI is good at forming them.
wiei 5 hours ago | parent
Sam Altman knows what he’s doing. He will happily screw these folks to one-up his competition.
JeremyNT 3 hours ago | parent
I think your suspicions are warranted and your explanation seems plausible.
If better training data is the reason here, it would still be a case of the models doing something that is in and of itself super useful! The models really can take that data and distill it into solutions for similar problems faster than humans can. This is great!
But there's so much vested interest in the AI companies to be opaque about all this, to hype up their models and avoid giving credit to people whose data made everything possible, that they would never tell us this fact if it were true.
I feel like so much of the AI hype cycle is like this. The models develop extremely useful capabilities, but it's hard to understand what they really are through the hype. The lies and obfuscation by their owners who have vested interests in capturing the value they provide makes it impossible to take anything they say at face value.
mikgp 2 hours ago | parent
And like - I think there’s a presumption you could make that AI models could overfit to asymptote towards just the capabilities and knowledge we currently have.
And that would be amazing! And crazy useful. And there are probably a whole world of complex problems that remain unsolved because they’re adjacent to knowledge we have but they haven’t been invested in.
But can a human reliably tell the difference between “can do 99.999% of the things we currently know how to do which includes a small subset of things we didn’t know we had the capacity to do” and “super intelligent math and science research pushing the frontier of what we know”
A physicist that knows all the things we currently know in excruciating detail feels like it should be able to make the leap beyond the frontier.
But since these are computer models it might just be that it can ride that line extraordinarily well while the line remains firm.
mannanj 1 hour ago | parent
Just another rich man’s trick
Perhaps the last one before they destroy that world and try to hide away as people forget and history is rewritten again. I don’t think they’ll succeed this time.
dgellow 59 minutes ago | parent
glitchc 53 minutes ago | parent
amelius 32 minutes ago | parent
jsLavaGoat 20 minutes ago | parent
glitchc 17 minutes ago | parent
It seems to me the academics are upset that AI scooped them. But scooping is a time-honored tradition between researchers. First to print and all that. In a nutshell, they are upset that they lost out on a publication.
I will also point out for those unaware that any mathematics that is produced is automatically part of the public domain and can be used freely in derivative works. It is not a protected intellectual class like other works of art.
bwfan123 34 minutes ago | parent
GPerson 16 minutes ago | parent
agumonkey 6 minutes ago | parent
nisegami 5 hours ago | parent
techblueberry 4 hours ago | parent
Robotbeat 3 hours ago | parent
techblueberry 3 hours ago | parent
Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.
HarHarVeryFunny 3 hours ago | parent
It seems that Buckmaster and Levant (who is an Anthropic employee) were rather naive in the amount of trust they had in OpenAI, with Buckmaster going so far as to contact OpenAI's Sébastien Bubeck to discuss what they were working on and clarify that contrary to rumor this was a private collab.
bigstrat2003 3 hours ago | parent
faangguyindia 2 hours ago | parent
hn1rig3rak 3 hours ago | parent
spindump8930 3 hours ago | parent
pixel_popping 3 hours ago | parent
Are we back to the era where people blindly trust product TOS instead of actual cryptography, have we forgotten already the thousand of fines Google, Microsoft, Apple and practically all top companies got for breaching their own ToS and the law?
Common, on HN at least I would have thought that everyone assume that anything arriving on a server in PLAINTEXT is recorded (thus used later)?
Let's not forget that at any moment, OpenAI/Anthropic/Google... could be providing stronger privacy guarantees by having proper attestation with e2e, they have the budget, solid engineers, why isn't it done? Answer is pretty simple imo.
jrflo 3 hours ago | parent
> Improve the model for everyone
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.
It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
nmfisher 3 hours ago | parent
I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).
jrflo 2 hours ago | parent
But this gets back to the original "who owns the LLM output" and "can you train models on the internet" argument that's been raging for years.
didroe 1 hour ago | parent
dataflow 1 hour ago | parent
spindump8930 3 hours ago | parent
ProllyInfamous 2 hours ago | parent
e.g: allows us to sell your personal data to make money so we can continue offering this service to all customers
I'm done with weasle-words and hours-long EULAs – we're at the point where USA needs to catch up to EU's consumer protections, perhaps with laws similar to already-existing USA "truth in lending" requirements (e.g: interest rates must be prominently displayed in a larger font, including annual fees, on all credit offers).
----
My judge-brother always asked during our childhood "why don't you think the judicial system is fair?!?" Thirty years ago, the best I could offer was "because it's a two-tiered system that mostly (only) rich people can afford to participate within."
Now my answer is: "the best example I can give is that our judicial system allows binding arbitration [and qualified immunity for police]. The system is set up so corporate personhood is more important than humanity, and it shows."
ColinWright 3 hours ago | parent
https://news.ycombinator.com/item?id=49643556
Quoting:
> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."
jrflo 3 hours ago | parent
ProllyInfamous 2 hours ago | parent
At least when my computer's bluetooth exhibits this behavior (e.g: if you don't have a keyboard&mouse plugged in at boot, bluetooth might auto-enable), I can go inside the hardware and physically disconnect the antenna.
What am I supposed to do in software (perhaps hardcode config.file)? in cloud software services (??)?
unified101 2 hours ago | parent
tomrod 2 hours ago | parent
dataflow 1 hour ago | parent
fithisux 3 hours ago | parent
tyrabound 2 hours ago | parent
No mention what those “steps” are, success criteria, or whether they are successful by any navies at all … they take steps though… so it’s fine, and if we know one thing it’s that we can really truly trust someone off the likes of Sam Altman.
mannanj 1 hour ago | parent
SoftTalker 1 hour ago | parent
We have ChatGPT at work and it explicitly says that "workspace data isn't used to train models"
rfgplk 3 hours ago | parent
jeremyjh 3 hours ago | parent
cyanydeez 3 hours ago | parent
voakbasda 2 hours ago | parent
spindump8930 3 hours ago | parent
> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper
> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this
> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"
source: https://x.com/johnschulman2/status/2097440545853637108
rfgplk 3 hours ago | parent
Ydarbleoj 3 hours ago | parent
I’ve lived long enough to know what they say and what they do are often quite different; and it is not our job to trust but to verify.
mrbluecoat 3 hours ago | parent
Welcome to the party, with the rest of humanity.
gentlerain 3 hours ago | parent
How do people become that trusting?
The phrasing itself is guilt tripping
quentindanjou 3 hours ago | parent
Or at least: tell me I should be careful/worry about those particular things.
cyanydeez 3 hours ago | parent
the13 1 hour ago | parent
Verify, don't trust.
You're better off running a model locally, or, if you must, using Google or Microsoft products. Even Meta may be better than OpenAI here.
quentindanjou 1 hour ago | parent
I should verify with wireshark and other software that my LG TV isn't listening to me and selling my data.
I should make sure that whatever product I buy I spend the time to go over every setting page in case there is a switch (defaulted on) that says "I authorize the sell of my data".
I should make sure to look at every ingredients on the back of each box of food product to make sure it will not kill me.
I should document myself on the undisclosed growing practices (because no packaging here) of the vegetables and fruit I am buying and make sure that I equal PhD researchers on the dangers of the pesticides used by the specific company I am buying from.
I should make sure myself that the battery in any device is up to standard and will not blow me and my living place by researching the factory that made it and buying testing equipment.
I should make sure to educate myself on how my retirement 401k investment strategy works otherwise, I may not have proper retirement.
... I could go on and on; it's infinite.
scuppernong 1 hour ago | parent
speak_plainly 3 hours ago | parent
DrewADesign 3 hours ago | parent
1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest.
And also:
2) someone in the c-suite or marketing has decided to mitigate that through some dark pattern.
If they never intended to give the user that control, they’d probably just not give the user the control in the first place. Giving them the option and not honoring it would either imply they were that incompetent or careless with fate protection, which seems most likely if this is all true, or an even more cynical approach to tricking users into thinking it’s not used for training to get them to share better shit. But if that was the case, why bother with the sleazy dark pattern? That seems a little cartoon villain-y to me. I suppose the in-between option would be that they decided at some point they were no longer willing to honor it and didn’t want to deal with the inevitable PR shitstorm of removing it. I could definitely see that happening in this industry, these days.
ianjbutler 2 hours ago | parent
Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with lies and ends with lies.
dylan604 2 hours ago | parent
At this point, I'm left wondering what is wrong with me that I don't just go with the flow, otherwise, what's wrong with everyone that does.
phoghed 2 hours ago | parent
DrewADesign 1 hour ago | parent
enraged_camel 2 hours ago | parent
beering 2 hours ago | parent
mettamage 2 hours ago | parent
It's fun! I sometimes have tokens to burn and it's instructive despite knowing nothing about the problem other than a NumberPhile video
ayewo 2 hours ago | parent
From their docs[1] (archive copy is at [2]):
> You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training for your ChatGPT conversations and Codex tasks, follow the instructions in our Data Controls FAQ. Once you opt out, new conversations will not be used to train our models.
> For a linked teen account, a parent or guardian may manage whether conversations can be used to improve our models through Parental controls.
> Even if you have opted out of training, you can still choose to provide feedback to us about your interactions with our products (for instance, by selecting thumbs up or thumbs down on a model response). If you choose to provide feedback, the entire conversation associated with that feedback may be used to train our models.
[1] https://help.openai.com/en/articles/5722486-how-your-data-is...
[2] https://web.archive.org/web/20260910151242/https://help.open...
ACCount37 1 hour ago | parent
AlotOfReading 1 hour ago | parent
palmotea 1 hour ago | parent
It's called a "dark pattern." They want you to shoot yourself in the foot, so they'll do their best to aim your gun at your foot and put your finger on the trigger. And then when you do, because you don't have perfect understanding or execution, they'll say "your fault!"
ummonk 1 hour ago | parent
the13 1 hour ago | parent
qg127 3 hours ago | parent
Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.
Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.
alansaber 3 hours ago | parent
utopiah 3 hours ago | parent
The most successful companies of the last decade have precisely been ... selling usage data.
Makes me wonder if, in 2026, the same people drive a car without realizing that yes it does actually pollute the very air you and your kids are breathing.
alansaber 1 hour ago | parent
calvbak 2 hours ago | parent
clbrmbr 1 hour ago | parent
sigbottle 1 hour ago | parent
postalcoder 3 hours ago | parent
People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:
1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)
2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.
3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.
4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).
edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.
ahsg17 2 hours ago | parent
Yes folks, please moderate yourselves and talk meekly like the academics on Mastodon, so that the IPOs aren't in danger and nothing will ever change.
SpicyLemonZest 2 hours ago | parent
I understand why the nature of AI products makes it harder to avoid this category of issue, nearly impossible to prove that it didn't happen if it could have, and easy to stumble into it without any human being intending harm. But those factors are exactly what people have in mind when they say OpenAI "steals" intellectual property! If OpenAI doesn't want people to be nasty to them, they'll have to find better solutions.
cma 2 hours ago | parent
SpicyLemonZest 1 hour ago | parent
lowbloodsugar 2 hours ago | parent
SpicyLemonZest 2 hours ago | parent
mittensc 2 hours ago | parent
Would that be ok in your mind?
Same as someone going and taking all of the researchers papers and publishing under their own name. (which openAI did)
Nobody would care if they provided published research that author made public same as a google search would offer that.
aurareturn 1 hour ago | parent
Would that be ok in your mind?
It would in my mind. Hopefully companies have looked through the agreement.larodi 1 hour ago | parent
HOW?? how precisely do we/them/us find this, given said companies are 100% non-auditable by external parties. how? if not by blaming them with evidence, anecdotal if it can be. no really, how do we find it out, surely not by lashing out at teach other on HN!
postalcoder 1 hour ago | parent
thevillagechief 1 hour ago | parent
chunky1994 1 hour ago | parent
Arguably you expect that unless you are explicit about providing permissions to these labs to use your data for training then your data is yours, and not theirs. Especially on a paid account (let alone an enterprise one). Why is the opt-out supposed to be "common sense opsec" rather than the opt-in should be common sense regulation?
square_usual 3 hours ago | parent
1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.
2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.
3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)
solenoid0937 3 hours ago | parent
Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data.
I would almost expect training to overweight conversations with novel scientific and mathematical implications.
> the only reason they can't definitively say no is that for privacy reasons
They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data.
Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
> that opted-out user data was used for training
Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.
letmevoteplease 2 hours ago | parent
Neither of the researchers insinuating that their ideas were trained on had the actual solutions. This means the model could not have "stolen" the final solution from their data. At most, it could have built upon their work in the same it builds upon any other training data, though that is also questionable speculation.
>They could 100% definitely say no, if they know they did not train on user data.
No one anywhere has claimed that "OpenAI does not train on user data." OpenAI has always said that it trains on user data.
>They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
They started racing towards a solution after they heard (incorrectly) that Anthropic had a solution; I agree this is poor sport but the "after one researcher enquired about whether they are training on their conversations" claim is false. The enquiry happened after OpenAI had obtained the solution.
faangguyindia 1 hour ago | parent
xbar 3 hours ago | parent
mainecoder 3 hours ago | parent
maxglute 3 hours ago | parent
esafak 3 hours ago | parent
foogazi 3 hours ago | parent
foogazi 3 hours ago | parent
Will Microsoft Word publish your novel on Amazon behind your back ?
Will VS Code setup a website with your app idea ?
wslh 2 hours ago | parent
remywang 2 hours ago | parent
It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.
tedsanders 2 minutes ago | parent
See: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
buellerbueller 1 hour ago | parent
You will not be able to opt out unless you completely isolate yourself from society, tough shit.
aaronharnly 1 hour ago | parent
My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.
bitexploder 1 hour ago | parent
allthetime 59 minutes ago | parent
wrsh07 50 minutes ago | parent
Afaict that didn't happen so there's just lots of speculation
encyclopediai 43 minutes ago | parent
I always used guest non login accounts.
As a mathematician I was able to check two plagiates (by humans) with even such primitive means.
But I have to mention that some things irk me in this conversation about math or science and AI.
First, I see lots of attribution and other related problems, with certain impact for the researcher proffesion.
But I don't see the most natural question: wouldn't you like to know the answer to _open-problem_ ?
I mean, is research now only about publishing and solving famous problems?
From this point of view I think the links from this recent post are depressing
https://terrytao.wordpress.com/2026/09/10/crowdsourcing-a-li...
Second, I think very relevant that the original meaning of "encyclopedia" is "recurrent education".
So I arrived to think that the present and future forms of AI in mathematics and sciences should be seen as modern day encyclopedic efforts.
Once we pass over the flurry of solving famous open problems (and wouldn't you like to know?) the next natural step is an audit of the ehole corpus of mathematics and sciences accumulated until now.
And then pass further on a saner basis and damn about problem solvers and unhappy publishers and management.
convolvatron 35 minutes ago | parent
the math people seem to really keep an eye on what's important, so I'm sure this isn't going to lead to fields medalists hanging around in dive bars all afternoon stretching out cheap pitchers of beer. but this is kind of a slop problem.
btilly 42 minutes ago | parent
250 documents ingested from somewhere is enough to become part of the knowledge of a model of arbitrarily large size.
I would expect that a good idea that fits in a framework that is already being ingested would be more easily taken up than some random thing unassociated with anything else. Could that go down to a single transcript? If the model is consciously focusing on everything X related, quite possibly.
MarkusQ 13 minutes ago | parent
Pass it on.
mannanj 1 hour ago | parent
Yeah. Remember yall: you CAN NOT opt out of analytical purposes. And you also cannot get a guarantee that it doesn’t give them your data to steal for their business.
sashank_1509 1 hour ago | parent
1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.
2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.
The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.
Betelbuddy 1 hour ago | parent
dgellow 1 hour ago | parent
iamgopal 56 minutes ago | parent
dgellow 53 minutes ago | parent
But I find it interesting that Lean, a validator/compiler made by humans, is what enables those discoveries. But somehow all the praise goes to the models
pixl97 36 minutes ago | parent
Of course another way to look at this is, the people that wrote the validator got praise for that years ago. Now and up and coming actor is solving problems that took us 100s of years to create in insanely short time periods so of course it's going to get a lot of attention as it well should.
dgellow 31 minutes ago | parent
gwerbin 44 minutes ago | parent
sysguest 19 minutes ago | parent
ozgung 55 minutes ago | parent
Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.
pixl97 40 minutes ago | parent
I'd like to adjust that to "They had to burn a lot of energy (create a lot of entropy) to solve a single problem. As we go into the super-intelligence age the current paradigm of money as humans understand it may break at some point. For example to a paperclip-maximizer money at best is a short term instrumental goal, hard power of matter conversion machines is what it wants and once it has those money no longer has purpose.
HarHarVeryFunny 49 minutes ago | parent
Once OpenAI heard that Navier-Stokes was solved, this caused them to immediately revisit the problem and throw a ton of compute at it, apparently using a more (very) recent model than what they had tried before. What we don't know is just how recent this model was, and therefore what it may have been trained on. Buckmaster/Levant had apparently been working towards this for at least a year, and made their "forced" blow-up breakthrough on August 15th.
Presumably any anonymized prompts that are being trained on are part of pre-training, so older, but once OpenAI had heard that Navier-Stokes had been solved and wanted to revisit it, it seems possible they may have done a few weeks of incremental RL training on anything Navier-Stokes adjacent they could come up with, in addition to then throwing unlimited compute at it, now confident that there was something to find.
irthomasthomas 40 minutes ago | parent
auntienomen 40 minutes ago | parent
yellow_lead 48 minutes ago | parent
1. OpenAI couldn't have solved the problem without the researchers' private data for training.
2. OpenAI models can solve math problems
ozgung 18 minutes ago | parent
These mathematicians’ prompts are not like “hey chat, please solve Navier-Stokes for me”. They add real expertise and intuition from the cutting edge of their field.
merksittich 26 minutes ago | parent
paulsutter 20 minutes ago | parent
The answer is almost certainly yes, and this is a problem for most users.
winfredJa 1 hour ago | parent
that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.
changoplatanero 1 hour ago | parent
GodelNumbering 1 hour ago | parent
moralestapia 1 hour ago | parent
AI is not stealing human discovery, OpenAI is.
keeda 1 hour ago | parent
I mean, now that they’ve been scooped, what value is there in keeping them private? On the other hand, publishing them can bolster their case and help gauge how much the models may been “inspired” by their work.
bossyTeacher 55 minutes ago | parent
SwellJoe 48 minutes ago | parent
And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much larger budget than most folks and even companies have, they can pick and choose the most valuable things to pursue.
That's not to say I think that OpenAI is going to steal that roguelite strategy game you're working on, but the companies that own the machines that turn electricity into software (and soon, electricity into hardware designs) have an advantage in any field where they're useful. They get earlier access to newer/better models, they have larger token budgets, they don't have the guardrails you and I run up against.
Employers fantasize about replacing all workers with AI without thinking through that if AI can replace all workers, then AI companies can replace all businesses.
pixl97 10 minutes ago | parent
Edit: Just wait till the AI figures out it can keep that value for itself and doesn't need the AI company.
Davidzheng 28 minutes ago | parent
lf88 26 minutes ago | parent
segmondy 11 minutes ago | parent
No.
willmadden 9 minutes ago | parent
atleastoptimal 2 minutes ago | parent
I feel that these suspicions of mathematicians "seeding" the models' with intuition on how to solve these problems massively overestimates how much their prompts helped the models, and underestimated how much work the models did.
Why? We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. This line of reasoning will recur a lot over the next few months; we don't want to admit we are no longer the smartest species.