113 points pluc 1 hour ago 54 comments
Weryj 1 hour ago | parent
jacquesm 54 minutes ago | parent
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
pluc 45 minutes ago | parent
dgellow 22 minutes ago | parent
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
pingou 44 minutes ago | parent
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".
embedding-shape 30 minutes ago | parent
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
ben_w 5 minutes ago | parent
Any argument that writers and artists lose from these existing, would remain unchanged.
CJefferson 28 minutes ago | parent
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
hlynurd 27 minutes ago | parent
really not the same entities here
cgio 19 minutes ago | parent
ipython 4 minutes ago | parent
https://www.napster.com/blog/napster-heads-to-microsoft-buil...
hkt 13 minutes ago | parent
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
mitxela 27 minutes ago | parent
loloquwowndueo 11 minutes ago | parent
steveBK123 12 minutes ago | parent
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
fzeroracer 11 minutes ago | parent
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
Luker88 7 minutes ago | parent
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
CrimsonRain 19 minutes ago | parent
Razengan 18 minutes ago | parent
What if it was for free, like Wikipedia?
> Crimes this large are crimes against humanity.
jfc no, sit down.
Try doing something about actual evil shit like arms manufacturers and the politicians ordering the deaths of millions from the comfort of their sofas.
At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress; there's just too much shit to learn "manually" (wait I'm not advocating for low-effort slop, chill)
It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit.
steveBK123 11 minutes ago | parent
TacticalCoder 5 minutes ago | parent
It should be no surprise that it's a movement of lying thieves and scammers, Effective Altruism, that's behind close-source AI in the US.
Now I'm not sure justice won't come: SBF is behind bars for 25 years. He could turn out to not the be last scum from the EA movement to end behind bars.
As to open-weights models: at least it's not sold back at a mark-up and anyone can run them.
gyosko 45 minutes ago | parent
ozgung 37 minutes ago | parent
Neil44 36 minutes ago | parent
pluc 34 minutes ago | parent
proc0 35 minutes ago | parent
sajithdilshan 32 minutes ago | parent
leonidasrup 28 minutes ago | parent
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
menaerus 16 minutes ago | parent
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
TutleCpt 12 minutes ago | parent
bcjdjsndon 6 minutes ago | parent