102 points ChrisArchitect 1 hour ago 49 comments
Onavo 42 minutes ago | parent
It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.
I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.
croes 40 minutes ago | parent
Onavo 39 minutes ago | parent
simonw 37 minutes ago | parent
xp84 22 minutes ago | parent
celsoazevedo 15 minutes ago | parent
faefox 26 minutes ago | parent
bonoboTP 12 minutes ago | parent
Using the information for training purposes is not the same thing. Not legally the same and otherwise.
drdexebtjl 35 minutes ago | parent
xp84 24 minutes ago | parent
This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.
imglorp 4 minutes ago | parent
Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage.
The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.
KPGv2 3 minutes ago | parent
Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.
simonw 42 minutes ago | parent
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've ålso already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
packetslave 40 minutes ago | parent
bsimpson 19 minutes ago | parent
gambiting 16 minutes ago | parent
ValentineC 15 minutes ago | parent
unkeen 10 minutes ago | parent
luckylion 25 minutes ago | parent
Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.
toomuchtodo 24 minutes ago | parent
https://en.wikipedia.org/wiki/Tragedy_of_the_commons
(no affiliation)
ronsor 15 minutes ago | parent
On the other hand, the Internet Archive is a non-profit offering a free public resource.
toomuchtodo 12 minutes ago | parent
bradly 1 minute ago | parent
> Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly
CqtGLRGcukpy 38 minutes ago | parent
BeetleB 27 minutes ago | parent
I've not been able to access web.archive.org from my work computer - I always get the 429 error.
But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.
dotmanish 24 minutes ago | parent
flexagoon 23 minutes ago | parent
timpera 25 minutes ago | parent
Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.
xyst 9 minutes ago | parent
Are the abusive bots in the room with us?
ignoramous 3 minutes ago | parent