- cross-posted to:
- webdev@programming.dev
- cross-posted to:
- webdev@programming.dev
The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. “Scraping” is facilitated by design!
Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.
That’s true but they probably didn’t account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves for profit.
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.
Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.
That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API
That’s simply not true. These bots are essentially DDOSing the entire internet, API or not.
Okay, if efficient APIs existed and they weren’t incompetently failing to use them, it wouldn’t be a problem. Happy now?
(I should’ve addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn’t be necessary because, as you can see from this text, it would’ve cluttered up my writing with double the words.)
Well they had an API, but…
But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.
It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”
I think it would be a problem because the scrapers are hammering all types of websites from small forums to reddit with tens of thousands of unique ip addresses at a time. Websites that have neither the money, hardware, or protection had to figure out solutions really quick or suffer what is essentially a constant ddos attack. This is the reality of the web now, it’s just an incredibly hostile place.
Reddit does have RSS feeds
Semantic web was a separate initiative by Tim Berners-Lee when the web already used un-semantic HTML. And it never went anywhere.
This is because appending
site: reddit.comto a search query is basically a surefire way to find results written by genuine humans.The article is from 2026, not 2016? That bot-ridden Reddit? Am i in the wrong film?
Meanwhile, I use
-site:reddit.commore and more.
We are not the same. ;)I use Google hit hider by Domain userscript (also works on most major search engines), which also let’s you remove AI barf sites.
100% bot ridden. However site:reddit.com is used for the types of questions where otherwise, you just get page after page of SEO sites, which are even wose.
wrong film
Second reality on your left, past the one where Santa Clause rules the world
Hey, those are real human made bots. Not shitty ai generated bots.
Or that site:url doesn’t return results it used too anymore.
Reddit seems to be in the business of extracting as much value as it can from said forums without completely destroying them.
I call bs. It’s been completely destroyed for a while.
I think they’re more upset AI companies scraped 'em and they didn’t get paid.
They told Google to pound sand and their stock took a hit. Fuck em
You mean Reddit told Google? They actually have an agreement with Google where the latter get a direct feed of new posts and comments, and index them pretty much immediately. So not sure where you got the ‘told Google to pound sand’ idea.
https://www.cnbc.com/2026/07/22/reddit-stock-google-ai-content-deal.html
They threatened to cut Google off from training AI and their stock dropped.
They are in the business of attracting new users into their new algorithmic engagement hellhole now. Show any propensity of interest, and you will get sidetracked to the most godawful side-communities that seem to have emerged to engage as many victims as possible. They do not want to focus on their old users as anything less than the content they already made that makes reddit show up as free advertisement to their new base in search engines. It’s all a game of “it’s the algorithm’s fault so you can’t blame us” now.
I agree with you. But what’s important is if they can sell anything. They destroyed their own product but they’re still pretending it has value, and maybe they can fool some investors into paying for script and AI slop.
While I agree, More celebs than ever have been asking for comments and ideas on their reddit pages.
If you consider scraping a threat then yes plain HTML might as well be giving up. The advantage of new reddit for that is quite clear: they can collect a bunch of data about your browser before deciding if you are a bot and if the rest of the page should load. The embedded recaptcha call in the screenshots is a pretty good hint. I suspect blocking trackers on new reddit will break as soon as the scrapers move over.
As for why they don’t just kill old reddit: a significant chunk of their active posters use it and are attached to it. So if they kill it entirely they will lose content. Posters are of course logged in so this change is less likely to affect them.
if you are a bot
Or they’re bouncing banned people.
Fuck Reddit and Fuck Spez.
even logged in I can’t access old reddit anymore.
if you use the reddit enhanced suite extension you can get redirected
If I wasn’t permabanned I might bother trying work arounds.
I just checked my old.reddit login on the desktop… still working just fine today. I don’t think you need work-arounds, you just need to access it without any work-arounds.
You think that if I do the thing that’s not working that it will somehow work.
old.reddit redirect addon just got a update that deals with that problem.
https://addons.mozilla.org/en-GB/firefox/addon/old-reddit-redirect/
Thanks. That was too easy not to do.
working for me
Fuck you spez.
How dare you question King Steven the Turd, Greediest of Pigboys? If he proclaims HTML to be unsafe, it must be so. That’s a King’s job, to tell the Landed Gentry how to behave…
I’m not forgetting or forgiving that shit either.
I don’t even understand how there are still so many comments on that site. How is anyone even accessing it anymore? I just assume it’s 100% bots
The site is most hostile to visitors who aren’t logged in, and the users who comment probably mainly visit the site while logged in.
Nah I have been banned loads for insanely minor stuff. I got a sitewide ban having appealed a ban for some star trek opinion I cant remember.
My account was eight years old with a decrnt history and contributions. They give mods too luch sway because mods are losers doing work for free.
Inertia, just like Facebook.
I now browse Wikipedia. Please don’t screw me over Wikipedia, I donated five bucks to one of your nags once.
For anyone who doesn’t know, you can download Wikipedia and host it yourself! I got the top 50k version (~7G) on my RPI3 and now no matter what fuckery they pull or the government pulls, I’ve got a pretty decent source of general information.
Could you share what you did to achieve this? I’ve been planning on doing just that, and have it auto-update every week or so (keeping the previous versions archived, of course) by using kiwix-serve for a static ‘.zim’ file and maybe a cron job for the auto-update. But if you have a better solution, I’d love to know. The deployment I am planning is kind of convoluted to be honest.
No, that’s exactly what I did, I’m running kiwix-serve, but I’m not going to bother updating it because I’m really worried about information degrading now that fascists are basically calling the shots on everything (and WP’s jackboot co-founder has a hard on for it). If I feel enough time has gone by to warrant an update I’ll just do it manually.
I am currently pretty tapped out on all storage and backup drives but if I had space this post would have motivated me. Just sayin
How do you do that?
You run a server on a machine inside your house, it can be any computer on your local LAN/wifi, but it’s obviously best if it’s a machine that’s always on. I use a Raspberry Pi 3B+ (these can be had for about $50) that I have plugged into my wifi router and it’s running a little program called Kiwix-Server (free and open source). You download the WP file (it’s a huge single file with a .zim extension) and point the server to it and boom.
Wow, neat! Thanks for the explanation
Yesterday, I visited new Reddit. No VPN, Chrome in incognito mode, so no extensions, clickedon a link directly from Google search. Got a message that my access was blocked for security reasons. Copied the link to Firefox (with uBO and a few privacy-centric extensions), changed it to old reddit (where I was already logged in), and it worked just fine. I found out that when Reddit kills old reddit, I won’t even have the choice to switch to the new one (not that I ever would) because I’d be blocked anyway.
It works through Tor, by the way. I had the same experience can’t look at it but I can through Tor. They even have a .onion.
seems like Old reddit is account-walled as of today maybe
reddit stinks super bad
a corpse of a platform
They’ve already disabled old.reddit for me. That makes it unusable, and thank you, Reddit. I actually used a domain blocker to block reddit, but generally would still be tempted to peek. Now, that is no longer the case. I, for one, am wholly in support of Reddit’s new anti-advertising stance!
I haven’t maintained my personal website in ages. That, of course, means it loads instantaneously, in comparison to all these Library of Congress websites.
funny, that comes just on the heels of my deciding that reddit is unsafe.
You’re just now realizing it?
not really but it make for good joke pacing.
i’ve been here for a while. and not there for a while.














