- cross-posted to:
- webdev@programming.dev
- cross-posted to:
- webdev@programming.dev
Reddit seems to be in the business of extracting as much value as it can from said forums without completely destroying them.
I call bs. It’s been completely destroyed for a while.
Whether or not it requires a login for old.reddit.com depends on client IP. I see this on a few networks, not on others.
reddit stinks super bad
a corpse of a platform
If you consider scraping a threat then yes plain HTML might as well be giving up. The advantage of new reddit for that is quite clear: they can collect a bunch of data about your browser before deciding if you are a bot and if the rest of the page should load. The embedded recaptcha call in the screenshots is a pretty good hint. I suspect blocking trackers on new reddit will break as soon as the scrapers move over.
As for why they don’t just kill old reddit: a significant chunk of their active posters use it and are attached to it. So if they kill it entirely they will lose content. Posters are of course logged in so this change is less likely to affect them.
if you are a bot
Or they’re bouncing banned people.
funny, that comes just on the heels of my deciding that reddit is unsafe.
You’re just now realizing it?
I now browse Wikipedia. Please don’t screw me over Wikipedia, I donated five bucks to one of your nags once.
For anyone who doesn’t know, you can download Wikipedia and host it yourself! I got the top 50k version (~7G) on my RPI3 and now no matter what fuckery they pull or the government pulls, I’ve got a pretty decent source of general information.
How do you do that?
I am currently pretty tapped out on all storage and backup drives but if I had space this post would have motivated me. Just sayin
Could you share what you did to achieve this? I’ve been planning on doing just that, and have it auto-update every week or so (keeping the previous versions archived, of course) by using kiwix-serve for a static ‘.zim’ file and maybe a cron job for the auto-update. But if you have a better solution, I’d love to know. The deployment I am planning is kind of convoluted to be honest.
No, that’s exactly what I did, I’m running kiwix-serve, but I’m not going to bother updating it because I’m really worried about information degrading now that fascists are basically calling the shots on everything (and WP’s jackboot co-founder has a hard on for it). If I feel enough time has gone by to warrant an update I’ll just do it manually.
Whats reddit?
typo, it’s spelled “read it”.
The entire fucking point of the Web was to make information as easily-accessible as possible, structured and semantically tagged, and consumable by humans and further machine transformation alike. “Scraping” is facilitated by design!
Using Javascript to deliberately break that is evil and every programmer who participates it is a piece of shit. No exceptions.
That’s true but they probably didn’t account for AI data scrapers ramfucking your server so they could steal all the value you assembled for general consumption and serve it themselves for profit.
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API. The “ramfucking” is caused by the attempt to block bots; it is entirely self-inflicted.
Remember, it’s all our content to begin with and Reddit does not have any right to try to lock it up for itself.
That doesn’t mean I like all the AI bullshit going on, BTW. But the problem is the generation of the slop, not the data accessibility.
Well they had an API, but…
The scraping wouldn’t be a problem if Reddit simply provided an RSS feed or other data-efficient API
That’s simply not true. These bots are essentially DDOSing the entire internet, API or not.
Okay, if efficient APIs existed and they weren’t incompetently failing to use them, it wouldn’t be a problem. Happy now?
(I should’ve addressed that in my previous comment, as I was aware of how one of the Lemmy instances was taken down by scrapers the other day despite the fact that they could easily get all the content simply by consuming ActivityPub directly. But I was naively hoping it wouldn’t be necessary because, as you can see from this text, it would’ve cluttered up my writing with double the words.)
Reddit does have RSS feeds
But we all know AI companies are unethically scraping and selling shit back to us right? I just really need people to acknowledge that.
It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”
I think it would be a problem because the scrapers are hammering all types of websites from small forums to reddit with tens of thousands of unique ip addresses at a time. Websites that have neither the money, hardware, or protection had to figure out solutions really quick or suffer what is essentially a constant ddos attack. This is the reality of the web now, it’s just an incredibly hostile place.
even logged in I can’t access old reddit anymore.
old.reddit redirect addon just got a update that deals with that problem.
https://addons.mozilla.org/en-GB/firefox/addon/old-reddit-redirect/
if you use the reddit enhanced suite extension you can get redirected
If I wasn’t permabanned I might bother trying work arounds.
I just checked my old.reddit login on the desktop… still working just fine today. I don’t think you need work-arounds, you just need to access it without any work-arounds.
You think that if I do the thing that’s not working that it will somehow work.
working for me
How dare you question King Steven the Turd, Greediest of Pigboys? If he proclaims HTML to be unsafe, it must be so. That’s a King’s job, to tell the Landed Gentry how to behave…
I’m not forgetting or forgiving that shit either.
I get blocked and asked to login to reddit no matter if it’s old or new. safereddit still works though so I’m using that for now.
afaik supposedly it was because a large chunk of the bot network went through the old.reddit portal over the standard reddit portal.
That’s just an excuse by reddit. Moving to the new style allows them to choke down and control the way that posts and replies are displayed and nested. This is good for them, because it allows them to offer white glove PR services to paying customers. It also obfuscates useful user supplied content so that it can be sold wholesale to anyone who has the money to buy it. That’s more important to them than offering a good user experience and useful website to the proles.
Anyone still posting on reddit (who isn’t a bot) is working for free for an unscrupulous company.
Only because the bots were already set up to do it that way. It’s quicker to use what exists when it works.










