Hacker News new | ask | show | jobs
by dlenski 10 days ago
> I think at this point I'm ready to give up on Reddit, most of my questions are already answered by LLM.

That's a bizarre and self-sabotaging path to take. Almost all LLMs get a substantial portion of their knowledge from Reddit. In order to evaluate the quality of those LLM responses, you'll have to click through to the Reddit threads and critically evaluate the quality of the underlying discussions.

If you just blindly accept LLM responses without evaluating the sources, you'll be accepting a lot of garbage, and it'll show.

> And the discussions quality on Reddit are abysmal, its filled with bots.

I spend quite a lot of time reading and writing on Reddit, and while I encounter plenty of low-quality posts and comments, I'd say that essentially none of them are written by bots.

6 comments

>I spend quite a lot of time reading and writing on Reddit, and while I encounter plenty of low-quality posts and comments, I'd say that essentially none of them are written by bots.

Err... I spend quite a lot of time on Reddit as well (including moderating a medium size subreddit) and I can assure you that there are tons of bot posts and comments all over the site right now. If you think this, you're probably not very good at identifying this kind of content.

It sounds like maybe your moderation is working!

I spend most of my time on subreddits related to the city where I live, and tax+financial planning forums.

I'm a moderator on a moderately popular and active subreddit, and I can assure you that there is an unreasonable amount of bot activity on reddit. They range from just spam (usually obvious) to bulk astroturfing. We try to stop what we can, but plenty gets through.
Any ideas to mitigate?
Inside of Reddit itself? Not with the tools we're given from Reddit. We know Reddit has internal tools to block spam/bots, so I believe what we see is the most sophisticated ones (they're not too sophisticated mostly, but they're not the zero effort bulk spam that reddit catches easily).

Something that would make things far easier for mods would be if the posters hashed IP was shown/made available to mods. Thats mostly for more localized bad actors/astroturfers and not large spam/scam rings with proxy networks.

Things we've done that have made a difference with the tools we've got from Reddit:

1) we don't allow new accounts or ones with negative accounts to post in our subreddit

2) any posts that match criteria or hit the front page are automatically put in a mode where you must have a positive subreddit karma over a threshold to post. This helps filter out brigading and low effort shit posting from front page plebs

3) we're using a reddit app that allows us to remove posts from people that have posted in identified subreddits. Right now we're just experimenting with this and we're in a data collection mode, but we can see that the people its catching (while some of our most prolific posters) are also some of our most problematic posters, so I think we're going to go from data collection to just straight up having the bot remove comments from people that post in those troublesome subreddits.

I'm curious what you've found the troublesome subreddits to be. I think I have a pretty strong idea what they are, but some validation would be nice.
I'm not comfortable disclosing them because we don't want our 'secret sauce' to be leaked and let people know we're on to their BS.

I will tell you that it's mostly the ones you expect if you're heavily engaged with Reddit and moderation. But there are several new gathering places that I didn't know about until we started spelunking on the most annoying posters histories.

This is one of the places where a tool that scrapes reddit and lets you see deleted comments is absolutely clutch--these people are very good about removing the worst of their vile spewing on reddit, or changing their comments to look innocuous after people have gotten worked up. It's all about making people they disagree with look like bullies and idiots.

> remove posts from people that have posted in identified subreddits

So you're actually censoring opinions, not fighting spam, just like every other reddit mod.

We're maintaining a space where people with actual skin in the game can have conversations about hot button topics. We do not maintain the subreddit for large numbers of outside agitators to parachute in, shout down local discourse, and then fuck off immediately.

There is room for people on any side of an argument to come in and discuss. They can do so in a civilized, adult manner, or they can fuck right off. Call it censorship if you like.

> There is room for people on any side of an argument to come in and discuss.

Obviously not, when you ban them immediately just for participating in another subreddit, regardless of what they wrote there.

Call your buddy and come up with a meeting spot.
This is my Hail Mary hope with LLMs. They're obviously going to get to the point that any sort of reasonable validation of a person online becomes impossible. LLMs could ultimately destroy online chatter. I like shitposting as much as the next guy, and probably a good deal more, but in reality the sort of discourse you get online is so mediocre relative to real life. It's just somehow quite addicting for whatever reason.
It’s just crazy to think that Reddit and the massive volume of text it allowed to be created helped lead us here….
No new valuable training data has been created on Reddit since 2023.
What's your source for this claim?

I've read new, valuable-to-me human-written information on Reddit since 2023, so I very much doubt this is true.

You can ask gpt/claude to verify multiple sources online and never to guess. You can even set up a project that always does that with anything you ask it. It's probably better as reddit can be very opinionated.
I agree with them. In many of the subs I used to frequent, the outrage is high and the signal is low. LLM really is the best way to consume the content there. Retrieve the new posts of the day/week, filter out the outrage and politics and low-effort Reddit in-jokes/copypasta, then summarize the various comment threads. My time is too valuable to sift through the utterly mind-numbing nonsense to get the occasional new piece of info. If you want to dig deeper, ask the LLM to dig in. Or have it regurgitate verbatim a pre-filtered selection of comments.

Reddit is a good source of new happenings and discussions, but I’m willing to spend some tech company’s compute to clean it up for me. For example, I often want to branch out from a particular musical artist. Reddit has many good recommendations. LLM is perfectly capable of gathering these.

To many bots on reddit, so your solution is to go directly to the bot!!?

If bots bury the signal in noise then an LLM is only going go back and give you the summary of the noise.

The bots (and humans) I’m trying to weed out are the pot-stirring permanently outraged/bitter commenters. For example on a “car overturned on X road” post, people getting into fights about whether BMW or Subaru or Tesla drivers are worse. That’s all noise. I want to know when it happened, which direction of traffic, if it’s cleared up yet, etc. LLM works well to remove the off-topic garbage. If someone is botting the actual content (recommendations, reviews, etc) and the up/downvoters are unable to identify it as a bot, I probably won’t be able to either. So I’m not worse off using the LLM.

Worst case I get some pre-digested garbage and I waste less time, best case the LLM pulled out a few good nuggets of info. It’s not a perfect system yet. There’s a lot to tune with how deep to go, what constitutes garbage vs signal, and so on. At least it’s fun :)