Hacker News new | ask | show | jobs
by tomhow 17 hours ago
It's one thing for us to detect and autokill generated comments on our own site; it's our site and we can set the rules and run software on our own servers to process the comments and handle things in the way we and the community are happy with.

It's a big additional leap to for our software to try to reach into others' sites, get through any anti-bot defenses they may be running, try to scrape their content and evaluate on whether it's sufficiently human-authored to be on HN.

There's generally a wider range of LLM involvement with a long-form post than the typical, relatively brief HN comment, which then opens the way for more debate on HN about "how much" LLM influence the post has and how much should be allowed on HN. Part of what we're trying to optimize for on HN is minimizing offtopic/meta discussion, so we don't want to encourage this kind of debate.

Our heuristic about article quality is largely unchanged from before LLMs were an issue: if an article is badly written, it shouldn't be on HN, and should be flagged.

2 comments

You should talk with dang, as the email I received about this from him, to my eyes, does not agree with your stance here in the long term. I don’t want to get into the interminable blood quantum debate over Llm authorship either, however, I see substack doing something about it, and, as a long time hn reader, it makes me want to spend more time over there than here.
We're talking about it all the time :) My comment above doesn't contradict the email. I didn't say we're not wanting/planning to do anything about it, just that there's more to it than plugging in Pangram. If Substack is being more proactive about it on their own site, that's great. It would make life easier for all of us if all the major content platforms cleaned up their own sites. We're already proactive about detecting/autokilling genai comments posted to our own site. It's detecting genai content on 3rd party sites that introduces more complications.

In the meantime, please feel free to flag items that are badly written/unpleasant to read, and email us if something is on the front page that shouldn't be there.

Looking forward to seeing how you handle it! I understand the complications of it as an engineer, I hope you figure out something soon!
There are AI-detection tools which exist. Even if those won't run against all URLs, using them where possible should provide some utility. HN already penalises sites based on various criteria, and if it takes hand-pasting some examples from a site to find that it is/isn't using AI slop, that's another option.

As to what should be tested: front-page items, possibly even a subset of those (top 10--15 of 30). That's going to be a limited set of items per day, though more than just 30. (I don't know how many items cycle through the front page on a daily basis, though I believe daily submissions as of 2022 were about 1,000/day (<https://web.archive.org/web/20220116193045/https://whaly.io/...>)).

Working this into the HN story-processing lifecycle might be a good call.

I'd much prefer not seeing a bunch of AI slop in submissions, by way of generated output. AI as part of the resarch process I think I could live with.

AI-generated content seems, definitionally, not to be intellectual in nature, and would seem to go against HN's prime directive. It also seems to make HN lose its collective mind, which has long been another mod consideration.