Hacker News new | ask | show | jobs
by dvt 1196 days ago
> So I can easily imagine a near future where the web will be flooded by LLM output or at least by content heavily inspired or edited by LLMs.

To be fair, we're already there, and we've been there for at least 10 years now. I'd wager >75% of the internet is garbage: auto-generated blog posts, programmatically-permuted ads, YouTube videos that mainly regurgitate other sources. Email is mostly garbage and the only reason it's usable is because spam filters have gotten pretty good. Even non-trivial amounts of heavily-curated social media (Twitter/FB/IG) is purely spam.

4 comments

Can you give an example of any significant scale of fully automatically published blog posts from 10 years ago? As far as I know, most of these crappy articles were content farms, often using templates and outsourced labor, but not automatically generated content.
There's a vast amount of automatically translated websites, which IMO fall into this category.

Some product comparison websites also seem to be built based on automatic sentence generation from tables with specs.

It's not the same as buzzfeed style content farming but it was a sign of the things to come.

"Can you give an example of any significant scale of fully automatically published blog posts from 10 years ago?"

For some reasons I never bookmarked those sites, when I left them in a rage and disgusted about so much information garbage. So it definitely has become way worse, but also 10 years ago I remember that pattern. Most often when looking for alternatives of software, then you were always a click away of being on a nonsense site, automatically filled with all the relevant keywords and lots of things to accidently click on, but nothing useful. Some of it might have been manual edited, but for the most part, I could allmost see the algorithms that filled those sites up with "content". There were just really primitive - so things will get interesting when this will gets combined with LLMs big scale.

Stockmarket related content
Spam is already easily generated so AI won't change that.

Misinformation and manipulation is based on a small number of posts being shared and upvoted en masse, so AI won't help there.

However, social hacking and fraud involving actual dialogues with people is currently labour intensive and low yield. AI will definitely enable more of those attacks to happen automatically; and conversely, also help anti-fraud companies create puppet accounts to waste the fraudster's time, and thus the game of cat and mouse continues.

The key difference is that LLMs will allow an unforeseen degree of interactive and custom spam that will feel less and less like spam and more and more like a natural and even useful interaction, until you just realize it was an elaborate scheme to subtly increase your preference for brand A over brand B
Or it will still end up in a bin called statistical spam filter.

It's relatively easy to detect when someone is trying to sell you something. Does not matter if it's a custom written text or not. We had the true spam filters going 2 nines accurate already. (Google's is somehow misgauging spam or not learning my particular mix of spam well.)

LLMs can also be used to identify spam, not by language, but by the actual intent and "is this email something I want to read".

Open question: is there any case where LLMs can be used for malicious purposes, but LLMs can't be used to defend against it?

> LLMs can also be used to identify spam

It will be fun to watch the arms race where the spam generator need to conceal prompt injection attacks meant to circumvent such filters while at the same time be too subtle for a humans to pick up

i agree, but this stuff could very likely be a huge force multiplier.
Yeah, I remember reading about things like sports reports and weather being generated by computers ages ago in the likes of SciAm or New Scientist. I don't recall if they used the term AI, I think they did but this was a long time ago.
A couple of financial reporting sites use machines to write articles on small cap stock ticker quarterly reports. It ends up being pretty generic but occasionally is nice to have at a glance human readable summary for when random tiny biotech is suddenly in the news out of nowhere.
I mean, sports reports are on the same level as airport announcements. There is really no need for a person to waste their life away just reporting plain boring numbers.