Hacker News new | ask | show | jobs
by adamddev1 5 days ago
> In addition, Codeberg conflates “having a community” with “being legitimate software worth hosting”

I don't think this is a fair summary. I don't see this conflation in Codeberg's blog post. The words in quotation marks don't actually exist in Codeberg's blog post.

It seems that what they are saying is: In a free hosting platform where we share all the resources, it's not fair that some people using LLMs are using up disproportionately huge hosting resources to handle the unnaturally large amounts of resources they can churn out with these tools.

I'm a solo dev, and I couldn't find the judgement against solo developers or any slight against people who don't have a team. The concern is that some people don't gobble up an unfair amount of resources.

I personally have been moving to self-hosted Foregejo for ultimate freedom. But I also think it's fair for some providers have the freedom to take a bit more of a harsh (or based - depending on your perspective) stance against LLM generated code. I mean, a lot of people have been bemoaning the fact that the human-generated material has been increasingly hard to find in a sea of AI-generated material. It could be nice to have a corner of curated, human-made code still available on the web.

10 comments

> It seems that what they are saying is: In a free hosting platform where we share all the resources, it's not fair that some people using LLMs are using up disproportionately huge hosting resources to handle the unnaturally large amounts of resources they can churn out with these tools.

Then instead of addressing the excessive resource usage problem indirectly via banning AI-assisted apps, they could've addressed it directly by imposing limits/quotas.

It's abundantly clear (and not that they are hiding it either) that their issue is with AI itself so anything else is just a secondary concern if not even a smokescreen.

---

> I'm a solo dev, and I couldn't find the judgement against solo developers or any slight against people who don't have a team.

Quoting their blogpost [0]:

> ## The development team of none

> Using LLMs to work with your code gives you a kick of adrenaline. You can develop at a rapid pace, build things as if you had a large team. Only that you have none. In fact, you are (often) alone, working with a statistical machine that turns energy into code.

I couldn't care less about what they think about solo developers, but it's just one bad excuse after another and that's what people are criticising. Open Source is one person. [1]

---

[0] https://blog.codeberg.org/protecting-our-floss-commons-from-...

[1] https://opensourcesecurity.io/2025/08-oss-one-person/

> they could've addressed it directly by imposing limits/quotas

For one, they've already done this to an extent. For two, they're also doing this by deciding what projects are and are not welcome on their service in the first place.

If 9 people eat an apple or two a day but a tenth person wants to eat 100 apples a day, we don't have to set some apple limit. We can tell the tenth person to pound sand and keep giving everyone else the apples they want.

Having a hard quota means each spammer will fill up the quota, then get some error they don't know how to fix and stop. This means the quota has to be low.

Just saying "no spam" means you don't have to waste a single byte on spam and everyone else can have a higher quota.

Resources were never part of the original discussion. It was about values. This is the author of the ToU addition the day it landed: https://mastodon.social/@gedankenstuecke@scholar.social/1169...

He also has a lot of posts since: https://mastodon.social/@gedankenstuecke@scholar.social

If you can’t keep up with the load, say that. Limit new signups and/or implement rate limiting, and ideally scale up if you can.

DO NOT try to spin that into taking some form of moral high ground. DO NOT make yourself the arbiter of what is and is not considered a worthwhile project. That is not neutral. That is not free.

They're not obligated to deal with the higher load brought on by people's LLM coded projects, CI pipelines etc. In order to keep things freed up and flexible for others, I think it's fair for them to say, "in this community, we're going to save the space for more hand-coded projects."

It's not fair for everyone to have to eat the costs of AI (like communities getting bulldozed for data centers) and just accept that it has to be used.

Of course nobody is obligated and I never claimed that. In fact, I fucking hate the LLM hypetrain. But this decision is decidedly non-neutral and non-free.
Fair point, but I mean perhaps it's not fair for them to have to impose usage limits (which would limit everybody) when they could just make a decision to avoid certain types of projects (which they feel strongly about.)
What would a neutral free SaaS look like to you?
Simple: Provide the same freedoms to everybody. Do not get involved unless absolutely required to by law. That entails not restricting access for specific projects, and not revoking freedoms for any reason. Everybody should get the same treatment.

For example, in the case of being overloaded by AI contributions, a neutral solution would be to implement the same rate-limiting for every project. Set a sensible limit, and enforce it universally. Do not discriminate specific projects.

Why should GNOME and Forgejo, which are great projects, have to suffer the same limitations as a blob of non-working output from a Markov chain?

It is like saying Gmail shouldn't delete spam, but instead, should apply a limit of 1 message per day per destination per source address to all messages.

Most vibe projects aren't even projects in the way that spam isn't even mail. The message says it's from my bank but it isn't. Why should I have the same limits on messages that say they're from my bank and messages that are actually from my bank?

Agree. Providers who don't want to deal with LLM generated code should be respected as part of the free market. Let the market decide which direction is for the better or at least allow for niches that cater to certain audiences.
It's wild to me that nobody is considering enforcement. Part of the reason this is a bad rule is because there's no way to actually validate what is an LLM coded project. And the AI train is not nearly at its destination. Things are going to change dramatically over the next four years. My prediction? This is not going to age well.
It'll be the same as literally every other ToS ever: if they detect a problem and trace it to you, and you're blatantly violating the terms, they kick you off. If they detect a problem and trace it to you, and you're subtly violating the terms, they lock out whatever is causing the problem and get in contact with you.
You're trying to sidestep how enforceable something is, and instead implicitly admit that enforcement will be inconsistent. That's my issue. Not all ToS are this way, and they're better for it.
Perfect enforcement is called tyranny. Anything that matters isn't enforced perfectly.
On the contrary? Selective enforcement is tyranny. I'm sure the projects that they disagree with politically are far more likely to be "AI generated" than the projects that they agree with politically.
That is fine unless you claim to be a neutral and free as in freedom service provider. You are clearly not at that point.
Define resources and create a better limit then. 10 small code-only projects will take orders of magnitude less storage, bandwidth and cpu than one project with committed image and sound assets.
It's a community code forge. If someone poured their heart and soul into making an awesome game with kick-ass image and sound assets, the community probably wants to support it.

Their stance is obvious: if you just want dumb git hosting with a freemium model, use something else.

The resource limit are quite clear: From each according to their ability, to each according to their needs.
And as long as it's hand-made, that's fine.
What is hand made code? Does that mean no auto complete? No linters?
If a carpenter uses a table saw, is he still a carpenter?

What if he just throws wood randomly at the table saw and tries to sell whatever pieces come out?

(Never do this with a table saw, btw. They can, and will if you don't use proper procedures, throw wood sideways at speeds that can punch through walls and impale people)

This analogy doesn't make much sense to me. Is the table saw analogous to LLMs because they both work with the appropriate substrate (wood/text)? I don't see how an LLM is like a table saw at all.

Is a musician less of a musician if they use a pre-recorded back-beat and sing on top? Are you not a musician if you use the chord mode on a keyboard?

Is code not hand-written if my editor automatically closes all of my html? What if my editor flags all of the errors and I manually correct them? And what if it auto-corrects at the press of a button?

There is a line and I'm asking about it earnestly, but the table saw analogy doesn't reach it IMO.

What if I barely know the rules of chess and I use a chess computer to play all my games for me; I win the world title this way. Am I a chess Grant Master?
I think you already know the answers to all of these questions.
This is an excellent analogy that shows how LLMs are fundamentally different than linters, auto-complete, or transpilers.
People downvote this kind of take because you're striking at the heart of the enforcement problem and also the definition of what a project that is majority built with AI really means. What they really want is to ascertain the amount of effort from a human that went into the project and they're using the percentage of code generated by an AI as a proxy metric. But it's a bad proxy, both because it doesn't capture human effort and because it is not measurable.
Even human effort is a terrible proxy because effort is a gradient that depends on skill. An expert exerts low effort for the same result a novice would labor over.
Maybe it shouldn't be viewed as some kind of reward but merely as a spam reduction measure. If you put in effort, it isn't spam.
That's a good point. I think the core metric is somehow quality, but that's notoriously difficult to measure.
> People downvote this kind of take because you're striking at the heart of the enforcement problem

That's not your problem.

What conversation are you trying to have? I don't even use Codeberg; I've spent years thinking about how the idea behind rules manifests when those rules are actually put in place and enforced. I think a lot about second and third order effects. That's my interest here: an org decided to put these particular rules in place, and based on my experience, I don't think it's going to age well.

But I guess you're right, as far as you go: none of this is my problem. I'm just a curious observer.

> people using LLMs are using up disproportionately huge hosting resources

So block huge resource use? What if people use LLMs with disproportionately tiny resource use?

Making a principled ban against LLM would seem to allow for more freedom and flexibility with usage limits. People who are creating more by hand would still be allowed to have more breathing room. With a hard limit of resources everyone would be more restricted. That's also a solution, but it's just a totally different approach. The LLM ban could help to allow for freedom without the resources getting overwhelmed.
Only if you can reasonably separate LLM from non LLM code, which is in no way a solved problem. Even if one agrees with their priorities 100%, it's just asking for non-automated enforcement. It's difficult to execute, and either lead to completely arbitrary enforcement, or a lot of resources being spent in exchange for this freedom you mention. It's just a very hard practical line to try to draw.

And let's be real, most of the costs LLMs incur on hosting providers have little to so ln who makes commits. A popular project than bans LLMs from contributing, but is used often by LLMs that are using the project as a library is indistinguishable from one that allows LLM contributors.

So from a practical perspective, it's a head scratcher

It's true that an LLM ban is really hard to enforce, but that doesn't mean it's worthless. If i make an image editor and say "It is my policy that serial killers and pedophiles cannot use my software" i can't really enforce that, but it still makes a (somewhat useless) statement.

In my opinion, not allowing LLM code at least signals that codeberg (and forgejo) itself is unlikely to be majority LLM code, which raises its stock in my book; It signals what kind of people the maintainers are.

> what kind of people the maintainers are.

Can you elaborate on this? It's new that we're judging what kind of people we're dealing with based on what tools they use.

I really meant -- are these maintainers passionate about the design and maintenance of their tools, are they trying to slop up a demo fast enough that some VC buys them out, is this project a goal in itself or the means to some other end, etc...

Open source software that exists for its own sake is exceedingly rare, and understanding the goals of a community before becoming dependent on it is a good idea

If you made this statement and you catch a serial killer using your software, it affects the options. A more reasonable example may be: Microsoft may not reverse engineer this. They still can do it privately - you can't enforde it - but if you catch them, then you have more options besides saying "welp, I said they could"
> Only if you can reasonably separate LLM from non LLM code, which is in no way a solved problem.

It's trivial to detect when someone's coding with a chatbot - they tell you, repeatedly.

Nobody who doesn't code with a chatbot worries that it's hard to tell.

Apparently some people have compared it to some other unlikable groups of people (I know this because a vibecoder complained about the comparison someone else made)

If someone has to ask whether their software is mostly AI-written, then it is.

If someone has to ask exactly how racist you have to be before you can be called a nazi, they're probably a nazi.

If someone has to ask how attracted to underage girls they can be before they're a pedophile, they're a pedophile.

If someone has to ask how responsible you have to be for a death before it's a murder, they're probably a murderer.

Because people who aren't these things clearly know that they aren't.

> principled ban against LLM

So was it about high resource use, or a principle?

> People who are creating more by hand would still be allowed to have more breathing room.

What about the breathing of people creating by hand with a sprinkle of an LLM?

> but it's just a totally different approach.

Indeed, a better one where you manage directly the issue you claim to care about

My interpretation is that the resource usage should be roughly proportional to the human activity involved. That wouldn’t restrict resource usage by any individual project, which might be a popular project with many contributors and a lot of activity, but the resource consumption needs to be justified by corresponding human effort and interest, which LLM usage doesn’t exhibit.
https://blog.codeberg.org/protecting-our-floss-commons-from-...

In this corresponding blog post, one point made is that LLMs tend to LARP big-project infrastructure on small projects. They use the resources for something like GNOME but they aren't actually GNOME so they don't have any excuse for it.

Both. They noticed the high resource usage of LLMs and decided that LLMs, in general, aren't worth supporting.
The main problem is that so many LLM users are selfish bastards and they won't respect the spirit of such rules and will only stop when the hard disk quotas force them to.

So now your solution entails lowering the quota for everybody because allowing in LLyou must assume eve

How does adding a few words to TOS stop "selfish bastards"?
Gives them a line to justify removing projects in the future if they are caught violating the ToS.
Do we have a tool that can detect whether a project is majority LLM-written? This is an unenforceable rule, unless it's just coming down to vibes and they're going to decide what projects to ban based on their feelings.
They are monitoring their resource usage. If they find a resource leak and it's traced to an LLM project, they will delete it.

Terms of use should be read as a statement of intent, not a series of if-then statements.

> It is an unenforceable rule

Are you sure? It's pretty easy to find telltale signs of LLM usage, like the CLAUDE.md file, the '.cursor' directory, the 'vim'-sized codebase done in a week.

I think you are wrong. A better way is to have a free tier and charge for usage.
Codeberg is a registered non-profit charity whose purpose is to provide free services for free software projects. You can make a different site that works the way you want it to.
Yeah, I really wish all the people whining about Codeberg would give us all pointers to free alternatives that don't have these restrictions.

And, if those alternatives don't exist, well, you now understand why Codeberg is taking the position they are.

Yeah, the natural solution here is to only offer some amount of hosting for free and then charge for larger amounts of code to host. Those quotas are something that is empirical and observable and can be managed. But instead, they're implementing a rule that is unenforceable because you can't actually validate whether or not the majority of the project was written by an LLM. They just chose a suboptimal path and pissed a bunch of people off in the process, and their terms of service makes this out like some kind of copyright or security issue. It's just silly.
Terms of use are a statement of intent, not an if-else ladder to be executed by a computer.
A more relevant quote is "to enable equal opportunities regarding the access to knowledge and education." from their mission statement. As a gemeinnütziger Verein (charitable association) their charta is legally binding.
All animals are equal, but some animals are more equal than others...
The needs of the many outweigh the needs of the spicy autocomplete.
well one can argue that many more people will write code with AI assistance than without it.. here i see the opposite: The wish of few outweigh the need of many.
Can one argue with a straight face that people will learn to code better by using a predominantly AI-based workflow? Or even learn by reading vibecode uploaded by someone else?

Hell, I'd argue that Vibecode without the prompt that generated it isn't even real open source.

Can you argue that you could learn to code better, copying and paste code from stack overflow? Who wants to learn, learns.. we did without Internet, and I still learn with AI.. btw, is in the "Satzung" from Codeberg e.V something about "learn to code?