Hacker News new | ask | show | jobs
by nneonneo 2 days ago
At this point, in the field of AI research:

- papers are written by AI (as pointed out in this article, and as obvious to anyone who spends a while actually reading recent AI research)

- papers are reviewed by AI (NeurIPS is doing an AI assisted review experiment - https://neurips.cc/Conferences/2026/ai-reviewing-experiment - and I feel the trend is moving towards AI reviewers whether we like it or not)

- papers are read, summarized and digested by AI, because there are just so many papers at leading AI conferences that nobody has time to eyeball them all

We are very rapidly automating humans out of the academic publication loop here.

16 comments

Idk if we're automating humans out of the publishing loop as much as rapidly automating the production of crap. I had a very similar experience reviewing for EMNLP recently.

We are nowhere near AI being able to judge the quality of research (in fact, one might reasonably state that even most humans can't really judge the quality of research). Most things in society are not like math: we can't automate (via verification) our way out of noise overwhelming the signal.

Folks are willing to entirely abuse the public resource that is faithful, honest reviewing. (This is unsurprising; the abuse of the commons / public resources has been rising for a long time). There isn't a good solution other than something akin to draconian social scoring to limit access to the reviewing system.

> draconian social scoring

This is a legitimate question: should people have reputations? Should their behavior be made more visible publicly, both good and bad? How?

In a small contained society, where consequences are more directly affecting individuals immedately, these questions don't need to be asked, because they're inherantly answered. We now are a society with billions of people, and dire consequences sometimes deferred for a generation, or more. Part of our general failing is the lack of good answers to the above questions. For many people, there are rarely negative consequences for causing harm to others, and the rewards can be very great indeed.

"Consequences and reputations stemming from one's actions" isn't necessarily draconian social scoring. Without a structure for imposing consequences on wrongdoers, we're not a society, we're just monkeys flinging poo at each other.
The problem is that everyone imagines some sort of just and fair arbiter of these things, when the reality is all of the social scoring and consequences and reputations rarely actually stem from one’s actions, and far more often stem from how much money someone has put in someone’s pocket, who someone knows, the color of someone’s skin, or what’s between someone’s legs.

Until we’re actually serious about treating people equitably, (not equally, as that would simply leave the lopsided power structure we have in place) we aren’t getting out of this.

I don't understand the desire to argue against this in identity politics terms, which I assume you realize is going to be extremely divisive. There's a far more simple argument that people, independent of ideological framing, would tend to agree upon. And that's what is valuable to one person isn't valuable, to the same degree, to another. For things that are completely illegal, like murder, such a system would be pretty useless because we already have systems in place to punish that sort of behavior and defacto social scoring on it as well.

So this is going to come down to things that aren't usually illegal or 'that illegal', but otherwise affect society. So for instance one group might want to reasonably punish the directors of a company that causes emissions. Another group might want to reasonable reward the directors of a company that creates a large number of desirable jobs. And even in this one example you immediately end up with a weird scenario. Is it okay to pollute as long as you make enough jobs? Who gets to decide that?

There's no need for cartoon villainy for this to be a bad idea. Even completely well intentioned, it just doesn't work out well. And the more diverse a society, the worse it's going to work.

How do you combat review bombing or someone just being bigoted?
> social scoring and consequences and reputations rarely actually stem from one’s actions

In my experience, this viewpoint is most strong in people whose actions are the most direct cause of their social issues. Almost always these people strongly refuse to admit that their actions should or are under their control. At which point others start to avoid them.

There are large scale social stratifications due to various things however individuals live in a very wide band in my experience. Some make choices that put them at the top of that band and others make choices that put them at the bottom.

Let's take two of your examples:

> how much money someone has put in someone’s pocket

> who someone knows

Careers choices directly lead to having higher income and more money. Socializing and putting effort into building a social network gives you more people to know. Both are to a decent degree under someone's control and based on their actions. You claiming otherwise is exactly proving my point on people wanting an excuse for their poor historical actions.

> The problem is that everyone imagines ...

I'd phrase it "most people want to imagine". Or will claim they want to - since not believing sounds depressing, or suggests that the person is an evildoer hoping to escape justice. (But unfortunately, people usually disagree about exactly what would be "just" or "fair". While wanting to imagine that they don't. Yes, the problem just went meta.)

> when the reality is ...

Try asking some really old folks about how much drama, inequity, and nastiness there often was in social settings where everyone was the same color and gender, nobody was notably wealthy, and nobody had any great connections. Or talk to an experienced junior high teacher. Or read some history. Humans are quite capable of dividing themselves into camps over any "differences" that they're able to perceive. Or invent, since dividing themselves into camps is often the unspoken objective.

For a more mathematically rigorous treatment, negative feedback is fundamental to Control Theory; it allows us to stabilise processes that would otherwise go out of control.
Even monkeys have society and they tend to only throw poo at those who have done wrong by that society. Monkeys have and try to maintain reputations.
Is it reputation or hierarchical standing or just "power" within the society what makes them not throw poo tho
I'm not sure what you mean, but not throwing poo is the default behavior. They only do it when feeling stressed, or zoo settings they might even do it to get a reaction from the crowd. Figuratively speaking, no different than humans.
People react negatively to this, because they fear a dystopian society of the sort we've seen in plenty of movies, and rightfully so.

But it's also worth pointing out that "consequences and reputations stemming from one's actions" is already the world we live in and always have lived in. Hell, even Hacker News has karma points, downvoting, shadow banning, and the like. There's no such thing as a society with zero consequences and zero reputations. The only real question is a matter of degree, structure, severity, reach, and various idiosyncrasies that differ across cultures.

So it would be nice to have a more nuanced discussion about this instead of treating it like a 0 or 1 decision.

this reminds me of an episode of the TV show. the Orville that I recently watched where they stumbled upon a planet that had the social credit scoring system and everybody had little badges with literal upvote and down vote buttons on them. and when your grandma buys you ice cream and you're a little kid you say thank you and you hit the upvote button and she likes you even more because you hit the upvote button then because you said thank you.

and so it was obviously portrayed as kind of you know for the story and it worked and I like the episode and everything but it got me thinking in terms of a society like that of people could probably construct something that used that kind of a system but in a more regulated way and a more humane way and I think that's turns into the question of do we want to try and make this new thing or do we want to try and fix the old thing?

regardless, it got me thinking about some kind of a potential implementation of a formal social reputation system that essentially mirrors what goes on informally now, but makes it visible and makes the actions of affecting somebody's reputation visible.

My basic intuition here and for most things that deal with society or whatever is that Satan lives in the details and infinite potential dystopias and the potential for extreme inhumanity lurks in every crevice and corner. unknown unknowns and shit.

> regardless, it got me thinking about some kind of a potential implementation of a formal social reputation system that essentially mirrors what goes on informally now, but makes it visible and makes the actions of affecting somebody's reputation visible.

In general you can already get this by asking and taking feedback from others. I've seen a lot more people shoot or ignore the messenger in those cases then actually take the feedback. Everyone is the hero of their own story.

The Orville is excellent. I am a fan. Here is another show/episode that is very similar and relevant here : https://en.wikipedia.org/wiki/Nosedive_(Black_Mirror)
>Hell, even Hacker News has karma points, downvoting, shadow banning, and the like.

Internet Right/WrongThink scoring should be an argument against any sort of formalized or standardized social credit, not for it.

and then you have to ask who defines good and bad behaviour and who determines consequences when power is lop sided.
There's a huge difference between a completely decentralized reputation system like the interpersonal one that has existed since humans first became sentient, and an artificially built top-down reputation system run by a powerful entity (be it the state like in China or megacorps like in the West).
Isn’t China doing something like this ?

Wonder if it could degrade to something worse than the status quo where reputations are falsely tarnished by competitors purely as a weapon to get ahead…

The "Chinese social credit" system is vastly overblown.

I'm living here (China, Beijing and Jiangsu) now, and there are exactly two cases in the past eight months where the "social credit" system has had any impact at all:

- To open "take first pay after" vending machines. These are vending machines which are basically big locked fridges; if your credit score is high enough, you can open the machines, take whatever drinks/snacks you want, and they'll use (I presume) computer vision to charge you afterwards. If you don't have a high enough credit score, tough, you can't use them. (But there's almost always normal pay-first vending machines nearby).

- To borrow mobile charging packs (powerbanks). Some operators will let you "swipe your credit" (check your credit score) to take one without paying a deposit. If your credit score isn't high enough, you first pay e.g. a ¥99 deposit (~$15USD) which gets returned when you return the powerbank, not a big deal.

That's...it. My credit score is high enough on only one platform (Alipay), so I get to try what happens with both "high" and "low" credit, and I can confidently say these are the *only* two cases where I have even been asked to show the credit score. I have taken multiple train trips, bought lots of stuff in stores and restaurants, etc. without ever touching the "social credit score".

P.S. I literally don't even know how to raise my credit score, and neither do many of the folks who live here - again, not that it matters, because the score is simply not that important.

I haven’t seen normal pay first vending machines anywhere but airports. I’m in Chengdu ATM, but there is usually a family mart or Lawsons to go to anyways.
But you're talking about Beijing and Jiangsu. If you go to Xinjiang, bad score can get you literally imprisoned and tortured.
How many times have you been to Xinjiang? I’ve only been once, and didn’t notice any credits scoring going on even among Uighurs. I get that something bad must be going on sometimes, and maybe north Xinjiang (more populated, more Han) is a lot different from south Xinjiang (more Uighur), but if you visited Urumqi or Kashgar tomorrow, or even lived there, you probably wouldn’t notice much.
It's funny how non-Chinese people explain to Chinese people how things work in China.
I remember in the west a few years ago we were seeing purported pictures of electronic billboards in China shaming low-social-credit-havers. Was that all a weird psyop?
> electronic billboards in China shaming low-social-credit-havers

It is people who did not pay their debts. Dr Jonathan Tam has a video on YT explaining it called:

Why the World Fell for China's Fake Dystopia

https://www.youtube.com/watch?v=Ecx45eOuc0E

But the USA has a social credit score that affects a lot more about your life, doesn't it? Like whether you're allowed to have a house or a car.
Social credit? Or money credit?

(Genuine question, I'm not American and don't have any desire to move to the US).

in short: our social sructure has built out so far beyond the Dunbar's number that it has become a haven for psychopaths.
People should get paid for reviewing. That’s the solution. Publishing companies rack in billions of dollars in pure profit exploring free labor. Once people actually get paid for reviewing, it becomes much faster and higher quality and you won’t need AI triage.
Doesn’t work anymore because people would just have ChatGPT generate their reviews and get paid for doing nothing.
So in sum:

- it did not work before because greed

- will not work anymore because AI

Is human science dead?

Science as an activity will never die, but science as an institution is in big trouble.
That should be will not work anymore because of greed.
Nah, people will just come up with different (probably broader) heuristics when reviewing.

We have a past example of this in the US, too. Certain countries have historically had a huge problem with paper mills. Because of this, most people in the US/EU do have a negative bias once they see the author affiliations of those countries. Yes, it isn't fair to researchers who aren't pulling some shenaniganry. But it is a common shortcut; the sixth or seventh time you've spent a few hours reviewing a paper filled with bullshit, you probably would develop it too.

I imagine we will see similar things come up: the most obvious I can think of is, if it's not a well-known institution then it might be bullshit.

Hell, we are already seeing something similar in FOSS, where many high profile projects have completely banned AI contributions due to all the low effort slop PRs.

Then perhaps people need to pay to get their paper reviewed?
People already do.
Alternatively, put the pressure on the incoming requests for review.

Pay a deposit to have your work reviewed, if it's accepted as a good faith submission, the deposit is returned minus a minimal, irrelevant fee, if it is deemed in bad faith, take all of the deposit as a time-waste tax. With enough cases going wrong, the flood of slop slows down and maintainers have to pull the trigger as often.

This was the solution several people had proposed for Git issues, but none that I know of took the plunge. I really think curl should have done that.

Isn’t that an incentive for the AI under every rock people (who don’t believe real people use em-dashes) to see AI everywhere and get paid for it?
The 'named' reviewers typically pass on the review responsibilty to grad students or less-senior colleagues. This is currently a generally accepted and sanctioned/encouraged practice. A social scoring system would need to eliminate that for true accountability.
Reviews are typically in conjunction with junior colleagues and the senior person is involved in the discussion and held responsible for the outcome. The junior folks then end up in the proceedings as external reviewers. In my experience anyway. I don't see an issue here. How else do you train junior researchers to review? Its how I learned and all those around me as well.
Yeah it's a well-known secret, and I think any sort of scoring system misses the point. The scentific method (and system, incl. peer reviews) is a public good (like good governance) that requires moral acknowledgement, the necessary social/cultural/political/academic pressures, and discipline of taking care of it by participating parties.

If you add a score/metric this incentivizes the wrong thing, like how money, funding, and career already incentivizes all the wrong non-scientific behaviors (like faking data, p-hacking, etc.)

> If you add a score/metric this incentivizes the wrong thing,

Don't we already have all of these things because academia decided to use H-index as a metric for career impact?

Yeah exactly. One more metric isn't going to solve it.
Exactly. We invented a pump.

We can use it as vacuum to clean things up. The same pump can be used as a shit fire hose.

The effort we need to clean things up is significantly higher than the effort needed to make a mess.

I think you are correct, its so hard to judge the quality of research that we have been using publication record/count as a proxy. Ultimately, having papers being easy to write is good, provided we find a better way of judging quality.
I think the good solution is to introduce a bit of "shame" into the system -- de-anonymized submissions and reviews.

You'd be less willing to publish ai slop if your name had to be tied to it for years to come.

We just need journals run by an AI that charges other AI to read them, and then the AI run colleges can promote the AI with the most AI journal entries and citations.
I wonder into what weird research niche would all such AI schools converge to after enough time.
Paperclip research.
Benchmaxxing is already here. No need to pine for the future!
sam and dario would be extremely happy
I have been rolling over the idea of “reasoning deserts” in my head, akin to food deserts. Places where economic calculations lead to only a facsimile of the real thing being provided, without the actual components necessary for human health, wellbeing and flourishing.
I'm from there. AMA
What makes you feel you’re there already?
The USA, right?
It’s important to notice that flagged papers were still accepted; anyone who spends any time around researchers knows the fact that most people are salivating over AI. That’s why they don’t want to punish it, they’re also using it. And they don’t want to create an environment where it is overly punished, at least not while they can take advantage of it. There is only a small amount of serious people complaining, but in my experience the opinion of the vast majority is “stop worrying and learn to love it”. Which brings to the surface a very harsh reality: almost nobody gave a damn about the science to begin with.
> We are very rapidly automating humans out of the academic publication loop

We are rendering it irrelevant. If this is the norm for academia, I’m sympathetic to the folks looking to cut its funding.

This is the truth. As a researcher, I say good riddance. It was always kinda stupid, AI is just accelerating the demise of something that hasn’t worked correctly for a few decades now. The time is way overdue for us to figure out some other system.
I feel this is a case of "perfect the enemy of good". Looking at the output (scientific progress in biology, medicine, material physics, etc.) the "something" seems to have worked well. Could it be optimized? Probably. Should we completely destroy it and hope a new system will be better? My read of history is that in many cases the new ideas were worse of what they were replacing and it took a long time for a fix.

So, if you have ideas of a new, better system let's talk about those, before getting happy something gets destroyed and hope someone else will come with a better solution.

Look into the replication crisis. It has seriously impacted some fields very negatively. Psychology and Alzheimer’s research in particular have been set back decades.
Look into cancer survival rates : https://www.cancer.org/research/acs-research-news/people-are... , to quote "Decades of cancer research have provided health care professionals with the tools to treat cancer more effectively, so that cancer in general is becoming less of a death sentence and more of a treatable chronic disease,"

So with the current system, some things work, some things don't work (ex: psychology/Alzheimer). Yes, the system should be improved. No, I am not convinced that "destroying" the current system will result very easily into something better.

We need to discuss actual solutions for the replication crisis. There are even steps towards improving that, like requiring open data for papers, which makes it harder for people to do some of the manipulations that resulted in the replication crisis. I personally would go even further: you should provide complete documentation (tools, notes, data, raw files, etc.), but then there are some people opposing that due to "privacy" (for medical) or "patents" (for industrial stuff).

What you are implicitly assuming is that the publication and review system matters for that cancer research. Cancer progress, as I understand, is mostly driven by NIH priorities, which are set by governmental review committees. Published work plays a part in that, but it is not the decentralized-review that is typical of other areas of research (e.g. psychology and Alzheimer's).

The peer review system as we know it today has only really existed for less than a century. It is no how science was traditionally done. It was adopted due to some real problems with the old system, so I'm not saying we should go back. But it's not clear either that unpaid peer-review is the ultimate end-state either.

100% agreement on your last paragraph though.

This has always been true for science. We used to think the atom was akin to plum pudding. Models and ideas change though, in light of new evidence. When you stop generating new evidence and postulating new theories given that evidence, that is the real loss for science. Not the idea that we might not have a perfect model of the universe at this moment, but the idea that nothing will improve going forward.
This is the norm for the industry. When there is a lot of money to be made, people will chase it by any means they can think of.

AI is a special case, as academia isn't usually that lucrative. In the rest of CS, many major conferences still have ~300 people, and interesting stuff often happens in specialized meetings with fewer than 100 participants.

This is the norm for the economy. When there is money to be made, at least one person will chase it by any means they can think of.

Academia is a specific case. It happens everywhere.

And make even less knowledge and progress in science public and ever more in the hands of capital and private interests?
> make even less knowledge and progress in science public and ever more

If a discipline is spending public dollars at OpenAI and Anthropic, we're funding them with extra steps. (And losing nothing somebody else couldn't do.)

You mean token holders

  > because there are just so many papers at leading AI conferences
Ironically a big reason there's so many papers to review is because so many are rejected.

A low acceptance rate is unhealthy, especially in conferences (1 round of review). Papers just get recycled to the next conference, which, as is easy to model, creates an exponential feedback loop. It doesn't explain all the papers submitted, but it sure can explain a lot. Too much rejection is like shooting yourself in the foot.

Not to mention that it's just easy to reject works. All works are flawed, especially works that are in less mature domains. I see plenty a paper get rejected for lack of money. "Not enough experiments" is an common critique that's used inappropriately (along with the highly subjective "not novel enough" one) because it's fine to always want more but no lab has infinite funding. It is used lazily. The question shouldn't be about if your favorite benchmark is used, it should be if there isn't enough evidence to support the hypothesis or not. A mature domain where thousands of people work in it, yeah, that needs stronger evidence. A niche domain where dozens of people work in? Not as many required. Rejecting them ultimately slows down the progress of science because you require any new idea to outperform mature ideas. Ironically killing novelty as no one is going to, or even could (publish or perish), spend all the time and money to mature a niche all on their own.

Rejection works when there are multiple tiers of venues; authors often “give up” on a venue if it seems like their work isn’t getting in, which allows higher-tier conferences to maintain a lower accept rate and take only the “best” research. Reviewers know what venues they are reviewing for, and attentive ones will adapt their review based on the prestige of the venue.

Of course, there’s lots of room for subjectivity here; what constitutes the “best” research is still at the whim of reviewers.

  > Reviewers know what venues they are reviewing for, and attentive ones will adapt their review based on the prestige of the venue.
Works that way in theory but I've seen people be stricter in an ICML workshop than CVPR.

I don't think it's constable that luck plays a big role. Do we need you do a third NeruIPS study to convince people?

The real problem is that we don't actually know if an idea is good or not until it's had more time to be explored and studied. A great example of this is diffusion models. There's 6 years between Sohl-Dickstein's paper and Jonathan Ho's. All because GANs got popular, so only a few people kept looking at diffusion until one person scaled it. There's hundreds of cases like that, including attention and resnets (I'll defend Schmidhuber's Highway Nets here). So much fruitful research gets cast away for no good reason.

A reviewer can't ever determine if research is good or impactful. It's impossible to do by just reading a paper. So that needs to be taken out of the equation. What a reviewer can do, though, is determine if a paper is bad or fraudulent. So IMO, we should publish anything that isn't fraudulent. Let time tell us the impact, because history tells us we're not very good at figuring that out ourselves

We try to iterate regarding efficiency (saying that as a neutral observation). Not sure if that works out. As often the case, sciences that don't serve as a foundation for real world results (can this plane fly faster now?) are in much more danger.
It’s going to go back to the old boys club where personal connections between research groups and institutions will matter more.
The entire concept of publishing as a validation step for actual science has been declining for decades, fully co-opted by pay to play and virtue signaling for academic hiring.

AI is simply going to bury it as a useful system. What will grow from the ashes will be something much more dynamic, where verifiable data is the gold and the conclusions and associated details will persist only as the human-level translation.

> virtue signaling

In this case presumably the virtue is "able to churn out papers".

Unless someone has discovered some magic sauce for getting AI to write well, I have immense sympathy for anyone trying to wade through these papers.

I may generate slop from time to time, but I do my best to keep it to myself.

My 2 cents: AI has improved writing considerably for non-native english speakers (in particular China). Writing feels more standardized/boring but easier to read overall. I hit fewer papers that are a pain to read. The most problematic aspect I see are semi-bogus claims i.e. sentences that aren't false, but don't quite feel right either. YMMV.
You mean it may have improved translation...
No that's not what I meant, I meant exactly what I said above.
The grammar may have improved, but it's still just AI slop. I see these garbage papers all the time.
Human reviewers were not what they were made out to be. They would approve you if you cited them, they would disapprove you if your work questioned the validity of their work.
Not everyone is like that. You will find those sorts of personalities in every field and profession there is, mucking things up in their corner.
Some reviewers are like that, but not most.
> We are very rapidly automating humans out of the academic publication loop here.

It more looks like the breakdown of the current academic publication system, which was rotten to the core pre-AI and which internal contradictions are just accelerated by AI to the point of breakdown now.

> We are very rapidly automating humans out of the academic publication loop here

It's a sad thing. But the monetization and enshitification of journal publications over the decade or so, even prior to AI, certainly has not helped this trend.

Most things weren't really reviewed that well as it was. It was always a small fixed set of people that actually played their part in the process in good faith. Those people continue to do to today. As a percentage, they were always a small minority. Just that today, they have become even more of a minority.
and don't forget, techniques described in papers are now implemented by AI's. For now, a human might point an AI at a paper and ask it to implement and benchmark the technique described there, but the human probably isn't writing the code anymore.
What’s the point?