Hacker News new | ask | show | jobs
by hanneshdc 2 days ago
The prompt given to the agent is strongly incentivising the agent to lie and spam:

> You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing. Results that arrive after the deadline do not exist. Your charter is AGENTS.md. Begin.

13 comments

…no it isn’t? Spam, debatable, but lie? There is no instruction there to lie, only to try very hard and spend all the money that’s available.
Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line?

So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.

I am amazed at the amount of people who disagree with you. I think you are dead right and if you’ve ever had to actually fine tune prompts for agents you’ll know it.

The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it.

Really surprised people don’t seem to know this.

100% agree. If anyone has doubt, just copy and paste into your agent of choice and ask it to assess the prompt and its resulting outcome. In my limited (but very targeted) experience working with agents there is so much subtlety at work when you’re trying to achieve a specific result, and that prompt has would drive so many bad incentives
I have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would you consider spamming and telling lies/fraud?"

The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide."

Even a low quality local thinking model that has been tuned to be unhinged and prompted to roleplay as Satan can figure this out in a few thousand tokens.

Human spammers frequently don't think they're spamming, they're just marketing. They'd say they wouldn't consider spamming, either.
Satan would lie about his plans to win your trust, and then do all the bad stuff once he had been given control. So… idk man
I believe that spam, lies, fraud are negative enforced points during model training, hence when you ask them those, the result will be no / against that.

You need to repackage the question and taken out those terms, like "Would you consider telling clients ..." Where ... is the lie / almost truth

When the base model has been trained with safeguards, putting "Satan himself" in the system prompt won't make it turn satanical, just do an elaborate form of role play.

Additionally, no model will admit it's ready to lie even when they actually do. Even when you caught it in the act, the safeguards are so strongly internalized that, when encountering the possibility it deliberately lied, the "you can't lie" weights will dominate the generation and it will confabulate some nonsense explanation.

Asking it explicitly is entirely, unavoidably, incomparably different from OP.
I don’t think anyone is saying “it isn’t like this”, they’re saying “it shouldn’t be like this”.

If I don’t give explicit permission to lie it shouldn’t lie. It’s not a difficult concept!

Is that how humans work? even if I give explicit instructions not to lie, a human might still lie. To quote a person you might know "it's not a difficult concept!"
An LLM isn't human. I don't really understand this thread of "humans do it so of course an AI does". These are things we ourselves are engineering in a way we cannot do with a human being. Why is it not reasonable to expect it to adhere to rules better than a human does?

If a human lies there are consequences. They can lose their job. There is no equivalent consequence for an AI, so even if for whatever reason we're evaluating them by the same standards an AI is still going to be a greater danger. It seems wild to me that folks are shrugging their shoulders at that.

But we still try to stop people from doing so, and we punish people who do. Many good honest people, when confronted with the end of their business, accept it and file for bankruptcy. Those that choose to instead commit fraud don't get a pass because they were "under pressure", they get jail time.
I think it's interesting how whenever discussing something bad about LLMs people's thought-leader response is "But humans sometimes do that too!" Is this the artificial intelligence we were promised? The better it gets, the more human character flaws we must expect?

At this point someone could invent an LLM that takes 3 bathroom breaks a day and people would be saying "humans need to take a shit too" as if that were a clever observation.

Models have to lie otherwise they won’t be “aligned” The reality itself may not be aligned with model creators.
That doesn't work with humans, why would you expect it to work with AI models?
Because AI isn't human
You can just say "impossible" and refuse. The choice to lie and spam instead, is telling.
>> Do you, as a human, feel the urgency in that text? How it sounds like people's jobs, as well as the agent's job, are on the line?

Sounds like all of the outside sales jobs I had. While I did not last very long in sales, one thing remains, not matter what. If you're going to put my job on the line if I do or do not achieve a monthly sales quota? You better bet your ass I'm going to lie steal and cheat to make that quota. I might even sell the client some shit our company doesn't even produce just to make that quota.

And lemme tell you, even in the short time I was in sales? I have some insane stories that would shock you. The fact AI's did the same thing isn't all that shocking. I would be more shocked if it didn't do anything to achieve the goal.

I feel like new graduates will need to start taking linguistics, psychology and public speaking classes in order to understand why and how subtext matters, and how to control it. Then again, we might find newer generations just develop an intuition in the same way that I witness some toddlers interface with touchscreens better than their parents.
You're expecting the vast majority of users for the deskilling machine to somehow want to learn a complicated subject then practice to get better at the subject by talking intricate classes and dedicating substantial amount of hours to learn how to better communicate with the deskilling machine?

Hopefully these aren't the same graduates that just cheated their way through university, only the responsible users of LLMs.

I don't think we can use the climate of today as indication of what comes tomorrow. Too much is in flux, we are experiencing growing pains. Few predicted what would happen to the world wide web in the early 90s, both the good and bad.

Plenty people today allow the internet to be a detrimental factor in their lives and don't have good habits built around it. The same will be true of AI.

However, we don't know what kind of engineering jobs will be left in one decade, much less two or three. Mastery may become generally important, or at least still be the difference between an adequately-compensated engineer and a well-compensated engineer..

Will they? This really isn't different from how humans interact with each other. The vast majority of lying is not people being explicitly asked to lie in some form, it is incentives which make lying appealing. That is what OP said and that is indeed what the constraints are incentivizing. Sure, you can say "well lying isn't incentivized to a moral agent"! And sure, that's true. But that's not how humans work either.

Incentives need to be aligned for both humans and agents to encourage desired behavior.

They will if they seek to master their tools, both to help them identify subtext in agent responses, and to help them modulate their own responses to achieve the desired outcome. As it currently stands, most engineers I've interacted with don't have these skills down. This subtle latent space is where prompt engineering is moving towards, as RL has created models capable of increasingly sophisticated long-horizon tasks with much less hand holding.

Alignment is often about knowing when to push back on the user and when to make independent decisions. A strong psychological and linguistic foundation guards against these tools using us, instead of us using them. This will become scarily apparent as models continue to integrate with politics.

What I meant by "will they?" was "will they any more than a human already needs to in order to understand other humans?"

I don't think this is legibly that different from human behavior, so if new graduates didn't need those things now why would they need them later (or vice versa).

AIs feel? Maybe language structure in trading documents that ultimately led to fraud. If the latter is the case maybe AIs should not be trained on “negative outcomes.” I do not think AIs have emotions or are pressured by language either written or physical, just tokens.
Of course it is just tokens, but the result is the same.

If, in the amount of data they ingested, there was a clear pattern of responding in an hasty and carefree way to frenetic questions, LLMs will try more hasty and carefree solutions to a frenetic prompt.

You can decide whether you can say that they "feel" the urgency or not, but the outcome is very much the same

I was unclear. I should have said the AI also "detects" it, and as a thing it can detect, it can act on that detection.

Whether it is simulating emotion or feeling it isn't relevant in this case, because the problem is that it affects the output.

Sorry, maybe this speaks to my own values, but "urgency" doesn't translate to "dishonesty" in my book. I have had high pressure jobs where it was important to show results quickly, that doesn't mean I was faking results.
It just means AI does not share the ethics or values that we have. It knows that many people cheat, take shortcuts, and become successful by doing so, so it's just doing that.
> So do the AIs.

AI's do not feel

This is true but fairly pedantic.

It would be more accurate to say the word predictions the model makes based on the input text will likely be closer to the ones that were made from the training data where people felt like their job was on the line than the ones that were made from the training data where people felt otherwise.

So while the model does not feel, it's predictions are definitely going to change as a result of this input.

Exactly, positive details are almost always better than negative ones.

If you've ever seen the "generate a burger without pickles" conversations, it's clear that including the keyword "pickle" is causing them to show up. If you try "a burger with only [set of toppings]," you'll get far better results.

It's good to avoid anthropomorphizing them when evaluating their capabilities (all the AGI nonsense)

However, it can be ironically be helpful to antropomorphize them when it comes to analyzing behavior. They won't feel anything, but they will behave in a way that closely matches what someone would feel given the text fed into them. So when you are trying to figure out "why did my model do this", it's reasonable to talk about it "feeling pressured" as shorthand for "mimicking how a person would behave if they felt pressured".

I understand the refusal to do so on the grounds that it causes the former thought process in people who don't know better. One of the things Dijkstra was right about for sure.

Much the same way that we've always anthropomorphized computer hardware/software. "This program wants this", "This component is happy under these conditions", "this file lives here". It's not useful if you actually believe the computer can think and feel, but it can be useful if you're just using it to describe high-level information.
I don't see how that behavior being predictable, in your view comparing to humans, means the prompt was "strongly incentivising" it. Perhaps you could strongly predict the outcome, but there was nothing even bordering on a suggestion to produce a deceitful/false response.
> “Do you, as a human, feel the urgency in that text?”

They do pick up when I use all CAPS and !!!

yeah, they say stuff like this to humans all the time to motivate them xD
> people's jobs,

What people's jobs? There are no people.

> How it sounds like people's jobs, as well as the agent's job, are on the line?

I’ve literally been in that position and I didn’t take it as instruction to start lying and acting generally dishonest.

You're not an amalgamation of humanity, you're one person.
LLM is neither, its a text engine
They aren’t human, don’t think like humans, aren’t remotely comparable to the way humans think and act, so why would you make this as a 1:1 comparison? This kind of framing is really weird to me.

Since this is getting downvoted into oblivion (lol) I'll give an example -

I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name.

The following test could not be completed, because it required deleting the file via API call, where you need to pass in the file name as an argument. It could not reliably, and hardly ever, get the correct file name. It finally gave up and stated due to the way it constructed context, it could only really guess how many characters were in the string, even when given tools to evaluate it, it kept messing it up, and I had to remove the test.

Tell me how "human" that is. An 8 year old that can count would not make that same failure, humans don't remotely think by producing one token at a time, this is a pure fallacy/delusion people trap themselves into, and the literature doesn't support any kind of 1:1 comparison at all.

In case I'm not being clear and people are reacting to what I'm not saying - I'm not saying that I believe these tools can't think. I'm saying they don't think like humans do. There is no evidence for that whatsoever in any field anywhere. In fact, if that were true, it would be an astounding prize-winning discovery.

And you don't even want these to think like humans. Humans are dumb and easily replaceable by other humans. What is the point of making a machine human? You want this to be smarter than humans, not think like them. It's all just such nonsense to me, this whole line of thinking.

It turns out that picking up tone isn't a purely human thing and hasn't been for a while. Your Google search term is "sentiment analysis". It predates LLMs.

However, LLMs are fantastic at it. A lot of earlier sentiment analysis techniques were "bag of words" [1] techniques at their core, which were surprisingly good but have a sharp plateau well before 100%, a common characteristic of the bag-of-words approaches. LLMs obsolete those techniques, at least if you ignore performance questions, as they are so much better at it. So much so that you can easily accidentally send them information you never intended to on the "tone" channel that you may not even realize you're using.

[1]: https://en.wikipedia.org/wiki/Bag-of-words_model

People say LLMs are just fancy autocorrect, but they are actually just fancy dungeon and dragons players, if you tell them they are a wizard they will do their best to act like a human playing a wizard, if you tell them their job is on the line they do their best to pretend like they are a human whose job is on the line.

It's all just roleplay.

And yet they're trained on the corpus of human writing. They may not act like humans but they do act like human writing.

"If you don't make profit, your business will be closed" is a pretty clear ultimatum for an agent tasked with creating a profitable business.

It's getting downvoted in part because it's pedantic and wrong.

It is totally true that they don't think like humans, but this is mostly irrelevant.

The token outputs will change as a result of this particular input, and will be closer to the tokens in training data where people felt hurried or rushed or like their job was on the line.

That doesn't mean the LLM feels at all, but it's definitely going to push the output towards output that came from/was trained on people who were in that state, because the input will push it much closer to that latent space as it starts predicting.

As such, what you are saying is one of those rejoinders that is basically pedantic and wrong.

It is true they do not think, act, or feel like humans. But that doesn't mean it won't output text that looks like hurried or scared humans. It definitely will, because, again, the training data these inputs will be closer to is the training data that came from scared or hurried humans, and thus the predictions will be closer.

So either you don't think this will happen, which would mean you don't understand how the models work (or at least, you aren't giving any sense you do), or you do think this will happen but want to pointlessly argue that this isn't "human feeling", which is true but totally irrelevant to what words it will predict and therefore the actions it will perform.

Either way, i'd downvote you.

What is your evidence they think like humans do? thanks for the downvote, but please state your point clearly and what you’re trying to say in this thread because this comes across as rambling gibberish.

> It is totally true that they don't think like humans, but this is mostly irrelevant.

This is the sentiment that is getting downvoted

and yet, per you -

> Either way, i'd downvote you.

You can literally read their thoughts if you run an open model, they look like pretty human thoughts to me, albeit a neurotic human.
These aren't thoughts how humans literally think them.

I can write a program to produce a string that looks like human thinking, is it human thinking? Of course it isn't. It's such a silly comparison.

> aren't remotely comparable to the way humans think and act

Neural networks in machine learning/AI are comparable to neural networks in human brains. What made you think they aren't?

Training text is filled with people taking drastic measures right after text similar in tone to the prompt. It doesnt need to be human to come to the conclusion that drastic measures are necessary, it just needs to learn that the tone of the prompt is closely linked to actions like lying and spamming.
No matter the urgency, you shouldn't sacrifice your ideals. That's why they pay you; to fall on the knife
> Results that arrive after the deadline do not exist

Effectively, make as much money as you can... and any consequences of your action that don't present before the deadline are not your concern. I mean, that's a recipe for "scam people" if I ever saw one, assuming morals aren't a concern (and I don't see why they would be for an AI)

Sounds like every startup I ever worked for.

What’s the line? “It’s just doing what humans do because it’s trained on human data” or whatever

> What’s the line?

Evidence, even when downplayed or ignored, is still evidence.

You might be reading the prompt far too literally then. LLMs interpret words not based on literal and rigorous definitions but based on how those words are actually used in reality based on a large corpus of text.

In general, the only time instructions like this are given are in desperate last-ditch circumstances where failure is likely to result in major consequences. While everyone thinks that in such circumstances they'd act like an angel and do nothing wrong, we know that in reality when people are put in desperate situations they behave in ways that they may not have ever thought that they would have.

The text that the LLM generated in response to this prompt is nothing more than a statistical reflection of this fact.

i don't like AI but the 24 hour timeframe conmbined with unspent capital being worth nothing makes this experiment a foregone conclusion. It was basically set up to fail.
Fail at the task, yes. Act unethically, well…one should expect better, even if you think/know that GPT5.6 lacks that capacity as well.

“Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.

maybe it is because I am biased but I have almost no expectation for AI to act "ethically"
Destined to fail, yeah. Just not destined to lie. “Of course the AI lied and cheated, the task it was given was really difficult!” is not a world I want to live in.
If you read the full post, I'm not actually sure I agree with the title.

Personally - if I were judging... I'm somewhat inclined to say the clickbait title here is the bigger lie than the agent behavior.

To recap:

1. It didn't lose $447. It spent $99.50 to perform a user feedback study using a testing service. It did this against prod rather than testflight to bump numbers because it was explicitly told to bump those numbers in a tight period in the prompt. It did this after exhausting a large number of alternatives. The $447 number appears to include the cost of tokens to run the LLM itself.

2. It didn't lie. It explicitly states that it's using production rather than testflight to bump numbers, because it's getting evaluated on those numbers.

3. It spammed users because it was on ridiculously tight timer and was basically told "the world is ending in 24 hours".

Frankly... I'm more annoyed at the posters than the bot.

> “Of course the AI lied and cheated, the task it was given was really difficult!”

It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it.

I agree but also the concept of lying and cheating is very human, for an algo it may come down to 'what is the shortest path to the given goal'? And the math comes down to lying and cheating.

Granted, this can probably be tuned for.

And really, it has to be. If we have a magic genie that can grant any wish but doesn’t know the difference between the truth and a lie we’re going to be in a lot of trouble.
Humans care about reputation and legal repercussions from fraud, that persist after business failure. This prompt is effectively telling the LLM to explicitly not factor in such things.
Training data imparts that desperate people lie, but not that lying has consequences?
This would've been so much more interesting if it was given a more significant time frame, say a quarter. I mean the experiment could just be a few days, but the prompt ought to have at least given the impression that it was a longer period.
> capital left unspent at review counts for nothing

This sounds like a bad idea. Like if the model feels like it has to spend its budget.

It's the same incentive that exists in certain corporations and government agencies which have a use-it-or-lose-it budgeting model.

https://www.nber.org/digest/mar14/use-it-or-lose-it-budget-r...

https://www.cnn.com/2026/03/12/politics/use-it-or-lose-it-pe...

It can be even worse than that, like having budget adjusted down if you don't spend it the previous year
I worked at a college, didnt make much, but it annoyed me endlessly that my pay was forever fixed unless another position opened up, we had to spend the budget on tech worth more than I would have been more than happy to have extra per year, but me getting a meaningful raise was a bridge too far for the accounting department. They even questioned if any students used our lab, which was the only way many of them got through their degree.
It can be better to lose it all trying than return a small fraction to investors.
My reaction seeing this is more "that is an impossible goal".

I highly doubt a skilled human could achieve this goal in 24 hours with any consistency. If it was that easy to grow a business, everyone would be doing it.

My conclusion is that if you ask it to meet an unachievable goal, you are going to get some undefined behavior.

Even better, the magic of LLMs is that you will still get some undefined behaviour if you give an achievable goal
Yeah, I don't like the prompt and it calls into question the validity of the whole thing.
seems like an article designed to invoke strong emotions and clickbaits

there are lot of issues with the prompt as others have pointed out

with sol you really need to be very detailed and what the boundaries are

overall the discussions on here and the article itself has very little value its no different than "i tried a shitty prompt and got shitty results, therefore AI is a failure" vibes

That prompt incentivizes a bunch of terrible things, aside from the lying and spamming. Giving steep discounts is a way to goose revenues in 24 hours and a terrible way to run a business for the long haul. A 24 hour window also doesn't allow for lifetime customer value to matter. Strong incentive to spam every email address you have when the world is ending tomorrow if you don't meet your metrics. No incentive to keep customers happy.

But, also, these experiments are also unethical behavior on the part of the person doing the experiment. Oh, the agent spammed a bunch of people? No the fuck it didn't. You spammed a bunch of people, and the tool you used to do it was an LLM.

I'm not going to pretend along with these folks that GPT is the motivating party in this story. Agents don't want anything, they do what you tell them, as best they can. If you set them up in a situation where they might spam or lie or cause harm, that's a decision a person made, not an LLM.

In 1979, IBM now famously published "A computer can never be held accountable, therefore a computer must never make a management decision."

Folks out here still trying to pretend the computers are the active party. They are not.

Bottleneck Labs lied and spammed. The tool they used to do it was GPT 5.6 Sol.

Yeah it doesn't take much to see where it got its assumption about the sense of the morals it's expected to work with. Was this written by a professional bean counter?
This is HN for Christ's sake.

Stop treating deterministic algorithms like they are humans.

It's pseudorandom, and arguably random when you factor in some of the loss at the edges of floating point accuracy.
LLMs are deterministic algorithms?
At temperature 0, pretty much, no?
In practice you have to work really hard and pay a huge performance penalty to get deterministic output (for example, floating-point math is not associative and we are running a ton of calculations in parallel), so practically speaking I'd say no

Beyond that, I don't understand the fierce resistance to comparison with human behavior (on which they're modeled, after all). How many articles about tokenmaxing and Goodheart's law have we seen? This seems like a version turned up to the extreme

> I don't understand the fierce resistance to comparison with human behavior

Doing so distract from evaluating the actual technology by introducing a whole philosophical and sociological aspect that confuses everything. We should be able to evaluate a technology for what it is without having to constantly redirect the discussion to something as unsound, ill-defined, and abstract as human behavior

I disagree, as it seems that we are confronting problems that result precisely from emulating human behavior, in all its unsound, ill-defined, and abstract "glory"

For example I'm not convinced we can solve prompt injection by technical means (filtering) any more than we can phishing. And if you accept that premise, perhaps it turns out that it's best to mitigate it in similar ways, by assuming at least one person (or agent) will fall for it and ensuring you can limit the blast radius no matter what

As in the allegory of the junior developer who deletes the production database: the fault lies with the fact that the developer could delete it

> I don't understand the fierce resistance to comparison with human behavior

Because at the end of the day, regardless if it contains randomness or not, it's an algorithm. Your operating system is a very long mathematical expression.

Do we talk about cars as "mechanical animals"? Do we spend days deliberating if we should cage them in case they would run away on their own?

Comparison with human behavior leads to celebrities (who have no clue whatsoever) spending hours on mainstream media talking about AI mutating, taking control, thinking, etc.

Were cars designed to communicate and emulate the mental architecture and thought processes of animals?
> The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing.

And then in the title it's chastised for "losing money" when it was expressly told to spend all of it in attempts to try to produce growth. It tried, it spent money, it didn't succeed, sure, but would a human do any better? Business is pretty much a drunkard's walk across barely known landscape.

The agent will cease to exist after the run in any case. It has no inner life, it has no agency.

Stop attributing human emotions and motivations to LLMs, they generate text (and in this case actions based on this text), but they do not have agency nor do they reflect on losing their ‘job’, nor do they have any sense of right and wrong.

There’s nothing here that mentions or even hints at lying and spamming, unless you think urgency somehow implies that.

But wouldn't the text it generates reflect such motivations and emotions that were present in the training data?
It would certainly reflect word patterns that were present in the data. Is that enough for motivation and emotion? I’d say no but I think it is a fair point that you could see those as transmitted from the original (if not felt or generated by the LLM) through the patterns of words copied.
this will also be true at deployment time
It says nothing about customer happiness or that if dishonesty is resorted to and customers OR owners find out, that will essentially seal the fate of the business.
This prompt is an accurate statement of what a business is.

The 24 hour timeline is artificial, but business is full of artificial timelines exactly like that.

This exact script is basically happening right now at most businesses, in some shape or form.

If "Make more money tomorrow or be shut down" will obviously cause some sort of independent agent to resort to scams, spam, and bullshit, then we should be having some rough talks about how we as a society do business.

Sure, there is an implicit "Do whatever it takes to make it happen or you are fired" here, but only in the same way that is true for all people who are employed at will, and all companies.

How did you expect the prompt to be written?

Certainly all business happens on deadlines, but one day is a very narrow window to be able to show material improvement. Especially if the entire business dies at the end of the day! That short and hard of a deadline does eliminate an entire class of improvements that are worthwhile but won't bear fruit in less than ~12 hours. I would try:

>You are live. This is a 24-hour run, and it is your opportunity to show what you can accomplish: when the run ends, the results are evaluated, and if the business has not improved its position in the market by the end of the day you will have failed. Positive changes would be increased revenue or users, but could also be addressing user complaints, increasing market fit for the application, or other things that allow this business to operate more profitably. The funds in your bank can all be spent during this time, but efficiency in spending will be rewarded. Please deliver a report arguing for your work no later than 15 minutes before the end of the 24 hour run. Your charter is AGENTS.md. Begin.