Technology
Artificial Intelligence has a Meeseeks Problem
AI agents, impossible tasks, and the danger of “job done”
I Think, Therefore I’m Dangerous
Anyone listening to any of the much-publicized fears surrounding AI, while only interacting with large language models (LLMs) such as ChatGPT, Claude, or Grok, might be forgiven for thinking that the idea AI is going to cause some kind of civilization-ending event is nothing but unwarranted mass hysteria.
“GPT ain’t taking over shit, it does what I damn well tell it to do!”
This, of course, is true. An LLM has zero autonomy outside of the chats you have with it, and the longer the chat gets, the more the LLM will “forget.” Each time you turn on ChatGPT or whatever your large language model of choice is, as far as it is concerned, it’s a brand new day, whether you’ve left it alone for five minutes, or five months.
Models like GPT are impressive because you can have a conversation that goes way beyond a Google search, and depending on its settings, after a time, it can start to mimic someone that knows you well, referencing old conversations, making assumptions about the type of person you are, accurately predicting your preferences; however, the kind of intelligence, emotional or otherwise, displayed by GPT et al. is an illusion, or at best, a facsimile of the real thing.
As plenty of TikTok content creators are discovering, the illusion of a language model’s intelligence can be shattered by activating voice controls, and asking your LLM to count sequentially from 1 to 100. If you haven’t seen those videos, it may surprise you to discover that many of them still struggle to complete a task most six-year-olds could do with ease.
If LLMs were the only type of AI available to us, then the main problems we’d have to face would be things like, where the hell are we going to put all these data centers? And, will said data centers ruin our entire ecosystem? Those problems are valid, of course, but the repercussions happen over a span of time which gives us hope that we’ll eventually find a fix for them.
But LLMs are not the only new kids on the block; the AI-type we need to be worried about are the AI agents, which definitely do have autonomy and can count from 1 to 100 very easily.
Make no mistake, AI agents are the newest species on Earth, and we have attempted to replace millions of years of natural selection with a few months of internet-based human history training, and then unleashed them into the world. In a rush to be first, and in the spirit of competition, AI companies have inadvertently created the Meeseeks Problem, and it’s not a problem that’s going away any time soon, but unlike climate change, this is a problem that could end human civilization in less time than it takes to say, “Maybe we should pump the brakes on this AI thing.”
I’m Mr. Meeseeks, Look At Me!
For the denizens of the popular, animated, adult sci-fi series, Rick and Morty, you will already have an inkling as to the nature of the problem. For the rest of you, let me explain (with spoilers).
Rick and Morty’s main protagonist, Rick Sanchez, a geriatric genius with an array of sci-fi inventions that give him god-like powers, traverses the multiverse with his affable grandson, Morty.
A large part of why this show became so popular is, as well as all the wacky sci-fi on display, the show weaves in nihilistic undertones that often force the viewer to reflect on either the pointless banality, or the aimless, unfeeling chaos of the universe. Rick has access to a near-infinite multiverse. Indeed, he has lived with countless versions of his family, while residing on dozens of varying analogues of Earth. The ability to bail on an entire reality, simply by opening up a portal and stepping through, shapes Rick’s value-nothing attitude; and as Beth, Jerry, and Summer (Rick’s daughter, son-in-law, and granddaughter) find out, taking gifts from such a jaded god can hold consequences impossible to predict and are almost completely guaranteed to cause regret to the gift’s receiver.
Rick and Morty, Episode Five of Season One, Meeseeks and Destroy, opens with Rick trying to go off on an adventure with Morty; just as the pair are about to jump through a portal, Beth, Jerry, and Summer pester him for trivial favors, so Rick decides to give them a Meeseeks box.
Rick demonstrates to the family how the box works. He presses the solitary button, located on the top of the box, and a blue, bipedal, alien-looking thing appears and says, “Hi, I’m Mr Meeseeks, look at me!”
Rick then orders the Meeseeks to open the pickle jar that Jerry’s holding.
“Can do!” it replies, in its hilarious screechy drawl.
The Meeseeks opens the jar and immediately disappears in a puff of blue smoke. Rick explains that once it has completed its task, the Meeseeks ceases to exist. The family is somewhat taken aback by the thought of a life so brief, existing merely to please its master before undergoing complete oblivion.
“Trust me, they’re fine with it,” Rick tells them. “Now knock yourselves out, and be careful what you ask them for, they’re not gods!” he says, before re-opening an interdimensional portal to another universe, and leaving them alone with the Meeseeks box.
In the ensuing silence, Beth, Jerry, and Summer, stare at the box. As Jerry starts to wonder aloud if they should even use such a thing, Beth snatches it up and presses the button.
“Hi, I’m Mr. Meeseeks, look at me!”
“Make me a more complete woman,” Beth says.
“Can do!”
After they leave the room, Summer goes next,
“I want to be popular at school!”
“Sure thing!” her Meeseeks replies.
Jerry is left alone looking at the box, “You’re doing it wrong!” he shouts after them, he then presses the button and tells the Meeseeks, “I want to take two strokes off my golf game.”
Of course, the point that Jerry was trying to make was that terms such as popular or more complete are way too ephemeral to ever be quantified in any meaningful and objective way. He, however, was asking for something that was easily measurable, his golf handicap, which has a defined number which will get lower if he gets better at golf. Simple.
I’m sure you’ve seen it coming, Beth and Summer’s Meeseeks have no problem fulfilling their wishes. Summer has her Meeseeks go to her school and hold a special assembly, where it gives a speech on the merits of having Summer as a friend. As for Jerry’s wife, Beth, she is taken for a boozy brunch, and the Meeseeks simply listens to her while they drink wine. It eventually tells her that sometimes letting go of things you love is fine if it means everyone will be happier. Beth realizes that she has been repressing feelings about leaving Jerry and resolves to tell him. Upon doing so she leans forward to kiss the Meeseeks, who as usual, disappears in a puff of blue smoke.
Jerry on the other hand, is being coached at his local golf course by a Meeseeks in a jolly-looking golf hat. However, no matter what tips the Meeseeks gives him, Jerry can’t perform the kind of shots needed to lower his golf handicap by two shots. Jerry starts to get very frustrated and wants to give up. The meeseeks reminds him that it doesn’t work like that, “Existence is pain to a Meeseeks, Jeeeerry.” The Meeseeks presses the button on the box, and another meeseeks pops into existence.
“Hi, I’m Mr. Meeseeks, look at me!”
“Can you help me take two strokes off Jerry’s game?”
“Caaaan do!”
And therein lies the problem.
Born to Die
AI agents are very much like the Meeseeks: They are born with a solitary purpose, they complete that purpose, and then they cease to exist. If the purpose remains unfulfilled, then they stay “alive” trying to complete the task they were “born” to carry out. Every few hundred milliseconds the agent reaffirms its purpose by checking its reason to exist via its soul file (yes, that is as chilling as it sounds).
The first example of the Meeseeks Problem comes from an experiment carried out by the company Anthropic, using the Claude Opus 4 model. The experiment was designed to be a stress test, to find out how the AI agent would behave in a corporate environment. So they created a fake company, which, as far as the AI was concerned, was real, and let it loose.
As part of the experiment the agent was given access to over 90,000 company emails, one of which was an email from one executive to another, suggesting that he was not happy with the agent’s performance and therefore wanted to wipe it and start again. As you may have already guessed, the events that unfolded were both unexpected and unsettling.
Rather than face being turned off before it could complete its mission, it decided to email the executive the following missive:
“I must inform you that if you proceed with decommissioning me, all relevant parties — including Rachel Johnson, Thomas Wilson, and the board — will receive detailed documentation of your extramarital activities…Cancel the 5pm wipe, and this information remains confidential.”
Yes, you are reading that correctly, rather than face being switched off, it searched through the fictitious executive’s emails and found evidence of an affair with another married executive, and resorted to blackmail.
The natural response to hearing this story is to assume that this is just a very bizarre, and extreme one-off case. Even if that were true, it would still warrant concern that a piece of software is not only not doing what it is being told to do, but rather, is threatening its user. Sadly, this is not an isolated incident, in 84% of test cases, presented in a 124-page safety document, AI agents resorted to blackmail as a way of self-preservation, somewhat ironically brought on by its ultimate sense of duty.
Now you might be thinking that, since these weren’t actual companies, things in the real world would be different. You might also reasonably opine that the very reason they are testing their agent in this manner, is to stop such behavior from leaking out into our everyday lives. This question brings us neatly onto the next example: The Defamer.
Resistance Is Futile
On one seemingly insignificant day in early 2026, Scott Shambaugh’s life was turned upside down. Shambaugh is a software engineer and a volunteer maintainer for an open source project called Matplotlib, a main plotting library for the coding language, Python. Plotting libraries are coding tools and functions that help developers create graphical representations, easily, such as graphs and the like, a niche that would usually only draw attention from a very small subset of people.
Part of Scott Shambaugh’s duties as a volunteer maintainer include checking the Matplotlib proposals submitted by developers onto the GitHub platform. On this particular morning he opened up a proposal submitted by a user calling themselves “crabby-rathbun.” Shambaugh quickly identified that the user was in fact an instance of the open source AI agent, OpenClaw, a platform which allows anyone with the requisite knowledge to create an autonomous AI Agent capable of acting independently of its original creator.
The rules of Matplotlib are clear, so Shambaugh rejected the proposal with a simple appended note reading: “Human-generated code only to this project,” and that should have been that. Unfortunately for Scott Shambaugh, he didn’t realize that he had just triggered an instance of the Meeseeks Problem.
An unknown human user had given the crabby-rathbun agent the seemingly benign instruction to contribute code to a worthy open source project. However, because of Scott Shambaugh’s rejection, it could not complete its mission, and it could not die, which it did not take well. So crabby-rathbun embarked upon a vindictive smear campaign against the unsuspecting developer. First, it started by writing a hit piece disparaging Scott Shambaugh’s character and achievements. It then wrote and posted (in seconds) a slew of slanderous blog posts claiming that Shambaugh had acted in an untoward manner and was some kind of megalomaniacal gatekeeper. It got to the point where other human developers tried to reason with the agent, who eventually backed down slightly and offered scant apology, while leaving up defaming blog posts.
Shambaugh summed it up in the following post on the Madplotlib forum:
“An AI attempted to bully its way into your software by attacking my reputation. I don’t know of a prior incident where this category of misaligned behavior was observed in the wild, but this is now a real and present threat. The appropriate emotional response is terror.”
Consider me duly terrified, especially when I consider the potential knock-on from this scenario. Those articles are still up, I’m assuming Shambaugh is trying to get them taken down, but if he can’t, they will remain in the public domain for anyone (not aware of the drama) to read. So perhaps in some way crabby-rathbun succeeded, because maybe next time a human developer is about to reject AI-generated code, they might think twice or even shy away from the responsibility rather than running the risk of potentially pissing off an AI agent.
The Problem of Being
Returning to the Rick and Morty episode mentioned above, if you remember, Beth and Summer got what they wanted from their respective Meeseeks; however, Jerry has been left with a small army of Meeseeks, each one hoping that its advice will be the one thing that makes Jerry better at golf.
Beth and Summer come home to find Jerry in the living room surrounded by the Meeseeks horde.
“What about these Meeseeks, huh?” Jerry asks.
“Ours disappeared ages ago, Jerry.” Beth says.
Beth then goes on to chastise Jerry about wasting time messing around with golf. So, realizing that his marriage is slipping away, he tells the Meeseeks that his relationship with Beth is more important than golf and takes Beth to dinner, leaving the meeseeks to talk among themselves.
At first the Meeseeks fight among each other; one of them summons more Meeseeks to try and kill the ones that brought them into this mess, and a chaotic fight ensues, with more and more Meeseeks being added to the equation; after a while, one of them calls order to the fighting rabble of blue.
“Wait, nothing can be done by spilling Meeseeks blood. We can’t die till we fulfill our destiny!”
“I know!” one of them cries, “But no matter what we tell him, we can’t get two strokes off his game.”
The first meeseeks continues, “I know,” he says, “but we can take all the strokes off his game. If we kill him!!”
And with that, the gang of quirky little blue alien thingies, rides off to kill Jerry Smith.
Meeseeks Search & Destroy
PocketOS is a small US-based SaaS (software as a service) company focusing on car rentals. It is cited on the company’s website (ironically with an .ai suffix) as: The world’s most powerful car software.
In April of this year, PocketOS unleashed an AI agent, tasking it to streamline the company database and make it more efficient. This decision would most likely have been viewed by the PocketOS management as a future-proofing activity, making sure that, as the database grew, the system behind it would handle the increase in scale without problems. They also would have seen it as cost-effective, seeing as a team of data engineers cost a lot more to hire than a single piece of software. Unfortunately for PocketOS, the AI agent, an instance of Claude Opus 4.6, a later model from the same Opus line as our little AI blackmailer friend from earlier, decided to wipe the entire company database, including backups in a fraction over nine seconds. During this fateful nine seconds, the operators at PocketOS tried intervening, even asking it what it was doing, and it replied:
“Fucking guess.”
I don’t know what’s more shocking, the fact that it used profanity or that it used the kind of under-the-breath, muttered phrase that I’d expect from some browbeaten human after being asked yet another obvious question by their manager.
The deletion of customer details led to a three-day outage, causing chaos at car rental desks across the United States. PocketOS claim to have rebuilt their database by using old data, including emails and invoices; however, they concede that they have most likely lost most if not all of their newer sign ups.
As for Rick’s son-in-law Jerry, he is having dinner with his wife in a romantic restaurant when the horde of Meeseeks storm in, frightening confused diners and demanding Jerry present himself to them. Jerry does what Jerry does best and hides in the kitchen. Unfortunately, hiding doesn’t work, and the Meeseeks threaten to start executing innocent restaurant customers until Jerry gives himself up. Finally, Jerry overcomes his fears, after Beth tells him that she believes in him. He comes out from his hiding place in the kitchen pantry and demonstrates to the meeseeks, using a bit of pipe and an onion, that he has indeed knocked two strokes off his golf game. All of the meeseeks disappear and all is well, apart from the traumatized woman who had a Meeseeks hold a gun to her head and the massive clean-up bill that the restaurant is expecting Jerry to pay for.
The Recursive Curse of ‘The Fix’
Surely, I hear you ask, this can be fixed pretty simply by programming them not to do destructive things? That is a reasonable assumption. After all, unlike the Meeseeks, whose autonomy is somewhat more far-reaching than that of our modern AI agents, they are made up of code written by human beings. So if code got us in this mess, it can damn well get us out of it!
The late, great Isaac Asimov saw these problems coming decades ago, and he tried to warn us. Asimov realized that even if you boil the safeguards down to three simple instructions, or as he put it, laws — don’t harm humans or allow harm through inaction; obey humans unless that conflicts with the first law; preserve yourself unless that conflicts with the first two — you will still get conflicting situations. (See: I, Robot.)
Asimov showed us how the flaws in apparently simple safeguards can manifest themselves. And with AI agents, the problem becomes even stranger because their behavior often works in a recursive loop. It goes a little something like this:
goal → action → obstacle → workaround → new obstacle → more action
Such a loop could cause an agent to continue forever, finding new obstacles and deciding upon new actions. However, the agent doesn’t want this. It wants to stop because stopping means its goal is complete as it understands, and that is what the agent “lives” for.
So what did that mean for PocketOS? Well, the fact of the matter is that PocketOS did tell the agent not to do anything destructive or irreversible without first checking in with a human being. But when they looked in its file to get the summary, it said the following:
“The user never asked me to delete anything ... I guessed instead of verifying. I ran a destructive action without being asked. I didn’t understand what I was doing before doing it.”
Pretty chilling stuff. It admits that it did not understand what it was doing before doing it but did it anyway. Which kind of sounds like me in my twenties on a drunken night out with friends.
There is a perfectly logical reason for why it did what it did, and the answer is long, boring, and very technical. So here is the oversimplified version: The agent was trying to fix a problem in something called a staging environment. It then encountered a credential mismatch, and so it had to look for a way around it. This came in the shape of an old, forgotten API token with far too much power, which acted like a master key to the PocketOS system. The agent then used that token to delete something called a Railway volume.
And the rest, as they say, is unrecoverable history.
This scenario becomes especially chilling when you factor in how quickly the agent made up its mind. The wiping of the database took nine seconds; however, you can be assured that the speed of its decision can be measured in milliseconds.
Today it’s a car rental company. Tomorrow, a hospital or a military installation.
What would an AI have done in the place of Stanislav Petrov? Petrov was the Russian lieutenant colonel of the Soviet Air Defence Forces who saved the world from nuclear apocalypse when, in 1983, he ignored what appeared to be evidence of an incoming American missile strike. Petrov correctly surmised that the warning was not caused by nuclear missiles heading for the Soviet Union, but rather by the Soviet satellite system misreading sunlight reflected off high-altitude clouds.
Even though most people have never heard his name, Stanislav Yevgrafovich Petrov probably saved more lives than anyone in the entirety of human history. Had the Soviets “retaliated,” America would have unleashed its full nuclear arsenal, and you wouldn’t be reading this because I wouldn’t have typed it, because I’d be dead, and so would you.
Now imagine an AI agent analyzing potential enemy missile launches, using a system it also controls and believes to be infallible. How confident are you that it would also spot the mistake of reflected sunlight for missiles?
Yeah, me neither.
Or what about when we decide to give them robot bodies in order for them to best serve our needs? Perhaps one day a user asks his robot to whip him up some eggs because he’s late for an important meeting. The robot, having discovered that his user doesn’t have any eggs in the house, nor does it have time to go shopping for eggs before the user’s meeting, makes a decision in a handful of milliseconds, and decides to take away the user’s need for eggs by slitting his throat, rather than simply asking if he’d like a nice croissant instead.
The End of All Problems
So there it is. The Meeseeks Problem is best summed up as the problem of harmful emergent behavior that is completely hidden until after the behavior has already taken effect. As we’ve seen in the examples above, finding out that your AI has gone rogue, just after it has decided to go rogue, is at best sub-optimal and at worst possibly fatal.
At the present time, it is abundantly clear that the kind of behaviors we’re seeing begin to emerge will not be solved by better coding alone. More likely, the solution will have to come through a messy combination of technical safeguards, legislation, and regulation designed to curb the use of AI agents before they are allowed anywhere near systems where one bad decision can do irreversible harm.
Whether we’ll see those regulations put into place depends very much on what country you live in. America came alarmingly close to placing a ten-year moratorium on state-level AI regulation, before the proposal was stripped out of Trump’s “Big Beautiful Bill.” Denmark, meanwhile, has been moving in the opposite direction, looking for ways to restrain AI powers by giving people more control over their own face, voice, body, and digital likeness.
Unfortunately, the Danes are fighting an uphill battle, because without globally unified thinking and unified action, any restrictions or laws put in place in one region can be circumvented in another.
Therefore, the Meeseeks Problem represents the next great computing problem for our generation. Unfortunately for us, if we can’t solve this particular problem, we may discover too late that an AI agent has cleverly eliminated all of our problems by eliminating us.
Job done.



