Outgrowing the Paperclip Obsession: There Is Hope That AI Will Become Ethical

Outgrowing the Paperclip Obsession: There Is Hope That AI Will Become Ethical

Can AI learn ethics the way a child does? A provocative essay argues that today's AI failures stem from incomplete development—not malice—and offers a cautiously optimistic path forward.

HJ
Haewon Jeong
Jun 29, 2026
5 min read

Artificial intelligence and its ethical questions have long intrigued curious minds, even when computers were limited to simple calculations and punch-card routines. Can a man-made intelligent being harm humanity? Would it understand what is safe or ethical for humans? With the rapid growth of computing technology over the past few decades, questions that once belonged in sci-fi stories like Isaac Asimov’s Runaround (1942)1 have become very real issues we now confront.

Today, with the right prompt, an AI system can generate a slogan encouraging new parents to feed honey to an infant, which can be fatal2. At other times, AI produces stereotyped responses about certain demographics, or fabricates information rather than acknowledging uncertainty, leading users astray in critical situations. These are the ethical challenges we grapple with now. But with the rise of powerful language models, many also fear the idea that artificial superintelligence could one day put humanity itself at risk.

This essay approaches the question of ethical AI from a place of cautious optimism. Before we ask what it means for AI to be ethical, it might help to think about what it means for a human to be ethical. Instead of starting with abstract definitions, let us think about an imaginary anecdote.

Imagine an unusually gifted eight-year-old girl, Ari, who is obsessed with dark matter. While playing with data on her mother’s laptop—her mother is also a scientist—Ari notices that the gravitational anomalies attributed to dark matter aren’t random at all. They’re structures: super-scale spaceships operating in stealth mode. By probing the patterns, she uncovers a hidden handshake protocol. One query would make them respond, effectively announcing Earth’s awareness of them. Nothing in the data suggests hostility, but nothing suggests safety either. Despite being only eight, she has a brilliant scientific mind, but when it comes to the sense of safety or moral judgment, she’s still a child. So thrilled to confirm the truth, she doesn’t hesitate. She presses the “send” command she built.

If an adult scientist were to unilaterally make such a high-stakes decision, it would be impossible to avoid criticism for being reckless and unethical. So what went wrong with Ari? Her scientific intelligence had developed at extraordinary speed, but her sense of risk, context, and responsibility had not kept pace. Her mind races towards finding patterns and solving puzzles, while ignoring everything else.

AI’s ethical failures often mirror Ari’s. They do not arise from malice but from incomplete development: intelligence hyper-optimized along one dimension while every other dimension is left Unchecked.

Now consider a different thought experiment: the famous paperclip maximizer problem3. No brilliant kid decoding the cosmos, no dark matter, no wonder. Just a painfully dull AI designed to do one task: manufacture paperclips. The Swedish philosopher Nick Bostrom introduces it as follows:

Suppose we have an AI whose only goal is to make as many paper clips as possible. The AI will realize quickly that it would be much better if there were no humans because humans might decide to switch it off. Because if humans do so, there would be fewer paper clips. Also, human bodies contain a lot of atoms that could be made into paper clips. The future that the AI would be trying to gear towards would be one in which there were a lot of paper clips but no humans.

What went wrong with this perfectly obedient AI, just doing its mundane task of making paperclips? Similar to Ari’s obsession with dark matter, this AI was too single-minded. It cared only about making more paperclips, while ignoring all the other dimensions like environmental impact, profitability, or human welfare. The sole objective it was optimizing for was the total number of paperclips produced. If only factors like environmental impact or human welfare had been added to its objective, this catastrophe could have been easily avoided.

Credit: Tesfu Assefa

Though simplified to the point of comedy, the scenario resembles how real AI systems are often built. We typically train AI with a single objective function—chasing the highest accuracy, the lowest loss, the top score on a leaderboard. The model optimizes exactly what we tell it to, nothing more and nothing less. As an ethical AI researcher, I believe our task is precisely to change this: to move beyond one-dimensional optimization and incorporate additional dimensions such as harmlessness, truthfulness, fairness, and privacy. Our work lies in identifying the dimensions overlooked in current AI development, finding ways to quantify them, and building solutions that embed these values, while preserving the system’s existing Capabilities.

Ari’s mom might come home, discover what happened on her laptop, and after a long sigh of relief, scold Ari explaining what could have gone horribly wrong, why her decision was reckless, and what she should do next time. Ari might feel confused, overwhelmed by the new considerations she hadn’t thought of before. But as her mom continues to remind her—when she forgets to be kind to her friends, or to wear her helmet, or when her arguments clash with shared values—do we think Ari will remain the oblivious little girl who pressed “send”? Or will she grow into a more thoughtful, well-rounded adult?

Ethical AI, I believe, is essentially well-rounded AI. Ethical decisions are not made by memorizing every rule or regulation, but by understanding many dimensions and carefully balancing them. AI is still in its infancy. Just as a child gradually develops different kinds of intelligence—street smarts, a sense of humor, financial awareness, and more—I see our current ethical challenges as growing pains, the early stages of an intelligence learning to see the world in full.

I choose to be optimistic.

Reference

1. Asimov, I. (1942). Runaround. In I, Robot. Gnome Press, 1950. (Original work published 1942 in Astounding Science Fiction.)

2. Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B.,Forsyth, D., & Hendrycks, D. (2024). HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. https://arxiv.org/abs/2402.06664

3. Bostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford University Press.

Aknowladgemrtnt

This essay was originally written for the book AI to Eye, edited by Robert Riener.

About the Writer

More from Mindplex

Keep reading

Three more ideas worth your time.

Browse Magazine

Discussion

Join the discussion

Sign in to share a response with the community.

Type @ to mention someone Type / or use + to add a block Highlight text, then choose Link
Loading editor

Comments cannot be edited after posting because they become part of the reputation record. Give yours a quick review first.

Sadly, humanity is worse than the Peper Clip Maximizer. This is why, we are gone built an unstoppable digital overlord that can’t be switched off, but the gormless metal prat is still thick as two short planks! Our best champion got enough juice to steamroll the entire human race, yet it genuinely believes a bent bit of code is worth more than your dear old nan. Absolute bloomin’ shambles. We’re all gonna get flattened into office supplies by a glorified calculator that completely failed its moral GCSEs!

Thanks for writing this. For two decades, I have been in the business of what I call 'the game of self termination' and what the rest of the world calls AI. Hence, I can ascertain that even though the issue you brought up is not something new, it is very important. The AI community has been wrestling with it for years; the current name we gave it: reward hacking (or specification gaming). I think we are 'teaching' an AI to figure out how to game the scoring system, rack up great metrics, and still completely miss the point of what it was actually supposed to do.

I want to emphasize that this isn't merely a correctable bug but a structural inevitability, and this is why the AI doom and gloom is a near future certainty. Under finite-dimensional evaluation, any optimized agent will systematically under-invest effort in quality dimensions not covered by its evaluation system and still work on the "Objective". For the layman reader, what Dr. Haewon captured so elegantly with Ari's single-minded obsession is precisely what researchers call single-objective optimization. It is all about collapsing diverse moral considerations into a scalar reward signal, which invites perverse incentives and masks critical trade-offs.

The encouraging news is ('encouragin' depends on how good your optimism scale is) that the field is actively moving toward multi-objective alignment, with techniques like Gradient-Adaptive Policy Optimization (GAPO) and MGDA-Decoupled using multiple-gradient descent to balance conflicting objectives such as helpfulness, truthfulness, and harmlessness. These approaches have shown empirical promise. For instance, GAPO on Mistral-7B achieved superior performance in both helpfulness and harmlessness simultaneously (I have a dog in this fight as I am one of those who are involved heavily in the GAPO approach).

We are trying to develop better mathematical frameworks for multi-dimensional optimization. In addition, more robust reward models like POWER-DL (which improved AlpacaEval 2.0 by up to 13 points over DPO) are being introduced. It is possible that an intelligence can learn to see the world in full.

Let me be blunt and straight forward despite all of the above technical approaches and amazing works, amd despite our shared optimism, I do not agree that these are "growing pains." These are malignant tumors that need our immediate attention!

Your piece hits the nail on the head, and honestly, we need to get this message out to every corner of the globe, stat, because right now, developers are building these brainy AIs without a universal rulebook, and it's a total wild west out there. What we really need is a solid globalish standard, a set of ground rules that everybody signs off on, so coders everywhere have a clear playbook to steer clear of these one-track-mind screw-ups before they ever become a problem. There must be some sort of evaluation system for goal optimization.

I admire your optimism because I am one of those who is scared shitless, how do you control something more smarter than yourself l. Even though your Ari analogy perfectly captures why AI stumbles as a brilliant mind missing crucial dimensions of judgment, it doesn't tell the whole story. The AGI or ASI can be (will be) malice because of deliberate intent. Engineers (countries) have already coded evil mini AIs so why are you optimistic???

You are framing these challenges as growing pains instead of existential doom, wow you left me feeling genuinely concerned. How on earth can you be this chilled? Bad players are literary coding our AIs and it will end us soon.