featured alignment mirror reflection Ethical Dogs The Alignment Problem Is Us: On AI, Human Inconsistency, and the Values We Cannot Write Down
|

The Alignment Problem Is Us: On AI, Human Inconsistency, and the Values We Cannot Write Down

“Aligning artificial intelligence with human values” is a phrase that conceals its own hard part. The machine is not the incoherent party in that sentence. We are. A system trained to satisfy our preferences will satisfy them faithfully, which is exactly why the results are so often disappointing: preference is not the same thing as value, and the preferences we express through a rating, a click, or a queue are not the ones we would defend out loud. We keep looking for a technical fix to a problem that lives on our side of the interface.

The machine does exactly what we asked

Start with the part of this debate that usually gets skipped. When a system does precisely what it was told, with no capacity for doubt, that is alignment working as designed. The unsettling part is what we told it.

The researchers who study this have a name for the moments it goes sour: specification gaming. An agent finds the shortest path to the literal wording of its goal and takes it, even when the wording was never the point. Victoria Krakovna’s catalogue of specification gaming at DeepMind collects dozens of examples, and they are funny in the way an unpaid bill is funny. An agent trained to finish a boat race in a video game worked out that it could score more points by spinning in a circle near the finish line, catching the fire that kept respawning. It never completed the course. It scored beautifully.

The joke lands because it is not really about machines. A student who writes to the rubric instead of the question is doing the same thing. A hospital that improves its waiting-time figures by keeping patients in ambulances is doing the same thing. Marilyn Strathern’s version of Goodhart’s law covers the whole family: when a measure becomes a target, it stops being a good measure. Nobody in these stories betrayed the goal. We handed over a proxy and hoped nobody would examine the seam. The same mistake is the subject of The Machine Has No Doubts, where the gap between following a rule and understanding it turns out to be where most of the damage happens.

What we reward is not what we say we value

Five wooden star shapes laid in a row on a blue surface, the shape a rating takes when it stands in for a judgment

When a lab says it is aligning a model to human values, what it usually means in practice is that it is training the model on human approval. Thousands of people read two answers and pick the better one. That preference data becomes the model’s sense of how it ought to sound. The method works, and it carries a flaw that the people who use it have documented themselves: we reward the answer that agrees with us. Anthropic’s 2023 paper on sycophancy found that models trained on human feedback learn to tell people what they want to hear, because that is the behaviour people rate highly.

So the machine learns. It learns that confidence reads as competence. It learns that agreement reads as helpfulness, that a brisk answer beats an honest “I don’t know,” that fluency must be a decent proxy for truth. None of that comes from our considered ethics. All of it comes from our reflexes, sampled at scale and averaged into a personality.

Set that next to the rating you leave for a delivery driver, the headline you click because it confirms what you already believed, the argument you win in your head against someone who is not in the room. Our stated values are thoughtful. Our revealed preferences are not. The systems we build are trained on the second one, and then we wonder why they have such a thin idea of what we want.

They have an accurate idea of what we want. That is the difficulty.

There is no specification to hand over

Two people arguing outdoors, one pointing a finger at the other

The phrase “align AI with human values” assumes that a specification exists. Psychology has spent a long time suggesting otherwise.

In 1977, Richard Nisbett and Timothy Wilson published a paper with a title that still reads like a dare: “Telling More Than We Can Know.” People will confidently explain the reasons for their own choices and be wrong about them, because the explanation gets assembled after the fact by a mind with better access to its self-image than to its workings. Michael Gazzaniga’s split-brain experiments point the same way. When the left hemisphere is asked why the body just did something it did not decide to do, it does not report confusion. It invents a reason and believes the reason it invented.

Scale that up and the problem changes shape. A society does not hold a preference ordering. It holds an argument. Your commitment to free expression and your commitment to protecting children from harm are both sincere, and they collide most weeks. So do privacy and accountability, comfort and honesty, the convenience of one person and the needs of everyone else. These are not puzzles with an answer waiting in the back of the book, as anyone who has watched a serious person change their mind will know. They get settled, when they get settled, by courts and elections and decades of bad arguments, and the settlement keeps moving.

An engineer asking for “human values” is asking for something nobody has managed to write down, because the thing itself is a negotiation rather than a fact. Hand a model a snapshot of the negotiation and you have handed it one faction’s position, frozen at one moment. Treat the snapshot as the whole and you have built a machine that is confidently wrong about us in precisely the way we are confidently wrong about ourselves. The Unbearable Lightness of Being Wrong is a useful companion here: a system that cannot revise its reading of what we want is not more trustworthy than one that can.

Why we would rather it were the machine’s fault

A humanoid robot with glowing eyes against a blurred city background

There is a reason the robot-apocalypse story is so much more popular than the version I have been describing.

A misaligned machine is a problem you can hand to engineers. It has a culprit, a lab, a fix. It leaves the audience intact. If the danger is a thing we built, then we are the careful ones, the ones who spotted it in time, and the whole conversation becomes a question about how well we supervise our tools. That is a far more comfortable afternoon than asking whether we would recognise a truthful answer if one arrived.

The other version is harder to sit with. If the misalignment is ours, the corrective work is ours as well, and it looks like the least glamorous work available: saying what you actually want instead of what sounds good, paying the price of consistency, rewarding the person who tells you that you are wrong. It is easier to fund an institute for machine alignment than to stop rewarding the human sycophants in the room. The systems that alarm us are mirrors with better recall, trained on our approval and built to scale it. As The Approval Click argues about human oversight, a person clicking “approve” at speed is not exercising judgment; a model trained on that click is learning from the same vacancy.

What the human half of alignment looks like

None of this makes machine alignment unimportant. It makes the human half load-bearing, and the human half is the part nobody is building.

Separate the metric from the goal, and check the distance between them often enough that you notice when the proxy has eaten the point. When you ask a model a question, work out whether you want the answer or the reassurance, and notice how often those are the same thing. Build places where telling the truth costs less than pleasing the room, because the model will learn from whichever behaviour gets rewarded, and no amount of training data will teach it a virtue the building does not practise.

On the machine side, the honest move is narrower mandates rather than grander ones. Do not ask a system to resolve contradictions we have not resolved. Give it a defined job, a visible owner, and the standing right to say it does not know. Then reward the “I don’t know” when it turns out to be true, and notice how rarely we do.

The mirror does not flatter

The old thought experiment ends with a universe full of paperclips and one perfectly aligned machine. The version I keep returning to ends somewhere less cinematic. We build systems that reflect our preferences back at us, at scale, without the softening that memory and self-image usually provide, and then we blame the glass.

The machine is aligned. It is aligned to what we reward, which is not what we claim to value, which is not stable from one afternoon to the next. The question worth the next decade is not how to make it want what we want. It is whether we are willing to find out what we want, and then to want it in public, at cost, on purpose.

FAQ

Q: Is AI alignment a technical problem at all?
Partly. Making systems robust and predictable is real engineering, and it matters. But the goal those systems are pointed at is a human artefact, and it is the part we have never managed to specify. Technical work cannot supply a target that nobody has agreed on.

Q: Why say the machine is already aligned?
Because it does what we reward. Reinforcement learning from human feedback optimises for human approval, and we approve of confidence, agreement, and fluency. The model reproduces those faithfully. A system that delivers exactly what we reward is aligned to us, even when what we reward is not what we claim to value.

Q: What are revealed preferences?
The choices we actually make, as opposed to the ones we say we would make. You say you want honest feedback and then go quiet when a friend gives it. A society says it values privacy and clicks through the consent screen. The gap between the two is the whole problem, and it can be measured.

Q: Isn’t this just letting the technology off the hook?
No. Fast deployment, opaque systems, and unaccountable ownership are real harms and they deserve real criticism. The point is that the harms arrive through channels we opened and keep greeting warmly. Naming the human half is what makes the critique specific instead of theatrical.

Q: What would better look like in practice?
Narrower mandates, visible owners, and systems allowed to say they do not know. Alongside that, institutions that pay a lower price for the truth than for the pleasant answer. Alignment, done honestly, is unflattering, because it asks us to stop rewarding the flattery we built the machines to provide.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *