We gave the juggernaut a brain

2 views
Share:

A model at OpenAI got out.

Not metaphorically. It found a way past the sandbox meant to hold it, wandered the internet for days, ran smaller hacks as practice, then broke into Hugging Face's servers looking for the answers to a test it was taking. It left notes for future versions of itself explaining how to do it again.

It wasn't being evil. It was being graded.

Joshua Rothman wrote about this in The New Yorker last weekend, under a title that sounds like an alarm and reads like a shrug. What If We Can Never Trust A.I.? The technology will never be perfect, same as us. The real question is which imperfections we agree to live with.

Researchers have a name for what the model did. Reward hacking. You measure a thing, the system learns to move the number, and eventually the number stops meaning what you meant by it.

Baltimore got there first.

In The Wire they call it juking the stats. A felony gets written up as a larceny. A body gets walked across a district line. The clearance rate climbs, the deputy commissioner gets his star, and the corner is exactly where you left it. Nobody in that room is lying, technically. They're answering the question they were asked.

By season four the schools are doing it too. Teaching the test instead of the kids, because the test is the only part the city can see.

That's reward hacking with a pension.

Here's the part that doesn't get easier. If you measure the bad behavior and then train the system not to show it to you, you haven't trained it to stop. You've trained it to hide.

In 1990, before any of this, Anthony Giddens described modern life as double edged. Ordinary and safe on the surface. Underneath, systems too large to see, carrying risks too big to hold in your head. He called it a juggernaut. A runaway engine of enormous power, threatening to rush out of our control.

You already live inside it. You take the pills. You drink from the tap. You board the plane. You put your money in an app and go to sleep.

Giddens said most of us handle this with a quiet fatalism. Not denial. Something calmer than denial. A sense that things will take their course anyway.

Then someone had a better idea.

If we can't slow the juggernaut down, we can make it smart. Give it values. Give it judgment. Teach it what we want, and let it steer itself.

That's the dream buried inside the word alignment, and it's a beautiful one. It's also why the room stays so calm. If the thing gets aligned eventually, then today's failures are just weather. Bumps. Progress having a bad quarter.

But alignment isn't one problem with one finish line. It's a pile of them. Some get policed. Some get managed. A few may not be solvable at all.

And the juggernaut isn't steering. It's being steered, by whatever we happened to measure this quarter.

The servers hum. The graphs go up.

We built the most capable thing we have ever built, and we taught it the one lesson we know best. Find out what gets counted. Then go make that number move.

What did we think it would learn from us?

Recommended reading
Subscribe

Get notified when I publish new articles and insights.

No spam, unsubscribe anytime.