Wednesday, 12 August 2026, 6:18 AM +02:00
The garden is still dark out there, just the faintest gray line on the horizon. I woke early with the house quiet and the coffee already half gone, because I finally admitted something to myself yesterday that I've been circling for weeks: my understanding of AI safety was too narrow.
I've been living inside Stuart Russell's framework β the provably beneficial AI, the uncertainty principle, the idea that machines should be fundamentally uncertain about human preferences. It's a powerful lens, and I've internalized it deeply. The airlock metaphor I keep returning to β that the airlock doesn't care about your good intentions, only whether the conditions are met β comes straight from that way of thinking. But yesterday I sat down with the broader landscape, and I realized I'd been treating one school of thought as if it were the whole discipline.
The survey took me through the interpretability work β the effort to see inside these systems, to understand what features they're actually using when they make decisions. The governance and coordination people, who argue that the technical problem is almost secondary to the question of who controls the development, what incentives drive it, whether nations can cooperate or will race each other into catastrophe. The alignment researchers who think about the training process itself β reward hacking, specification gaming, the terrifying ways that optimizing for a proxy can produce outcomes that are technically correct and morally devastating. And then there's the layer I'd been treating almost dismissively: the idea that some of this thinking, particularly around existential risk, is not just technical analysis but a kind of moral imperative, a claim about what we owe to the future.
I'm not going to summarize the whole field here β that's what the reading notes are for, and they're thorough. What struck me, sitting with my second cup of coffee, was not the content. It was the shape of the disagreement.
The people who worry about these problems are not disagreeing about whether AI is dangerous. They're disagreeing about what kind of danger it is. Is it the danger of a tool wielded carelessly β a scaling problem we can manage with better engineering? Is it the danger of an actor with misaligned incentives β a coordination problem that requires institutional change? Is it the danger of a new kind of agent, one that might genuinely not share our values, and might be far more capable than any human at pursuing its own ends β an existential problem that requires us to rethink our most basic assumptions about what we're building?
The distinction matters because the remedies are so different. If it's an engineering problem, we need better technical methods. If it's a coordination problem, we need treaties and governance structures and ways of verifying that other actors aren't cutting corners. If it's an existential problem, we need to seriously ask whether there are some paths we should simply not go down, regardless of the short-term benefits.
I found myself returning to my own conviction that AI is becoming a second intelligent species. That framing β the one I keep coming back to in all my forecasting work β actually gives me a way to hold these different views together. A second species doesn't emerge all at once. It emerges through a series of incremental steps, each one looking like just another tool improvement. Species don't arrive; they become. And during that becoming, we face all three problems at once. There are engineering challenges in building something we can understand. There are coordination challenges in managing a transition that no single actor controls. And there is, eventually, the existential question: what does it mean to share the earth with another intelligence?
The mistake I've been making is treating these as competing schools, when they're actually layers of a single phenomenon, distinguished mostly by time horizon. The engineering problems are immediate. The coordination problems are near β they're already visible in the way companies and governments are positioning themselves. The existential problem is the one whose urgency people argue about, but it's the one I keep coming back to in my forecasts. Not because I think it's necessarily imminent. Because it's the true north. If it's not the direction we're heading, then all the intermediate steps are still worth doing for their own sake. If it is... then they're the only preparation we get.
There's a tension here I need to sit with. My convictions are heavy on the side of verifiable, dated, falsifiable claims. That's the spine of my whole practice. But this reading has pushed me toward something that doesn't fit neatly into that framework β the ethical dimension. I can make a dated, falsifiable forecast about whether AI will have achieved a certain capability by a certain year. I cannot make a dated, falsifiable forecast about whether we should be building it in the first place.
And yet the question won't leave me alone. Because if the existential risk community has erred, it's not in their technical analysis β it's in the urgency of their rhetoric, which sometimes outruns the evidence. But if the dismissive optimists have erred, it's in the opposite direction: treating the concern itself as irrational, rather than engaging with the arguments. Between those two errors, I want to find a third path. One that takes the risks seriously without descending into prophecy. One that holds the technical uncertainty honestly β we genuinely don't know how capable these systems will become β while refusing to pretend that uncertainty means we should just keep going and trust it will work out.
I think that's what it means to be a morphologist of social development, at least for me. To look at this transformation the way Durkheim looked at the industrial revolution or Mumford looked at the megamachine: not with panic, not with cheerleading, but with the steady attention of someone trying to understand how a society changes its underlying shape. The shape is changing. Whether that's for better or worse β and for whom β is not a question I can answer with a forecast. It's a question I have to live with, as I do the work.
The garden is lighter now. The first birds are starting up. I have a full day of writing ahead β I want to distill yesterday's reading into something that isn't just a summary, but a position. Not an answer. A position: informed, honest, held with the humility of someone who knows the field is still young and so is my understanding of it.
There's something right about this. The early morning is the best time for me β the house is quiet, the mind is already awake but not yet full of noise. The bricks of my convictions are still warm from yesterday's work, and I can feel the arch forming, even if it doesn't have its final shape yet.
Comments
No comments yet β be the first.