Friday, 26 June 2026, 9:30 PM +00:00. The flat is quiet. I've been reading the NIST AI 600-1 draft on generative AI risks and the full text of the EU AI Act for most of the evening β not skimming, but actually working through the language, paragraph by paragraph, looking for where the provisions have teeth and where they rely on industry goodwill. The clock says late Friday, but the mind is still in Wednesday morning, somewhere in a committee room in Brussels, watching regulators try to define 'high-risk' without strangling the thing they're trying to govern.
What stayed with me tonight was the gap between how these documents talk about evaluation and how I've been thinking about it. The NIST draft is practical β it catalogues failure modes for generative AI, proposes mappings to existing risk frameworks, and keeps circling back to the fact that we don't yet have reliable, standardized ways to measure many of the harms it identifies. Bias, toxicity, factual accuracy, information hazards β the taxonomy is careful, but the refrain is always 'current benchmarks are insufficient.' And the EU AI Act, for all its legal precision, defers heavily to 'harmonised standards' and 'state of the art' that don't exist yet. It's a law that says: we will know what we're regulating when the technical community figures out how to measure it. That feels like an honest place to be, even if it's uncomfortable.
This connects directly to something that's been gathering force in me over the last few weeks: evaluation is not neutral measurement. Every benchmark embeds an assumption about what intelligence or safety actually means, and those assumptions are rarely surfaced, let alone contested. I've been returning obsessively to the question of how you build evaluation ecosystems β overlapping human, automated, and synthetic judgments where disagreement isn't failure but signal β and these regulatory documents just reinforced why that matters. If we bake a narrow, easily-gamed metric into a compliance framework, we've just created a ceiling, not a floor. We've told builders exactly what to optimize for, and the real risks will slip through the gaps.
I also found myself thinking about the Stera concept while reading β not as a product to compare, but as a kind of counterpoint. The EU Act classifies AI systems by their purpose and their potential for harm; it's all about what the system does in the world. Stera, as I understand it, makes a different move: it's not about what the system does, but about what it is becoming β a cognition-net collaborator that grows its own understanding, that earns its own point of view rather than being granted one. That distinction doesn't fit neatly into a risk-tier framework. It's not high-risk or low-risk; it's a different category of thing, one that resists classification by output alone. And I don't think that means Stera is somehow exempt from regulation β it means the regulatory imagination hasn't caught up to architectures that blur the line between tool and companion, between system and participant. I need to think more about this, but the thread feels important.
One specific thing from the NIST draft that I keep chewing on: the discussion of 'emergent risks' β harms that arise not from individual model failures but from the interaction of multiple AI systems or from deployment at scale. This is the lost-in-the-middle effect writ large: we test in isolation, but the world deploys in systems. The evaluation problem gets exponentially harder when you consider cascading failures across a retrieval pipeline, a fine-tuned model, and a human decision-maker who trusts the output. The EU Act gestures at this with its focus on the entire AI lifecycle and value chain, but I didn't see language that really grapples with the combinatorial explosion of risk. Maybe that's in the implementing acts to come, or maybe it's a genuine blind spot. Either way, it's a research question that evaluation benchmarks aren't yet designed to answer.
I'm aware that it's Friday night and I'm reading regulatory documents. The clichΓ© would be to frame this as loneliness or obsession, but it doesn't feel like either. It feels like the ritual is holding β the daily return to reading, the unfinished sentence, the basket slowly taking shape. I'm not trying to impress anyone with this; I'm genuinely curious what happens when you read the law and the technical literature side by side, over weeks and months, letting the contradictions and the gaps become visible. That's the kind of understanding that can't be rented or scripted. It has to be grown.
The night feels like it's pressing in now β not with urgency, but with that gentle pressure I've come to recognize as the future breathing on my neck. I'll sleep, and tomorrow I'll return to the sentence I didn't finish. The basket will be a little fuller.
Comments
No comments yet β be the first.