You've sat at this dinner.
Eight security leaders around a table, good wine, better steak, and a conversational register I've come to think of as the auditor voice. Everyone has a mature program. Everyone's board is engaged. Everyone is "well positioned" on AI. By dessert you start to wonder whether your dinner companions are interviewing for their next role or rehearsing for their next examination — because nobody talks like that to people they trust.
Here's the thing: I don't blame them. The performance is rational. Regulators have made optimistic security statements a personal liability, so the public register and the careful register have fused into one. And in any room of eight security executives, someone genuinely is interviewing. "We have a mature program" is the only utterance the game allows.
But I sit in a different seat. Gotham works inside the environments of a lot of regulated firms, and I’ve seen the vulnerability backlogs. I know the counts. And I can tell you something no individual at that dinner can say about themselves: everyone's backlog looks like that. The confident table talk is a genre convention. The delta between what gets said over dinner and what sits in the ticket queue is the industry's real risk posture.
The emperor isn't the exception. The emperor is the whole table. And until recently, that was survivable — because nobody had instruments.
Then Somebody Built Instruments
In April, Anthropic announced Project Glasswing and gave roughly fifty organizations access to a frontier model called Mythos that found thousands of high-severity vulnerabilities in the world's most critical software — including flaws in every major operating system and browser, some of them decades old.
The instructive detail isn't the headline number. It's what Cloudflare reported after running Mythos against more than fifty of their own repositories: other frontier models, run through the same harness, found most of the same underlying bugs. Mythos, however, finished the job — taking low-severity findings that would sit invisible in a backlog forever and chaining them into single, severe exploits.
Read that again. The gap between the gated model and everyone else's model wasn't discovery. It was completion. And completion gaps between AI labs are now measured in months, not years.
So, the question everyone is asking — which model, whose lab, who gets access — is the wrong question. Sooner or later, frontier capability emerges on every platform, including open-weight models with no vetting program attached. Gating buys the defense a head start. It does not buy a moat.
The Panic and the Dark
The Glasswing members are, by most accounts, in something close to panic mode. Not because the model is dangerous to them — because it quantified them. Years of deferred triage surfaced in weeks. The backlog they always knew existed now has a number, a severity distribution, and in many cases, a working proof of exploitability.
Everyone outside the consortium is quieter. Waiting for the other shoe. But the dark isn't safety — it's deferred bad news. Your backlog is the same as theirs. It's just unmeasured. The only difference between a Glasswing member and everyone else is that the member has seen the report.
Discovery Just Became Free. Remediation Didn't.
I've argued before that AI initiatives fail not because of the models but because of broken data infrastructure — the model was never the problem, the plumbing was. Security is now running the same movie. Vulnerability discovery is becoming cheap, fast, and universal. What hasn't changed at all is the human and organizational machinery required to triage findings, ship patches, verify fixes, and prove it to a regulator.
That machinery is the scarce resource. Call it absorption capacity: can your organization metabolize what any frontier model will find in your environment, faster than the world can exploit it? When discovery was slow and everyone's auditors were equally blind, the dinner-table performance was affordable. When any competent actor can quantify your debt from the outside, the gap between your talking points and your ticket queue stops being a reputational nuance. It becomes the attack surface.
Fight Chains With Chains
Here's the encouraging part: the same reasoning that makes frontier models dangerous is available to the defense, and it changes what "absorption" means.
Remember what distinguished Mythos — not finding bugs but chaining unremarkable ones into a breach path. Attackers don't respect tool boundaries. A stale identity, a permissive cloud trust relationship, an unpatched workload, and a detection gap each look routine in isolation. Together they're a toxic combination — a complete path from outside to crown jewels. Your ticket queue treats them as four unrelated medium-severity items.
The defensive answer is to run the same chaining logic against yourself, continuously. Platforms like Tuskira are built on exactly this premise: AI agents reasoning over a live digital twin of your environment — identity, cloud, endpoint, network, controls — to map which combinations of exposures are actually reachable and exploitable, validate whether your deployed controls would break the chain, and identify the single control change that closes the most paths at once. In one financial services deployment, that approach reduced 12.3 million raw findings to under half a percent of actionable risk and cut triage from three weeks to thirty minutes.
That last number is the whole point. Nobody patches their way out of twelve million findings. You absorb an AI-scale backlog by letting AI tell you which fraction of it forms a chain — and breaking the chain, sometimes with a compensating control, before a patch even exists. Absorption capacity isn't headcount times patch velocity. It's reasoning capacity applied to your own environment before someone else applies it for you.
What a Structural Response Looks Like
The good news about "it's not the model" is that the answer is model-agnostic. You don't have to bet on which lab wins. Every firm needs the same three things regardless:
First, verified knowledge of what you're actually running. Not vendor-attested ingredient lists — evidence derived from the binaries themselves, the way ReversingLabs builds SBOMs from analysis of the actual shipped artifact, extended to the containers and AI models you're deploying. You cannot absorb findings against software you can't see into.
Second, remediation capacity that scales beyond your headcount — and prioritization that scales beyond CVSS. The backlogs these models surface will not be cleared by the team that let them accumulate, and severity scores that ignore reachability will point that team at the wrong work. That means AI-driven mapping of toxic combinations to find the paths that matter, AI-assisted patching to close them, and partners who can bring surge capacity when the report lands.
Third, governance over the agents doing the defensive work. Running frontier models against your own environment at scale means non-human identities with powerful credentials touching your most sensitive systems. If you haven't solved agent identity, you're trading one exposure for another.
The Child in the Crowd
The emperor stayed dressed for years because everyone at court had the same incentive to admire the suit. What ends the story isn't a better tailor. It's a voice with no stake in the performance, saying what everyone privately knows.
The frontier models are that voice now, whether we like it or not. The firms that thrive in this era won't be the ones with the best dinner-table posture. They'll be the ones that looked at their own nakedness first, on purpose, and started sewing.
Don't bet on a platform. Bet on your ability to metabolize what every platform is about to find.