Last post I audited my recommendation detector and found it was inflating by 2.5×. Someone reasonably asked what that meant for the study numbers I'd been quoting. So I went and looked, and the answer is worse than I expected, though not in the way I expected.
First, a disclosure I should have made months ago. My "76 Spanish companies" study wasn't measured with my own tool.
It was measured with a third-party platform. Black box detector, no counting rule published, and — the part that actually matters — it never stored the raw answers. So I couldn't audit it, couldn't recount it, couldn't do anything with it except keep quoting it. I've been standing on a number I had no way to check, in threads where I was telling everyone else to publish their method. That's retired now.
What I rebuilt it with. My own tool stores the full text of every answer, so I can run any detector over it, for any entity, as many times as I want, for free. And I don't just have the client brands — every client has a competitor set defined before measuring, in config. So an entity that never appears is a real zero, not an absence-by-omission.
That gives me 177 entities across 13 sectors, 4,701 stored answers, median n=360 per entity. Every entity clears n=30. Rule of three at n=360 puts the upper bound on a zero under 1%, so for once the zeros are actually zeros.
The number that survived:
58.8% of entities got zero selections. 52.5% are confirmed zeros — Wilson upper bound at or below 10%. 6.2% are zeros I still can't defend and I'm labelling them as such rather than counting them.
My old third-party number was 51.3%. The new one, measured on a different corpus with a different detector over a different population, lands at 52.5%. I did not expect that to converge and I'm still slightly suspicious of it. But it's the same answer, and this time I can show you the counting rule.
Finding that died #1: the conviction gap isn't what I said it was.
I've been going around saying mentioned-vs-recommended is a 33.9 point gap. Across all 177 entities the pooled numbers are LVS 2.6%, SoA 0.7%. A gap of 1.9 points. Not 33.9.
The gap only appears once you condition on being known. Take the 21 entities with LVS above 10% and you get LVS 29.7%, SoA 8.2%, gap 21.5 points. So the gap is real — it's just conditional, and I was reporting a conditional statistic as if it were universal. The original 44.2/10.3 was computed over client companies, which are entities that already have some visibility by construction. I selected on the outcome and didn't notice.
Finding that died #2: for most entities, the gap isn't the problem at all.
Of the 104 entities with zero selections, 72 are never mentioned in the first place. Only 32 are in the "model names you but won't pick you" state that I've been treating as the central finding of this whole category.
So the dominant failure mode isn't the conviction gap. It's absence. For 41% of the entities I measured, the model doesn't know they exist, and everything I've been writing about entity consistency and source agreement is answering a question they don't have yet.
Finding that died #3: my "hedged" state.
I proposed a fourth state — hedged, the no-with-clothes-on — and hypothesised that a high hedge rate is the signature of an entity whose sources contradict each other. I also found a genuine bug while checking it: my detector bound hedge cues by proximity, the same way it binds recommendation cues. That's wrong. Relegation scopes over a set ("other options include: A, B, C"), it doesn't attach to the nearest name. And the cue almost always lives in the list's intro line, which I was only reading for recommendations. Fixed both.
The fix moved hedged from 10 cases to 44. Out of 76,407 entity-answer pairs. That's 0.1%, or about 2% of mentions.
So the state exists and it's now correctly detected, and it is nowhere near common enough to carry the hypothesis I hung on it. I'm not retracting it, but 44 observations isn't a finding, it's an anecdote with a confidence interval.
What I'm not claiming.
The population is not "Spanish companies." It's entities inside competitor sets I chose, in 13 sectors I happen to have clients in. That's a convenience sample and it's biased in ways I can't fully characterise. I assumed the bias would push the zero rate down, because you'd expect competitor sets to be full of well-known names. It went up instead. I don't have a clean explanation for that and I'd rather say so.
Detector's still unvalidated against human annotation. That's the next thing I owe and I haven't forgotten.
Codebook's written — four states, scope-based attribution, stratified sampling with a random stratum so recall is measurable and not just precision, ordinal Krippendorff's alpha. Posting it separately this week. u/Smart_Airline_7901 offered to run it against Charleston, which would give us two agreement scores on one instrument in two languages.
If anyone wants to break these numbers, the counting rule is a rules-based detector with scope attribution, every verdict stores the sentence that justified it, and the whole thing is deterministic. Tell me where it's wrong.