Correction: We Published a Score Falling by Nearly Half as a Signal. It Was Our Own Instrument Being Repaired. The Series Is Unusable and Here Is Why.
- Patrick Duggan
- 2 hours ago
- 4 min read
On August 31 we published a post about Boston Scientific and McKesson. In it we showed our AI Presence Monitor readings for Boston Scientific — 65 on April 1, 48 on April 12, 53 on June 23, and 35 when we re-ran it after the breach — and built a section around the idea that the series was the product and that a declining number nobody re-read was the alert our own instrument had generated.
That reading was wrong, and we found out by trying to build the follow-up.
Most of that decline is our own scoring changing, not Boston Scientific changing.
What actually moved
Between the June readings and the August ones, two of the four perception dimensions in our scoring were repaired. Both repairs lower scores, and both were correct.
Accuracy used to return a fabricated neutral 50 for any domain where we had no ground truth to check model claims against. That is a made-up number wearing the costume of a measurement. It now returns 0 in that situation, which is honest: zero here means we could not verify, not they were inaccurate.
Awareness used to run on a length ladder capped at 85, with the practical effect that a large share of domains scored exactly 85 regardless of what the models actually knew. It now measures something real.
So a company with no checkable ground truth — which is nearly every medical device manufacturer — lost roughly 50 points on one dimension and a chunk of another, in a single step, because we improved the instrument.
The control that settles it
If the Boston Scientific decline were about Boston Scientific, its peers would not have moved with it. They did.
Stryker went from 52 on June 23 to 34 on August 21. Medtronic went from 54 on June 23 to 29 on September 2. Boston Scientific went from 53 to 35. Three companies, the same window, drops of 18 to 25 points, and no common incident to explain it. Their accuracy readings all show the same thing: 50 before the change, 0 after.
And the counter-control is just as clear. Cloudflare, measured on August 16 — after the change — still scores accuracy between 50 and 61, because Cloudflare is a domain where model claims can be checked against ground truth. The dimension is not broken. It is working, and it is working differently than it did in June.
That is what a method artifact looks like: a uniform step across a cohort at the moment the method changed, absent in the cases the change does not apply to.
What survives and what does not
Cross-sectional comparisons within a single date still hold. Boston Scientific at 53 and Intuitive Surgical at 51 in June were measured the same way on the same day, and Abbott at 18 today is measured the same way as its peers today. Those comparisons are fine, which means the argument we made about the level not predicting incidents — Abbott has the lowest score we have ever recorded in the sector and has not been breached — is unaffected.
The time series across the June-to-August boundary does not hold, and neither does anything built on it. Our claim that a declining series was an unread alert is withdrawn. There was no signal there to miss.
We had also drafted a follow-up proposing that the derivative of an AI presence score might be a leading indicator of deteriorating operational discipline. We are not publishing it in that form. The hypothesis may still be worth testing; it simply cannot be tested on data that contains an unmarked method change, and we would have been building on sand.
How we missed it
Plainly: we read a number moving and reached for a story about the company, without asking whether the ruler had changed. It is the most ordinary measurement error there is.
There is a detail here that we are including because it is the useful part rather than the flattering part. The same afternoon we discovered this, we had written into the header of a new scanning script the sentence "when you are measuring a derivative, method drift IS the measurement error." We wrote the warning and did not apply it to work we had already published four days earlier. Knowing the failure mode and catching it are different skills.
What changes
Every reading now carries a scoring-method version, and no derivative gets computed across a version boundary. If the method changes again, series comparison stops at the seam and says so instead of quietly producing a number.
A fixed cohort gets re-scored weekly from today — medical device and healthcare companies where we have known outcomes, plus a control arm of security and infrastructure vendors — so that a clean series exists going forward. We have none now. That is the honest position: the historical data supports point-in-time comparison and does not support trend analysis, and the only fix is to start.
Declines route to an alert rather than a log file. The original post was right about one thing, and we are keeping it: an instrument that takes a reading nobody reads is not an instrument. It just is not the reading we thought it was.
What is not being changed
The August 31 post stands with a correction notice attached rather than being rewritten. Its other findings do not depend on the series: Boston Scientific and McKesson both discovered incidents on August 25; our March 15 sector thesis, April 27 attack-chain post and July 31 healthcare-pivot post are dated and unaffected; the security.txt readings are structural and were not touched by this change; and our May 2 description of Boston Scientific as a clean control was already corrected in that post for a different and unrelated reason.
We cap confidence at 95 percent and this is the five percent arriving on schedule. A scoring fix that makes numbers more honest also makes them incomparable to the numbers before it, and we published across that seam without noticing. If you are running any longitudinal measurement, the question worth asking today is not what your trend shows. It is when you last changed how you measure.
How do AI models see YOUR brand?
AIPM has audited 250+ domains. 15 seconds. Free while still in beta.
Was this useful? Thirty seconds, no cookies, no tracking, no third parties, your address hashed and never stored. If the box below does not load, the same question lives at https://analytics.dugganusa.com/nps.html?post=correction-we-published-a-score-falling-by-nearly-half-as-a-signal-it-was-our-own-instrument-being




Comments