```html ```
top of page

We Read Our Own AI Report Card Out Loud. Then We Ran the Same Test on Cribl.

  • Writer: Patrick Duggan
    Patrick Duggan
  • Jun 30
  • 8 min read

Updated: 3 days ago

Microsoft started handing out report cards and most people have not noticed yet.


On February 11, 2026, Bing Webmaster Tools shipped a new section called AI Performance, in public preview. For the first time it shows publishers how often their content gets cited inside generative answers — Microsoft Copilot, the AI summaries that now sit at the top of Bing, and a handful of partner AI experiences. It surfaces the exact pages that get referenced, and it introduced a strange new unit of measurement called a grounding query: the reformulated question Copilot writes to itself, behind the scenes, when it decides it needs to go read the web before answering you. On June 16 they expanded it with intent labels, topic clusters, a Citation Share metric, and period-over-period comparison.


Translation: there is now an official, Microsoft-run scoreboard for whether the machines that are quietly replacing search can see you at all. So we did the uncomfortable thing. We pulled ours and read it out loud.



Our Report Card, Unedited


Here is aipmsec.com — our own AI-presence product, the brand whose entire job is helping companies show up correctly in AI. Bing's AI Performance export covers April 14 through June 28. Total Copilot citations in that window: eight. Seven of them landed on a single day, April 21. One more on May 13. Every other day is a zero, and the last six straight weeks — May 14 to June 28 — are an unbroken line of zeros.


Here is www.dugganusa.com, our main blog, 1,600-plus posts deep. Eleven citations total. One on April 13, one on April 29, a burst of seven on May 2, two on May 6. Then nothing. Seven and a half weeks of zero, May 7 through June 28.


That is not a typo and it is not a humble-brag with a twist ending. We sell AI-presence auditing, and Microsoft's own instrument says the large language layer barely cites us, and lately not at all. O'Toole's Axiom — Murphy was an optimist — applies to your own dashboards first.


So why publish it? Because the number that actually predicts whether a company will fix its AI visibility is not its citation count. It is whether the company is willing to look at its citation count. We look at ours. We are showing you ours. That is the entire thesis of the product in one screenshot: you cannot fix what you refuse to measure, and almost nobody is measuring this yet.



Then We Pointed The Instrument At Someone Bigger


Cribl is a very good company. It closed a 319-million-dollar Series E at a 3.5-billion-dollar valuation in August 2024, it became one of the fastest infrastructure companies ever to cross 100 million in annual recurring revenue, and it is trusted by something like half the Fortune 100. None of that is in dispute and we are not here to dispute it.


What caught our attention is the pitch. Cribl's whole 2026 narrative is about data accuracy for the AI era. Their framework states it as an equation — data integrity equals accuracy plus consistency plus context. Their messaging promises to transform raw telemetry into the AI-ready foundation your teams and your AI agents need to succeed. Their newer products lean on context-aware analysis that goes beyond simple pattern matching to ensure, in their words, precise, accurate identification. The thesis they are selling to every customer is: your data has to be accurate and richly contextual, or the AI cannot be trusted with it.


It is a good thesis. It happens to be exactly the question our AIPMSEC auditor was built to answer — just turned around and aimed back at the company selling it. So on June 30 we ran the audit. Same model council, same scoring, same day we scored ourselves. The instrument asks GPT-4o, Claude, Mistral, and DeepSeek what they actually know about a company, checks their answers against the verifiable facts, and then separately measures how machine-legible the company's own website is — because that legibility is how the models ground themselves in the first place.



A Word On How We Caught Ourselves


The first time we ran this, two of the five council seats came back empty — Gemini and Mistral both returned nothing usable for any of the three companies. We could have quietly averaged around the holes and published. Instead we treated our own instrument as the suspect, because a tool that measures honesty has no business lying to its owner. We found it: the Mistral key was misconfigured and Gemini's was dead. We fixed Mistral, re-ran the whole battery, and these are the corrected numbers from a live four-model council. Gemini stays benched until we sort its access out, and we would rather show you four working voices than five with two of them faking it. The correction mattered — it tied a race we had previously been winning. We are publishing the version that flatters us less, because that is the entire point.



The Receipts, Side By Side


Same auditor, four live models — GPT-4o, Claude, Mistral, DeepSeek — same morning. Every score is capped at 95, because we guarantee five percent of any confident-sounding number is nonsense, including ours.



Measure

dugganusa.com

aipmsec.com

cribl.io

Overall AI-perception

51

52

57

AI awareness (do models know you)

68

68

68

AIPM-NPS (would models recommend you)

minus 25

minus 25

plus 95

AI accuracy (do models get your facts right)

40

40

40

Site machine-readability

87

86

63

Schema.org structured data

85

85

5

Combined score

65

66

59


Read the top of that table honestly, because we promised honesty — and because the corrected run is harder on us than the broken one was. With a full four-model council, the recommendation gap is brutal: ask whether they would recommend Cribl for observability work and the council comes back a plus 95, near-unanimous promoters, while the same four models score us at minus 25. A minus 25 means the models either do not know us well enough to vouch for us or actively steer people elsewhere. We are a startup founded in late 2025. That is what the cold start looks like, and pretending otherwise would make the rest of this post worthless. The models will say something about all three companies — awareness lands at 68 across the board — but only Cribl has earned the recommendation. Earned. Bought with eight years and three and a half billion dollars, but earned.


On AI accuracy, the precise thing Cribl sells, the corrected council scores all three of us at 40 — a dead tie. In our first, broken run we had scored Cribl lower than ourselves and were tempted to make a meal of it; the fourth voice erased that gap, so we will not pretend it exists. The models get Cribl's facts about as right as they get ours. Fair.


Now read the bottom of the table, because that is where the conceit lives, and where the giant cannot hide behind size.


Structured data. Schema.org markup is the single most direct, machine-verifiable way for a website to tell a model this is who we are, this is what we make, in a format built specifically for machines to ground on. It is the literal accuracy-and-context layer, shipped as code on your own pages. Our score is 85. Cribl's is 5 — and that 5 is a floor, not a grade, because the auditor found exactly zero Schema.org blocks on cribl.io. We ship eight on dugganusa.com and eleven on aipmsec.com. The company selling data context to the AI era publishes a homepage with no machine-readable context at all.


A company whose pitch to customers is that raw telemetry lacks business meaning and must be enriched with metadata before a machine can trust it — service owner, criticality, region, business unit — publishes a corporate website with almost no business metadata for a machine to read. They left the AI crawlers fully welcome at the door, robots.txt wide open to GPTBot and ClaudeBot and PerplexityBot, an llms.txt file in place. And then gave those invited machines five out of ninety-five worth of structured ground truth to stand on. That is the conceit, measured: enrich your data so the AI can trust it, said the site the AI cannot accurately read.



What Both Halves Are Actually Telling Us


The Bing report and the AIPMSEC audit are two instruments pointed at the same disease, and they agree.


Bing says: even the company that sells AI presence is getting almost no Copilot citations, and so, we would wager, are you. This is not a DugganUSA problem. The generative layer is citing a vanishingly small slice of the web, and the days of zeros in our export are the days of zeros in nearly everyone's export. AI visibility in mid-2026 is a near-empty field, which means it is the most winnable field on the internet right now.


AIPMSEC says: the way you win it is not by buying awareness, it is by becoming legible to machines. Cribl bought awareness, and they earned the recommendation that comes with it — they are bigger, older, better-funded, the models know their name and will vouch for them. We have not bought any of that, and the council says so plainly, minus 25. But on the one thing that determines whether a machine can ground itself in who you actually are rather than just recognize your logo — structured data — the scrappy startup that built the auditor ships eleven blocks of machine-readable truth and the giant that preaches the gospel ships zero. Combined score, the one that weights machine-legibility into perception, lands us at 65 and 66 against Cribl's 59. We will take that trade every day, because awareness and recommendation compound slowly over years, and machine-legibility is a weekend of work.


We are not going to tell you we won. The models would recommend Cribl and would not yet recommend us, and we showed you the corrected number that makes that gap worse, not better. We are going to tell you something more useful: Microsoft just started grading this for free, the whole class is failing, and the answer key is structured data and honest measurement. We published our own failing marks — and our own instrument's failure, and the fix — to prove we read the test. Cribl, with a hundred times our resources and a thesis that is literally about machine-readable context, shipped zero machine-readable context.


Open your own Bing Webmaster Tools. Find the AI Performance tab. Read your citation count out loud. If it is a column of zeros, you are not behind — you are early, and so is almost everyone. Then go fix the layer the machines actually read. That is the whole game now, and the scoreboard finally exists.


Methodology, for the skeptics: Bing AI Performance figures are the official Citations and Cited Pages exports for each verified property, April through June 28, 2026. AIPMSEC scores come from our model-council audit run June 30, 2026, all three domains on the same configuration. Our first run that morning had two of five council seats — Gemini and Mistral — return no usable answer for any domain; we diagnosed it as a key misconfiguration in our own instrument, repaired Mistral, and re-ran the full battery. The figures above are the corrected four-model council (GPT-4o, Claude, Mistral, DeepSeek); Gemini remains benched pending access. The correction narrowed our accuracy lead over Cribl to a tie and widened Cribl's recommendation advantage — i.e. it made us look worse, which is exactly why we are publishing it instead of the first run. The structured-data and machine-readability scores do not depend on the council and did not move. All scores are capped at 95 on principle. We will happily re-run any of it on camera.




How do AI models see YOUR brand?

AIPM has audited 250+ domains. 15 seconds. Free while still in beta.



Was this useful? Thirty seconds, no cookies, no tracking, no third parties, your address hashed and never stored. If the box below does not load, the same question lives at https://analytics.dugganusa.com/nps.html?post=we-read-our-own-ai-report-card-out-loud-then-we-ran-the-same-test-on-cribl



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page