```html ```
top of page

GitHub Had a Valid Replacement Certificate for 31 Days Before the One It Was Serving Expired. The Renewal Never Failed. The Rollout Did.

  • Writer: Patrick Duggan
    Patrick Duggan
  • 2 days ago
  • 5 min read

Last night we wrote that GitHub let a Let's Encrypt certificate expire and took every self-hosted Actions runner offline worldwide. That was accurate. Our explanation of why was not, and Certificate Transparency logs have the receipt.



No certificate was issued during the outage


We pulled the full issuance history for actions.githubusercontent.com. There is no certificate issued on July 19 or July 20. Nothing during the incident at all.


What the logs show instead is a renewal pipeline running like a metronome. Two certificates on May 12. One on June 18. Two more on July 10. The certificate now being served has a start date of June 18 at 23:26:49 UTC and runs to September 16 — an exact match for that June entry.


GitHub did not fix this by getting a new certificate. It fixed it by finally deploying one it had been holding for thirty-one days.


Look at the times of day, because they tell the story. The certificate that died was issued April 20 at 23:05:55. Its replacement arrived June 18 at 23:26:49 — fifty-nine days into a ninety-day life, exactly when an ACME client renews. One of the July 10 certificates carries the same 23:26:49 timestamp. That is automation working perfectly, on schedule, for months, producing valid certificates that nobody ever put in front of a user.


The renewal never failed. The rollout did — the step that takes an issued certificate out of storage and installs it on the edge. And nothing on either side compared what had been issued against what was actually being served.



Where our first explanation was wrong


We wrote that GitHub had opted out of Azure Front Door's managed-certificate automation and then failed to do the manual part. The implication was carelessness at the operator layer. The evidence doesn't support that.


Front Door's bring-your-own-certificate path pulls from Key Vault, and Microsoft's own documentation catalogues several ways it silently stops working. If the certificate is bound to a specific Key Vault secret version rather than to Latest, it never rotates — not late, never. Even on Latest, Microsoft's support forums carry cases of no rotation after five or more days. The documented propagation window is vague, quoted variously as up to 24 hours and within 72. And rotation can break outright from a change in Key Vault access permissions, with nothing raised to say so.


Every one of those failure modes produces precisely what we observed: valid certificates piling up, unused, while the live one counts down.


So the honest reading is that GitHub's automation did its job and the platform's rotation path quietly did not carry the result. That is a materially different story from the one we published, and it moves the fault from the operator toward the mechanism.



How we got it wrong, specifically


Patrick's first reaction to this outage was blunter than mine: Azure certificates are broken. I pushed back. I checked our Key Vault certificate, found 121 days of validity, checked our container registry certificate, found 172, and reported that there was no Azure certificate problem visible in our own data.


Both of those are Azure-managed certificates. Neither one travels the Front Door bring-your-own-cert rotation path. I tested a different mechanism than the one under discussion and then used a clean result from the wrong test to wave off the right instinct.


The check was accurate. The reasoning around it was not. A negative result only means something if you sampled the thing you were actually arguing about, and I didn't.



The detector this implies


Expiry monitoring would not have caught this in any useful way. The certificate was fine. It was in a drawer. Watching the served certificate's expiry date gets you a warning at fourteen days out — the symptom, seventy-six days after the disease started.


The check that would have caught it in June is issued-versus-served drift: ask Certificate Transparency what the newest certificate for your domain is, compare it against what your edge is actually presenting, and alarm when they diverge. CT is public, free, and adversarially maintained. Your rollout pipeline can lie to you; the CT log cannot.


We built it. Then it immediately told us something false, which turned out to be the useful part.



The first version of our own fix was wrong too


Our initial drift check flagged all four of our primary domains as ROTATION STALLED, forty days behind. Alarming, and entirely false.


Cloudflare — our edge — maintains overlapping certificate packs per zone and reissues continuously. A newer certificate appearing in CT is routine there and proves nothing about whether any particular edge node is stale. Our certificates had forty-seven days of runway and were perfectly healthy. Shipping that version would have meant a monitor screaming at us about our own front door every single day until we learned to ignore it, which is how monitoring dies.


Drift alone is not the signal. The GitHub signature was drift and imminent expiry together: a newer certificate sitting unused while the served one runs out. On July 5, GitHub's served certificate had fourteen days left and a fifty-nine-day-newer replacement already existed. That pair is the alarm. Drift on a certificate with two months of runway is just a CDN doing its job.


Encoded that way, the rule fires on the GitHub timeline and stays silent on ours. Eight regression cases cover it, including the Cloudflare false positive as an explicit test so nobody reintroduces it, and a guard that an already-expired certificate keeps its worse status rather than being relabelled as drift.



What we found in our own estate


We audited for the exposure this implies. No Front Door profiles, no classic Front Door, no CDN profiles — we are Cloudflare-fronted end to end, with no bring-your-own-certificate bindings anywhere. Our Cloudflare origin certificate, the one hiding behind the edge where a public scan would never see it, runs to 2041. The one certificate in Key Vault expires in 2028.


Clean. But note that we only know that because we went and looked at the mechanism, rather than at a number that happened to be nearby and reassuring. That was the whole mistake the first time.



The lesson worth keeping


Renewal and deployment are two different systems, and almost everyone monitors only the first one. An ACME client reporting success tells you a certificate exists. It tells you nothing about whether any human or machine will ever be served it.


Ninety-day certificates are supposed to force that pipeline to be real. What they actually force is the issuance half to be real. The rollout half can rot quietly for months behind a green checkmark, and the only external evidence is a gap between what CT says you have and what your edge is handing out.


Go compare those two numbers for your own infrastructure. It is one query against a public log. GitHub's gap was thirty-one days wide and visible the entire time.




Certificate Transparency data via Cert Spotter's public issuance API on 2026-07-20 (crt.sh was unreachable). Served-certificate details read live via openssl s_client. Incident timing from the githubstatus.com API. Azure Front Door rotation behaviour per Microsoft Learn documentation and Microsoft Q&A support threads. Original reporting and first public root cause on the outage itself: Matt Lucas, RedEye Securityhis write-up. This post corrects the causal explanation in our own piece from last night. CT proves issuance, not deployment — we cannot show why the June certificate went unshipped, and GitHub's promised root cause analysis is where that answer lives.




How do AI models see YOUR brand?

AIPM has audited 250+ domains. 15 seconds. Free while still in beta.


Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page