Operations

Run a certificate expiry game day before a real expiry runs one for you

October 1, 202610 min readCertPulse Engineering

Most teams have tested certificate renewal. Almost none have tested certificate expiry, and those are different exercises. Renewal testing covers the happy path: the ACME client runs, the CA issues, the new cert lands, everyone moves on. Expiry testing covers what happens when that chain quietly breaks somewhere and nobody notices until the leaf's notAfter passes and clients start refusing to connect.

I've watched the second version play out the same way more times than I'd like. Someone rotates a token and the DNS-01 solver loses its API credentials. Certbot logs a failure twice a day into a file nobody reads. The renewal window, a comfortable 30 days, burns down to zero. Then at 2am a payment webhook starts failing with "x509: certificate has expired or is not yet valid," and the on-call engineer spends the first forty minutes working out which load balancer even serves that hostname.

Big companies aren't immune. Microsoft Teams went down in February 2020 over an expired authentication certificate. In December 2018, O2 and other carriers lost mobile data because of an expired certificate in Ericsson's SGSN-MME software. When Let's Encrypt's DST Root CA X3 cross-sign expired in September 2021, plenty of older clients broke even though every server certificate involved was valid. Anyone could have predicted each of these years in advance. None of them was really tested before production found it.

Why this gets worse every year

CA/Browser Forum ballot SC-081v3 is cutting maximum TLS certificate lifetimes in stages. The 200-day limit took effect in March 2026. Next it drops to 100 days in March 2027, then to 47 days in March 2029, when domain validation reuse also shrinks to 10 days. With 398-day certs you faced expiry once a year per cert, and a lot of organizations got by on calendar reminders and luck. At 47 days, every certificate renews about eight times a year. Each of those renewals is a chance to fail.

So expiry stops being a rare event someone handles heroically. It becomes a recurring operational condition, like disk pressure or a failed deploy. You rehearse those. Run a game day.

Choosing what to break

"An expired cert" is too vague. The failures that hurt are the ones where clients, monitors, and humans each behave a little differently. Here's my minimum set.

Expired leaf. The baseline. Do clients refuse, and does anything warn a human first?

Not-yet-valid certificate. Rarer, but it happens when a freshly issued cert hits a client with a lagging clock, or when someone generates a cert with a future notBefore by mistake. It shows you which clients check both ends of the validity window.

Missing intermediate. The classic "works in my browser" bug. Browsers cover for incomplete chains with AIA fetching and cached intermediates. curl, OpenSSL, Go, Python, and most Java clients just fail. If your monitoring runs in anything browser-like, it may never see this.

Hostname mismatch. Usually a cert reissued with a changed SAN list, or traffic routed to the wrong listener after someone touched the load balancer.

Untrusted private root. A client without your root in its trust store calls a service on your internal CA. Container base image updates cause this constantly, because they can silently reset the CA bundle.

Renewal succeeds, deploy doesn't. This one causes the most pain and gets tested the least, and it isn't close. Certbot writes a new cert to disk, but no deploy hook reloads nginx. cert-manager updates the Kubernetes Secret, but the app read the cert into memory at startup and never looks again. ACM happily renews a managed cert while an imported cert on the same ALB was never going to renew at all. Your issuance metrics look great and the endpoint still serves the old cert.

Use test infrastructure, not production certs

None of this requires touching a real production certificate.

badssl.com has public hosts for most of these conditions: expired, wrong host, self-signed, untrusted root, incomplete chain, revoked. It's a good first pass for seeing how a client library or monitoring tool reacts. It tells you nothing about your own routing, alerting, or ownership, though.

The Let's Encrypt staging environment issues real ACME certificates from deliberately untrusted roots, with much looser rate limits than production. You can run the whole renewal pipeline against it, DNS challenges, deploy hooks, and reloads included, without risking production issuance. And because the staging roots are untrusted, you get an untrusted-root test case thrown in.

For expiry itself, I'd go straight to an internal step-ca instance. It issues 24-hour certificates by default and can go much shorter, and the step CLI lets you set notBefore and notAfter explicitly. Want a cert that dies in fifteen minutes, or one that isn't valid until tomorrow? Done. No clock faking. The cert really does expire on schedule.

Injecting the failure safely

Three techniques, roughly in order of how much real infrastructure they touch.

Shift time for one process

libfaketime hooks time-related libc calls through LD_PRELOAD, so you can launch a single process that thinks it's 60 days in the future. Point it at a cert that's valid today and you see exactly how it handles expiry, with nothing else on the host affected.

The catch: it only works for programs that get the time through libc. Go binaries generally use syscalls or the vDSO directly, so libfaketime does nothing to them. If your infra tooling is mostly Go, that's a big hole. Statically linked binaries ignore it too.

There's also a common belief that a container can have its own wall clock. It can't. Linux time namespaces (kernel 5.6+) offset the monotonic and boottime clocks, not CLOCK_REALTIME, and CLOCK_REALTIME is what certificate validation reads. Containers share the host's wall clock. If a whole environment needs to live on a different date, use a VM with NTP off and the clock set by hand, or run libfaketime inside the container for the processes that respect it.

Point a canary at a deliberately broken cert

Set up a dedicated hostname in non-prod, something like expiry-canary.staging.yourcompany.internal, and serve whatever broken cert you're testing: a step-ca cert that expires mid-exercise, a chain with the intermediate stripped, a cert for the wrong name. Then aim a representative set of real clients at it. Service-to-service callers, webhook dispatchers, the mobile app's staging build, your synthetic monitors.

I like this better than fake clocks because nothing is simulated. The cert is actually broken, and every client's actual TLS stack has to deal with it.

Make renewal fail for real

To test the "renewal silently failed for three weeks" path, break renewal in non-prod and leave it broken. Block egress to the ACME directory endpoint at the firewall or with a DNS override. Better yet, revoke the DNS-01 solver's credentials, since that's closer to how it fails in real life. Pair it with short-lived step-ca certs so the gap between "renewal starts failing" and "cert expires" is hours, not weeks.

Then watch. Does anything alert on the first failed renewal attempt, or only once the cert is already dead? Those are two separate alerts. You need both.

Blast radius and rollback

Write these down before you start:

  • Scope: the exact hostnames, namespaces, and clients involved. Anything not on the list is out of scope, and if it breaks, that's a finding and your cue to stop.
  • Abort criteria, e.g. "any change in production 5xx rate" or "any page to a team that isn't participating."
  • Rollback: a known-good cert and the command to redeploy it, ready before you break anything. For blocked ACME endpoints, have the firewall rule ID or DNS override on hand. Do a dry run of the rollback first so you know it actually works.
  • A heads-up to adjacent teams and your security team. A weird internal hostname in CT logs or a sudden spike in TLS errors is exactly what they should investigate, and you don't want them burning a morning on your game day.

What to watch during the exercise

Give one person the job of watching and taking notes. Only that. They don't fix anything. You're measuring behavior, not proving you can eventually get it working.

Start with how clients fail. Some fail closed right away with a clear error. Some retry with exponential backoff for hours and look like a latency problem. And some have InsecureSkipVerify, or the equivalent, buried in a config someone copied from a 2019 Stack Overflow answer, so they never fail at all. That last group is a security finding, not a reliability one, and honestly it's often the most valuable thing the whole exercise turns up.

Health checks are where most teams get surprised. AWS ALB health checks against HTTPS targets don't validate the target's certificate. Kubernetes HTTPS liveness and readiness probes skip verification too. Your dashboards can be solid green while every real client hits handshake failures. Note which checks noticed and which didn't.

Then alerts. Did the expiry-threshold alert fire, and when? Did a renewal-failure alert fire on the first failed attempt? Where did it go: a Slack channel with 400 unread messages, an email alias forwarding to someone who left last year, or an actual pager?

Last, time to owner and time to runbook. Start a stopwatch when the first alert fires. How long until someone knows which team owns the cert? How long until they find a runbook, and does it describe how the cert is issued today? I've seen runbooks pointing to a certbot cron job on a host that had moved to cert-manager two years earlier. If finding the owner means digging through Slack history, that number is your result.

Turning findings into fixes

The scorecard

For each injected failure, record four yes/no answers with timestamps:

  1. Detected: did any system notice before a user would have?
  2. Alerted: did that detection produce an alert at the right severity?
  3. Routed: did the alert reach a human who owns the cert and can act on it?
  4. Resolved within target: did they fix it in your target time using the documented runbook?

Most first scorecards are patchy. Expired leaf detected and alerted, but routed to the wrong team. Missing intermediate never detected. Deploy failure invisible until clients broke. That's fine. The point is to get the gaps written down. Every "no" becomes a ticket with an owner and a due date.

How often to rerun

At minimum, before each SC-081v3 phase: once before March 2027 and again before March 2029. Shorter lifetimes squeeze every renewal window and break timing assumptions that used to be safe. Renewing at 30 days remaining is fine on a 398-day cert. On a 47-day cert it means renewing after just 17 days, which is a completely different pattern, and ACME Renewal Information (ARI, RFC 9773) may move your renewal times around anyway.

Beyond those dates, I'd run it quarterly, plus after any reorg, any certificate tooling migration (certbot to cert-manager, imported certs to ACM-managed), and any change to on-call routing. Ownership drift is the most common reason certs expire, and reorgs are when ownership drifts hardest.

Check that your external monitoring actually catches it

Internal checks only see what they're pointed at, so test whatever sits outside your infrastructure as well. If you use an external endpoint monitor, add the canary hostnames during the game day and confirm it flags every injected condition, not just expiry but chain completeness and hostname mismatch too. With CertPulse, for example, you'd add the canary endpoints, set alert rules for expiry thresholds and chain problems, and confirm the notification lands with the same person your internal alerts should have reached. Check the scan frequency while you're in there. A daily scan can take up to 24 hours to catch a cert that broke right after the last check, so make sure the interval matches the response time on your scorecard.

Game days tend to expose a blind spot in internal alerting, and the external check is there to cover it. It doesn't excuse your internal alerting from working.

The short version

Pick six failure modes. Break them on purpose in non-prod using step-ca short-lived certs, Let's Encrypt staging, and badssl.com, never real production certs. Watch clients, health checks, alerts, and people, and time every step. Score detected, alerted, routed, resolved. Fix the "no"s, then run it again before the 100-day phase lands.

It costs you an afternoon. Skip it and an expired cert will eventually run the exercise for you, probably at 2am.

This is why we built CertPulse

CertPulse connects to your AWS, Azure, and GCP accounts, enumerates every certificate, monitors your external endpoints, and watches Certificate Transparency logs. One dashboard for every cert. Alerts when auto-renewal fails. Alerts when certs approach expiry. Alerts when someone issues a cert for your domain that you didn't request.

If you're looking for complete certificate visibility without maintaining scripts, we can get you there in about 5 minutes.

Run a certificate expiry game day before a real expiry runs one for you | CertPulse