entra1-9 Exercise 3: Post-Incident Review ========================================================= Source: Microsoft Learn pages on SAML certificate management and ISV certificate rotation guidance; course chapters 4, 8 and 9. WHAT WENT WRONG (facts) - Entra's 60/30/7-day emails went to the administrator who added the application, who had left three months earlier. - Nobody else was watching the date. - The product accepts one certificate only, so the swap needed downtime. - The fix was done on a weekday, in a four-hour outage. SYSTEMIC FIXES What our team controls 1. Expiry register covering every SAML certificate, with an internal owner and a 90-day flag, so we do not rely on the client's mailbox. 2. Product change: accept more than one signing certificate, and read the Entra federation metadata URL regularly (at least every 24 hours), promoting the new certificate when it is activated. Removes downtime and removes the need for coordination. 3. Show the SAML certificate expiry date and the certificate in use in the product's settings page, with product-side warnings at 60/30/7 days sent to named client contacts and to our support mailbox. 4. A runbook with the create-install-prove-remove steps for SAML, and a scheduled maintenance window template for products that must swap. 5. Onboarding checklist item: record the client owner, a shared mailbox and notification addresses at set-up time. What we can only request of the client 6. A distribution list (up to five addresses) as the SAML notification recipient, verified in the portal after being set. 7. A named owner on each side, updated when people leave (leaver checklist includes "integrations they own"). 8. Agree a renewal date and window in advance, with an approver. 9. Optional: send logs to Log Analytics and alert on integration sign-in failures, so the first failure is noticed within minutes. TWO WITH THE BIGGEST EFFECT (a) Fix 2 (multiple certificates plus metadata auto-rollover). It removes the downtime and, if the client simply activates the new certificate, removes the need for our involvement. It turns a risky event into a routine one, and it works even when every human process fails. (b) Fix 1 (our own register with an internal owner and a 90-day flag). It removes the dependence on someone else's mailbox, so even a client who has lost track gets warned by us in time. It is cheap and fully ours. Runner-up: Fix 6 (distribution list), cheap and effective but depends on the client and can silently decay again. WHY THIS WORKS AS AN ANSWER --- Each fix addresses a cause rather than blaming a person, and the ranking is based on how much each reduces both likelihood and impact.