Triage Playbook
Microsoft Entra ID: Integrations & Access Troubleshooting
Course 1 · Chapter 10 · Capstone: A Triage Playbook
Nine chapters of parts; this one is the whole machine. It joins everything into a single path from "the client says they can't sign in" to root cause, fix, client message and prevention, shows it working on three realistic cases, and ends with the artefacts your team can actually use: a flowchart, an evidence checklist, communication templates, an escalation checklist and a one-page runbook.
The Playbook at a Glance
Read the tree top to bottom. Each branch names the chapter that explains it.
CLIENT CAN'T SIGN IN / INTEGRATION FAILING │ ├── 0. PRESERVE: save the error text, time + zone, a log download (Ch 7, 8) │ ├── 1. WHO is affected, and since when? (Ch 6 table) │ ├── everyone, started at one moment ...... credentials / disabled app / policy / consent │ ├── some users ........................... assignment (nested groups?) / group-targeted policy │ ├── only new users ....................... user consent disabled or tightened │ ├── only guests .......................... guest account / cross-tenant / untrusted MFA │ ├── one place or network ................. location Conditional Access │ ├── background/API only .................. app permission / workload policy / secret │ └── right after OUR release .............. new permission requested │ ├── 2. RIGHT TENANT? RIGHT APP? tenant card, client ID, both objects (Ch 1, 5) │ └── five-minute inspection: Enabled? Assignment required? credentials? URIs match? │ ├── 3. THE LOG: which sign-in tab? filter by correlation ID (Ch 7) │ ├── 7000222 / 7000215 / 50012 / 700027 / 700030 ... CREDENTIAL -> Ch 8 fix │ ├── SAML assertion rejected / cert date passed ..... SAML CERT -> Ch 8 fix C │ ├── 65001 / 90094 / 65004 ......................... CONSENT -> Ch 6 │ ├── 50105 ......................................... ASSIGNMENT -> Ch 6 │ ├── 53003 / 50076 / 50079 / 70043 ................. CONDITIONS -> Ch 6 │ ├── 50020 / 50034 / 500011 / 700016 / 90002 ....... TENANT / ACCOUNT / APP MISSING │ ├── 50011 ......................................... REDIRECT URI -> Ch 3, 5 │ ├── 7000112 ....................................... APP DISABLED -> Ch 5 │ ├── 700082 / 70008 / 50173 ........................ TOKEN, not integration (Ch 4) │ └── nothing found ................................. check other tabs, narrower window, request path │ ├── 4. CONFIRM IN THREE PLACES: log, portal, audit log (Ch 8) │ └── four stories: expired / renewed one side / deleted / not the credential │ ├── 5. FIX IN SAFE ORDER: create, install, PROVE, remove (Ch 8) │ └── adjust the rule for the integration, never switch security off (Ch 6) │ ├── 6. TELL THE CLIENT at diagnosis, plan, done (Ch 5-9 boxes) │ └── 7. PREVENT: register, owner, recipients, lead times, product (Ch 9)
The Evidence You Record Every Time
From Chapters 3, 7 and 8, one list. Copy it into the ticket before any change is made, and keep it in text so it can be searched later (redact tenant IDs, client IDs and names before it leaves your team):
| Group | Fields |
|---|---|
| The error | error name; error_description; AADSTS number; trace ID; correlation ID; timestamp (and time zone) |
| Who and what | tenant ID (tenant card); application (client) ID; resource; user or service principal; user type (member or guest) |
| How | pattern (OIDC, SAML, client credentials, SCIM); arrangement A, B or C; authentication requirement; client app |
| Entra state | Enabled for users to sign in?; Assignment required?; credentials with expiry dates and Key IDs; permissions granted; Conditional Access result and policy names |
| Changes | audit entries for the relevant days: who, what, when |
Case 1: The Overnight Job That Stopped
Fictional setting. Fabrikam's nightly sync into your product stopped on 2 October. Their users can still sign in. Arrangement B (the registration is in Fabrikam's tenant), client-credentials pattern.
| Step | Finding | Decision |
|---|---|---|
| 0 Preserve | Job log shows invalid_client, AADSTS7000222, 02:00:04 UTC, with trace and correlation IDs. You download the day's service-principal sign-ins | Evidence safe |
| 1 Who | Only the background job; interactive sign-in is fine | Credential / permission / workload policy suspected. Not assignment or user consent (no user involved) |
| 2 Right app | Tenant ID matches the tenant card; client ID matches. Enterprise app "Enabled for users to sign in? = Yes" | Not Chapter 5's master switch |
| 3 Log | Service principal sign-ins, Failure, 7000222, Conditional Access Not applied | Credential family, and not a policy |
| 4 Confirm | Portal: secret "prod 2025" expired 2026-09-28; no newer secret. Audit: no changes | Story 1: simply expired |
| 5 Fix | Client administrator creates "prod 2026-10" (under 12 months), enters the Value into your product, you prove with a fresh Success, remove the old secret a week later | Create, install, prove, remove |
| 6 Tell | "The credential expired on 28 September; nothing wrong with your users or tenant; replacement takes ten minutes." | Diagnosis, plan, done messages |
| 7 Prevent | Register row with owner and 90-day flag; shared mailbox; 14-day install step | Chapter 9 |
Case 2: SAML Sign-in Failing for Everyone
Fictional setting. From Monday 5 December every Contoso user gets a SAML error page in your product. Arrangement C. Your product holds one signing certificate.
| Step | Finding | Decision |
|---|---|---|
| 0 Preserve | The product's error says the assertion signature was rejected; user reports begin 08:40 | Save text and times |
| 1 Who | Everyone, starting at one moment, a date that looks round | Expiry or a change; test the certificate date first |
| 2 Right app | Enterprise app Single sign-on > SAML Certificates: the Active certificate expired 2026-12-04; an Inactive certificate valid to 2029 exists | Strong lead |
| 3 Log | Interactive sign-ins show Entra issuing tokens; the failure is on your side. (Entra records the sign-in; your product rejects the assertion) | Not a Chapter 6 gate |
| 4 Confirm | Audit log: the Inactive certificate was created in August by the former administrator. Your product still trusts only the expired certificate. Story 2: renewed on one side only | Fix C |
| 5 Fix | Agree a short window; download the 2029 certificate, upload to your product, Make certificate active back to back; test user signs in; remove the old certificate | Single-certificate product means a pair of steps |
| 6 Tell | "Your signing certificate expired on 4 December; a replacement already existed but our system hadn't been given it. We'll swap it in now, which takes about 15 minutes." | Plain statement, no blame |
| 7 Prevent | Support multiple certificates and read the federation metadata URL at least every 24 hours; notification list to a distribution list; register entry | The strongest fix: remove the human step |
Case 3: The Integration Is Healthy, but Access Was Removed
Fictional setting. After Northwind's security review on Thursday, a number of staff can no longer open your product. The credentials are fine and nothing expired. Arrangement C (SSO), OIDC for the API.
| Step | Finding | Decision |
|---|---|---|
| 0 Preserve | Users report an "Approval required" screen (new starters) or an error naming no assigned role (existing staff) | Two symptoms, so two causes |
| 1 Who | Some existing users; all new users | Some: assignment. New: user consent |
| 2 Right app | Properties: Enabled = Yes; Assignment required = Yes (was No last week). Users and groups: only "Support Staff" assigned | A setting, not a credential |
| 3 Log | Existing users: AADSTS50105. The new-starter screen is the admin consent workflow ("Approval required") | Chapter 6, gates 2 and 1 |
| 4 Confirm | Audit log, Thursday afternoon: Update service principal (assignment required changed) and a change to the tenant's user-consent setting | Story 4: not the credential. Do NOT rotate anything |
| 5 Fix | Their administrator assigns the "Sales" and "Operations" groups (directly, since nested groups don't inherit), and either grants admin consent for the app's permissions or names reviewers for the consent workflow | Adjust the rule for the integration, do not switch the controls off |
| 6 Tell | "Your tenant now restricts who may use applications and who can approve them. That is a sound control; we just need the right groups and approver on the list." | Treat the change as legitimate |
| 7 Prevent | Add "assignment and consent settings" to the onboarding checklist; ask the client to tell us before security reviews that change them | Chapter 9 |
Client Communication Template
A reusable shape for the three stages from Chapter 8, plus the update rhythm in between:
1. ACKNOWLEDGE (within the hour)
"We have your report and are investigating. First, can you confirm the time it started and
whether it affects everyone?"
2. DIAGNOSIS (when the log confirms it)
"We found the cause: <one plain sentence, e.g. the credential used by the connection expired
on <date>>. Your users' accounts and your tenant are not at fault. It is fixed by <action>."
3. PLAN
"We'd like <role> to <action> (about <n> minutes). Nothing is removed until we've confirmed the new
one works. We never need a password or secret value."
4. UPDATES (at the interval you promised)
"Status at <time>: <done / waiting on / next step>. Next update at <time>."
5. DONE
"Sign-in has worked again since <time>. We'll remove the old credential on <date> and have added
a reminder for <next expiry> to the shared mailbox you chose."
DON'T: blame a named person; promise a cause before the log confirms it; ask for a secret;
suggest turning off a security control; use acronyms the reader hasn't met.
Escalation Checklist
Escalate with the evidence list already filled in, so the next person doesn't start from nothing.
| Escalate to | When | Send |
|---|---|---|
| Our engineering team | The failure is in our product: assertion rejected, credential not used, validation message wrong, cached secret, no multiple-certificate support | Error text, version, tenant card, steps, logs from our side, the portal facts (dates, thumbprints) |
| The client's Entra administrator | A setting or credential in their tenant must change: assignment, consent, policy, new credential | Specific page names (Chapter 5), exactly what to change and why, and the safe order |
| The client's security team | Conditional Access, cross-tenant or consent policy is blocking the integration | Policy name from the log's Conditional Access tab, the integration's fixed IP or identity, and a narrow request |
| Microsoft support (via the client, unless we hold support rights in the tenant) | Credentials are valid, consent and assignment are in order, no policy applies, yet Entra still returns an unexpected failure | Tenant ID, application ID, correlation ID and request ID, timestamp with time zone, error code and trace ID, what was changed and when, and a statement of what has been ruled out. Microsoft's guidance for unresolved sign-in problems is to open a support request |
| Our incident process | Many clients are affected at once, a credential may be exposed, or a security control looks bypassed | Follow the incident workflow from the Incident Response & Ticketing Workflows course; treat any leaked secret as compromised and rotate it |
The One-Page Runbook
RUNBOOK: <PRODUCT> CAN'T SIGN IN VIA MICROSOFT ENTRA ID Owner: <team> Reviewed: <date>
1. PRESERVE Error text, time + zone, download Entra logs (Free keeps 7 days; P1/P2 30).
2. IDENTIFY Tenant card: client, tenant ID, app (client) ID, arrangement A/B/C, pattern.
3. WHO FAILS All / some / new / guests / one place / background / after release -> likely gate.
4. LOOK Sign-in log (right tab) by correlation ID. Copy: code, reason, app, resource, user, CA result.
5. INSPECT Right tenant? Enterprise app: Enabled? Assignment required? Credentials + expiry vs today?
URIs/thumbprints match our side? Permissions granted? (Read-only; change nothing yet.)
6. CONFIRM Log + portal + audit log agree on ONE cause. If not, keep looking (never "rotate and hope").
7. FIX Credential: create new -> install -> PROVE -> remove old. Secrets only via our product.
SAML cert: New Certificate (Inactive) -> upload to us -> Make certificate active -> test.
SCIM: new token -> Secret Token -> Test Connection -> check provisioning log.
Consent / assignment / policy: client admin changes; adjust, never disable.
8. PROVE New Success in the right sign-in tab; Key ID/thumbprint matches; real data flows; check next day.
9. COMMUNICATE Acknowledge / diagnosis / plan / updates / done (templates). No blame, no secrets.
10. RECORD Symptom, evidence, cause, fix, proof, prevention date, owner.
11. PREVENT Register row, owner + shared mailbox, 90/60/30/14/7 reminders, notification list verified.
ESCALATE Product fault -> Engineering | Setting/credential -> client Entra admin | Policy -> client security
Unexplained after 1-6 -> Microsoft support with IDs | Many clients/exposure -> incident process
NEVER Ask for a secret by email/chat. Remove the old credential before proving the new. Delete an app or
service principal to "tidy". Ask a client to switch off a security policy. Rotate a healthy secret.
Where This Meets the Rest of the Site
This course is the Entra-specific layer on top of general support skills taught elsewhere on the site: tokens, sessions and OAuth in Authentication & Session Security; certificates in HTTPS/TLS Fundamentals; reading logs in Logging & Log Analysis; diagnosing a failing application in Web & Application Troubleshooting; ticket handling in Incident Response & Ticketing Workflows; wording for non-technical readers in Customer/User Communication for Support; turning a procedure into a team asset in Documentation & Runbooks; and the habits that keep secrets safe in Security Basics for Support Technicians.
Make It Yours: Three Tasks
- Answer the four open questions for your own systems and write them at the top of the runbook: which integration type(s) you use (SAML, OIDC with a secret or certificate, SCIM, or another); whether you use arrangement A, B or C (or a mixture); what access you normally have in a client's tenant (none, guest, a read role, or screen share only); and whether the "validate the connection" check is in your product or in Microsoft's portal.
- Collect three real, redacted errors from past tickets and file each under the playbook's log-code branches. Anything that doesn't fit shows a gap worth fixing.
- Fill the first ten rows of the expiry register from your busiest clients, using Chapter 9's columns, and agree the owner for each.
Course Complete: What You Can Now Do
| Chapter | You can now… |
|---|---|
| 1. What Entra ID Is | Explain tenants, the rename from Azure AD, and whose side owns what; keep a tenant card |
| 2. How Applications Integrate | Tell an app registration from an enterprise application; recognise OIDC, SAML, client credentials and SCIM, and arrangements A, B and C |
| 3. What Passes Between Entra and Your System | Follow a sign-in step by step, read token claims, and know where each step fails |
| 4. Credentials and Expiry | Distinguish credential expiry from token expiry; know lifetimes, the two-sided trap and where to look |
| 5. Finding the Integration | Switch to the right tenant, find both objects, know your role's reach, and run the five-minute inspection |
| 6. Permissions and Consent | Separate permission, assignment and Conditional Access causes; read the consent and policy errors |
| 7. Reading the Sign-in Logs | Pick the right log and tab, filter by correlation ID, know retention, and capture evidence |
| 8. Fixing an Expired Credential | Diagnose in three places, fix each credential type safely and prove it |
| 9. Preventing It Happening Again | Run a register, use Entra's notices, choose lifetimes, and make rotation routine |
| 10. Capstone | Triage any "can't sign in" report with one playbook, communicate it, escalate it and document it |
Entra changes often. The concepts in this course (identity provider, two objects, credentials versus tokens, three gates, evidence first, create-install-prove-remove) will outlast the menu names; keep Microsoft Learn open when you need the current spelling of a portal path or a figure, and update your runbook when it changes.
Hands-On Exercises
All three use fictional organisations and values. Never put real secrets, tokens or tenant details into practice notes.
Run the playbook on a new case. Tailspin's partner firm Litware has guests in Tailspin's tenant who use your product via SSO. Since Monday morning, Litware's staff (and only Litware's staff) get blocked. Tailspin's own staff are fine. Two of the Litware guests' sign-in log entries show AADSTS53003 and the Conditional Access tab lists a policy "Require compliant device" = Failure; a third shows AADSTS50020. Work through every branch: who is affected, the likely gates, what you check in the portal, the story, the fix, the message to Tailspin and the prevention step. State which parts you can be confident about and which need more evidence.
📄 View solutionPersonalise the runbook. Write answers to the four open questions as if for a real or realistic product (or the fictional "Acme Support Desk": multitenant app, OIDC and SAML, a vendor support role with screen share only, "validate the connection" inside the product's admin page). Then edit three steps of the one-page runbook so they match: step 2 (what the tenant card holds), step 5 (what you can inspect yourself and what you must ask for), and step 7 (the fixes you actually perform).
📄 View solutionA diagnosis turns out wrong. In Case 1 you told Fabrikam "the secret expired", sent the diagnosis message, and they created a new secret. The job still fails, now with AADSTS7000215 (invalid secret). Write the correction message, state what you got wrong and what the evidence now says, and list what you will change in your process so a wrong early diagnosis is less likely. Then re-order the evidence you should have gathered before sending the first message.
📄 View solutionChapter 10 Quick Reference
- Path: preserve, who is affected, right tenant and app, the log, confirm in three places, fix in safe order, tell the client, prevent
- Code families: 7000222/7000215/50012/700027/700030 credential; 65001/90094/65004 consent; 50105 assignment; 53003/50076/50079/70043 conditions; 50020/50034/500011/700016/90002 tenant, account or app missing; 50011 redirect URI; 7000112 disabled app; 700082/70008/50173 token
- Four stories: expired, renewed one side only, deleted/replaced, not the credential (then do not rotate anything)
- Safe order: create, install, prove, remove; SAML: new cert Inactive, upload to us, Make certificate active, test
- Escalate with the evidence list: product fault to engineering; setting or credential to the client's administrator; policy to their security team; unexplained to Microsoft with tenant, app, correlation and request IDs
- Never: ask for a secret by email or chat, remove before proving, delete to tidy, ask to disable security, rotate a healthy secret
- Prevent: register, owners and shared mailbox, 90/60/30/14/7 reminders, verified notifications, product that accepts two credentials and reads federation metadata
- Make it yours: answer the four open questions, file three real errors, fill ten register rows