Triage Playbook

Microsoft Entra ID: Integrations & Access Troubleshooting

Course 1 · Chapter 10 · Capstone: A Triage Playbook

Nine chapters of parts; this one is the whole machine. It joins everything into a single path from "the client says they can't sign in" to root cause, fix, client message and prevention, shows it working on three realistic cases, and ends with the artefacts your team can actually use: a flowchart, an evidence checklist, communication templates, an escalation checklist and a one-page runbook.

The cases are fictional, and so is the product
All three cases, organisations, IDs and log lines are invented, built from behaviour documented in the earlier chapters. Your real error and your four open questions (which integration type you use, which arrangement, what access you have, where "validate the connection" runs) were never supplied to the course. Exercise 2 and the closing section show how to fill those gaps in yourself, which is what makes the playbook yours.

The Playbook at a Glance

Read the tree top to bottom. Each branch names the chapter that explains it.

CLIENT CAN'T SIGN IN / INTEGRATION FAILING
│
├── 0. PRESERVE: save the error text, time + zone, a log download      (Ch 7, 8)
│
├── 1. WHO is affected, and since when?                                 (Ch 6 table)
│   ├── everyone, started at one moment ...... credentials / disabled app / policy / consent
│   ├── some users ........................... assignment (nested groups?) / group-targeted policy
│   ├── only new users ....................... user consent disabled or tightened
│   ├── only guests .......................... guest account / cross-tenant / untrusted MFA
│   ├── one place or network ................. location Conditional Access
│   ├── background/API only .................. app permission / workload policy / secret
│   └── right after OUR release .............. new permission requested
│
├── 2. RIGHT TENANT? RIGHT APP?  tenant card, client ID, both objects   (Ch 1, 5)
│   └── five-minute inspection: Enabled? Assignment required? credentials? URIs match?
│
├── 3. THE LOG: which sign-in tab? filter by correlation ID             (Ch 7)
│   ├── 7000222 / 7000215 / 50012 / 700027 / 700030 ... CREDENTIAL      -> Ch 8 fix
│   ├── SAML assertion rejected / cert date passed ..... SAML CERT      -> Ch 8 fix C
│   ├── 65001 / 90094 / 65004 ......................... CONSENT         -> Ch 6
│   ├── 50105 ......................................... ASSIGNMENT      -> Ch 6
│   ├── 53003 / 50076 / 50079 / 70043 ................. CONDITIONS      -> Ch 6
│   ├── 50020 / 50034 / 500011 / 700016 / 90002 ....... TENANT / ACCOUNT / APP MISSING
│   ├── 50011 ......................................... REDIRECT URI    -> Ch 3, 5
│   ├── 7000112 ....................................... APP DISABLED    -> Ch 5
│   ├── 700082 / 70008 / 50173 ........................ TOKEN, not integration (Ch 4)
│   └── nothing found ................................. check other tabs, narrower window, request path
│
├── 4. CONFIRM IN THREE PLACES: log, portal, audit log                  (Ch 8)
│   └── four stories: expired / renewed one side / deleted / not the credential
│
├── 5. FIX IN SAFE ORDER: create, install, PROVE, remove                (Ch 8)
│   └── adjust the rule for the integration, never switch security off  (Ch 6)
│
├── 6. TELL THE CLIENT at diagnosis, plan, done                         (Ch 5-9 boxes)
│
└── 7. PREVENT: register, owner, recipients, lead times, product        (Ch 9)

The Evidence You Record Every Time

From Chapters 3, 7 and 8, one list. Copy it into the ticket before any change is made, and keep it in text so it can be searched later (redact tenant IDs, client IDs and names before it leaves your team):

GroupFields
The errorerror name; error_description; AADSTS number; trace ID; correlation ID; timestamp (and time zone)
Who and whattenant ID (tenant card); application (client) ID; resource; user or service principal; user type (member or guest)
Howpattern (OIDC, SAML, client credentials, SCIM); arrangement A, B or C; authentication requirement; client app
Entra stateEnabled for users to sign in?; Assignment required?; credentials with expiry dates and Key IDs; permissions granted; Conditional Access result and policy names
Changesaudit entries for the relevant days: who, what, when

Case 1: The Overnight Job That Stopped

Fictional setting. Fabrikam's nightly sync into your product stopped on 2 October. Their users can still sign in. Arrangement B (the registration is in Fabrikam's tenant), client-credentials pattern.

StepFindingDecision
0 PreserveJob log shows invalid_client, AADSTS7000222, 02:00:04 UTC, with trace and correlation IDs. You download the day's service-principal sign-insEvidence safe
1 WhoOnly the background job; interactive sign-in is fineCredential / permission / workload policy suspected. Not assignment or user consent (no user involved)
2 Right appTenant ID matches the tenant card; client ID matches. Enterprise app "Enabled for users to sign in? = Yes"Not Chapter 5's master switch
3 LogService principal sign-ins, Failure, 7000222, Conditional Access Not appliedCredential family, and not a policy
4 ConfirmPortal: secret "prod 2025" expired 2026-09-28; no newer secret. Audit: no changesStory 1: simply expired
5 FixClient administrator creates "prod 2026-10" (under 12 months), enters the Value into your product, you prove with a fresh Success, remove the old secret a week laterCreate, install, prove, remove
6 Tell"The credential expired on 28 September; nothing wrong with your users or tenant; replacement takes ten minutes."Diagnosis, plan, done messages
7 PreventRegister row with owner and 90-day flag; shared mailbox; 14-day install stepChapter 9

Case 2: SAML Sign-in Failing for Everyone

Fictional setting. From Monday 5 December every Contoso user gets a SAML error page in your product. Arrangement C. Your product holds one signing certificate.

StepFindingDecision
0 PreserveThe product's error says the assertion signature was rejected; user reports begin 08:40Save text and times
1 WhoEveryone, starting at one moment, a date that looks roundExpiry or a change; test the certificate date first
2 Right appEnterprise app Single sign-on > SAML Certificates: the Active certificate expired 2026-12-04; an Inactive certificate valid to 2029 existsStrong lead
3 LogInteractive sign-ins show Entra issuing tokens; the failure is on your side. (Entra records the sign-in; your product rejects the assertion)Not a Chapter 6 gate
4 ConfirmAudit log: the Inactive certificate was created in August by the former administrator. Your product still trusts only the expired certificate. Story 2: renewed on one side onlyFix C
5 FixAgree a short window; download the 2029 certificate, upload to your product, Make certificate active back to back; test user signs in; remove the old certificateSingle-certificate product means a pair of steps
6 Tell"Your signing certificate expired on 4 December; a replacement already existed but our system hadn't been given it. We'll swap it in now, which takes about 15 minutes."Plain statement, no blame
7 PreventSupport multiple certificates and read the federation metadata URL at least every 24 hours; notification list to a distribution list; register entryThe strongest fix: remove the human step

Case 3: The Integration Is Healthy, but Access Was Removed

Fictional setting. After Northwind's security review on Thursday, a number of staff can no longer open your product. The credentials are fine and nothing expired. Arrangement C (SSO), OIDC for the API.

StepFindingDecision
0 PreserveUsers report an "Approval required" screen (new starters) or an error naming no assigned role (existing staff)Two symptoms, so two causes
1 WhoSome existing users; all new usersSome: assignment. New: user consent
2 Right appProperties: Enabled = Yes; Assignment required = Yes (was No last week). Users and groups: only "Support Staff" assignedA setting, not a credential
3 LogExisting users: AADSTS50105. The new-starter screen is the admin consent workflow ("Approval required")Chapter 6, gates 2 and 1
4 ConfirmAudit log, Thursday afternoon: Update service principal (assignment required changed) and a change to the tenant's user-consent settingStory 4: not the credential. Do NOT rotate anything
5 FixTheir administrator assigns the "Sales" and "Operations" groups (directly, since nested groups don't inherit), and either grants admin consent for the app's permissions or names reviewers for the consent workflowAdjust the rule for the integration, do not switch the controls off
6 Tell"Your tenant now restricts who may use applications and who can approve them. That is a sound control; we just need the right groups and approver on the list."Treat the change as legitimate
7 PreventAdd "assignment and consent settings" to the onboarding checklist; ask the client to tell us before security reviews that change themChapter 9

Client Communication Template

A reusable shape for the three stages from Chapter 8, plus the update rhythm in between:

1. ACKNOWLEDGE (within the hour)
   "We have your report and are investigating. First, can you confirm the time it started and
   whether it affects everyone?"

2. DIAGNOSIS (when the log confirms it)
   "We found the cause: <one plain sentence, e.g. the credential used by the connection expired
   on <date>>. Your users' accounts and your tenant are not at fault. It is fixed by <action>."

3. PLAN
   "We'd like <role> to <action> (about <n> minutes). Nothing is removed until we've confirmed the new
   one works. We never need a password or secret value."

4. UPDATES (at the interval you promised)
   "Status at <time>: <done / waiting on / next step>. Next update at <time>."

5. DONE
   "Sign-in has worked again since <time>. We'll remove the old credential on <date> and have added
   a reminder for <next expiry> to the shared mailbox you chose."

DON'T: blame a named person; promise a cause before the log confirms it; ask for a secret;
       suggest turning off a security control; use acronyms the reader hasn't met.

Escalation Checklist

Escalate with the evidence list already filled in, so the next person doesn't start from nothing.

Escalate toWhenSend
Our engineering teamThe failure is in our product: assertion rejected, credential not used, validation message wrong, cached secret, no multiple-certificate supportError text, version, tenant card, steps, logs from our side, the portal facts (dates, thumbprints)
The client's Entra administratorA setting or credential in their tenant must change: assignment, consent, policy, new credentialSpecific page names (Chapter 5), exactly what to change and why, and the safe order
The client's security teamConditional Access, cross-tenant or consent policy is blocking the integrationPolicy name from the log's Conditional Access tab, the integration's fixed IP or identity, and a narrow request
Microsoft support (via the client, unless we hold support rights in the tenant)Credentials are valid, consent and assignment are in order, no policy applies, yet Entra still returns an unexpected failureTenant ID, application ID, correlation ID and request ID, timestamp with time zone, error code and trace ID, what was changed and when, and a statement of what has been ruled out. Microsoft's guidance for unresolved sign-in problems is to open a support request
Our incident processMany clients are affected at once, a credential may be exposed, or a security control looks bypassedFollow the incident workflow from the Incident Response & Ticketing Workflows course; treat any leaked secret as compromised and rotate it

The One-Page Runbook

RUNBOOK: <PRODUCT> CAN'T SIGN IN VIA MICROSOFT ENTRA ID               Owner: <team>   Reviewed: <date>

1. PRESERVE      Error text, time + zone, download Entra logs (Free keeps 7 days; P1/P2 30).
2. IDENTIFY      Tenant card: client, tenant ID, app (client) ID, arrangement A/B/C, pattern.
3. WHO FAILS     All / some / new / guests / one place / background / after release  -> likely gate.
4. LOOK          Sign-in log (right tab) by correlation ID. Copy: code, reason, app, resource, user, CA result.
5. INSPECT       Right tenant? Enterprise app: Enabled? Assignment required? Credentials + expiry vs today?
                 URIs/thumbprints match our side? Permissions granted? (Read-only; change nothing yet.)
6. CONFIRM       Log + portal + audit log agree on ONE cause. If not, keep looking (never "rotate and hope").
7. FIX           Credential:   create new -> install -> PROVE -> remove old.   Secrets only via our product.
                 SAML cert:    New Certificate (Inactive) -> upload to us -> Make certificate active -> test.
                 SCIM:         new token -> Secret Token -> Test Connection -> check provisioning log.
                 Consent / assignment / policy: client admin changes; adjust, never disable.
8. PROVE         New Success in the right sign-in tab; Key ID/thumbprint matches; real data flows; check next day.
9. COMMUNICATE   Acknowledge / diagnosis / plan / updates / done (templates). No blame, no secrets.
10. RECORD       Symptom, evidence, cause, fix, proof, prevention date, owner.
11. PREVENT      Register row, owner + shared mailbox, 90/60/30/14/7 reminders, notification list verified.

ESCALATE        Product fault -> Engineering | Setting/credential -> client Entra admin | Policy -> client security
                Unexplained after 1-6 -> Microsoft support with IDs | Many clients/exposure -> incident process

NEVER           Ask for a secret by email/chat. Remove the old credential before proving the new. Delete an app or
                service principal to "tidy". Ask a client to switch off a security policy. Rotate a healthy secret.

Where This Meets the Rest of the Site

This course is the Entra-specific layer on top of general support skills taught elsewhere on the site: tokens, sessions and OAuth in Authentication & Session Security; certificates in HTTPS/TLS Fundamentals; reading logs in Logging & Log Analysis; diagnosing a failing application in Web & Application Troubleshooting; ticket handling in Incident Response & Ticketing Workflows; wording for non-technical readers in Customer/User Communication for Support; turning a procedure into a team asset in Documentation & Runbooks; and the habits that keep secrets safe in Security Basics for Support Technicians.

Make It Yours: Three Tasks

  1. Answer the four open questions for your own systems and write them at the top of the runbook: which integration type(s) you use (SAML, OIDC with a secret or certificate, SCIM, or another); whether you use arrangement A, B or C (or a mixture); what access you normally have in a client's tenant (none, guest, a read role, or screen share only); and whether the "validate the connection" check is in your product or in Microsoft's portal.
  2. Collect three real, redacted errors from past tickets and file each under the playbook's log-code branches. Anything that doesn't fit shows a gap worth fixing.
  3. Fill the first ten rows of the expiry register from your busiest clients, using Chapter 9's columns, and agree the owner for each.

Course Complete: What You Can Now Do

ChapterYou can now…
1. What Entra ID IsExplain tenants, the rename from Azure AD, and whose side owns what; keep a tenant card
2. How Applications IntegrateTell an app registration from an enterprise application; recognise OIDC, SAML, client credentials and SCIM, and arrangements A, B and C
3. What Passes Between Entra and Your SystemFollow a sign-in step by step, read token claims, and know where each step fails
4. Credentials and ExpiryDistinguish credential expiry from token expiry; know lifetimes, the two-sided trap and where to look
5. Finding the IntegrationSwitch to the right tenant, find both objects, know your role's reach, and run the five-minute inspection
6. Permissions and ConsentSeparate permission, assignment and Conditional Access causes; read the consent and policy errors
7. Reading the Sign-in LogsPick the right log and tab, filter by correlation ID, know retention, and capture evidence
8. Fixing an Expired CredentialDiagnose in three places, fix each credential type safely and prove it
9. Preventing It Happening AgainRun a register, use Entra's notices, choose lifetimes, and make rotation routine
10. CapstoneTriage any "can't sign in" report with one playbook, communicate it, escalate it and document it

Entra changes often. The concepts in this course (identity provider, two objects, credentials versus tokens, three gates, evidence first, create-install-prove-remove) will outlast the menu names; keep Microsoft Learn open when you need the current spelling of a portal path or a figure, and update your runbook when it changes.

Hands-On Exercises

All three use fictional organisations and values. Never put real secrets, tokens or tenant details into practice notes.

Exercise 1

Run the playbook on a new case. Tailspin's partner firm Litware has guests in Tailspin's tenant who use your product via SSO. Since Monday morning, Litware's staff (and only Litware's staff) get blocked. Tailspin's own staff are fine. Two of the Litware guests' sign-in log entries show AADSTS53003 and the Conditional Access tab lists a policy "Require compliant device" = Failure; a third shows AADSTS50020. Work through every branch: who is affected, the likely gates, what you check in the portal, the story, the fix, the message to Tailspin and the prevention step. State which parts you can be confident about and which need more evidence.

📄 View solution
Exercise 2

Personalise the runbook. Write answers to the four open questions as if for a real or realistic product (or the fictional "Acme Support Desk": multitenant app, OIDC and SAML, a vendor support role with screen share only, "validate the connection" inside the product's admin page). Then edit three steps of the one-page runbook so they match: step 2 (what the tenant card holds), step 5 (what you can inspect yourself and what you must ask for), and step 7 (the fixes you actually perform).

📄 View solution
Exercise 3

A diagnosis turns out wrong. In Case 1 you told Fabrikam "the secret expired", sent the diagnosis message, and they created a new secret. The job still fails, now with AADSTS7000215 (invalid secret). Write the correction message, state what you got wrong and what the evidence now says, and list what you will change in your process so a wrong early diagnosis is less likely. Then re-order the evidence you should have gathered before sending the first message.

📄 View solution

Chapter 10 Quick Reference

  • Path: preserve, who is affected, right tenant and app, the log, confirm in three places, fix in safe order, tell the client, prevent
  • Code families: 7000222/7000215/50012/700027/700030 credential; 65001/90094/65004 consent; 50105 assignment; 53003/50076/50079/70043 conditions; 50020/50034/500011/700016/90002 tenant, account or app missing; 50011 redirect URI; 7000112 disabled app; 700082/70008/50173 token
  • Four stories: expired, renewed one side only, deleted/replaced, not the credential (then do not rotate anything)
  • Safe order: create, install, prove, remove; SAML: new cert Inactive, upload to us, Make certificate active, test
  • Escalate with the evidence list: product fault to engineering; setting or credential to the client's administrator; policy to their security team; unexplained to Microsoft with tenant, app, correlation and request IDs
  • Never: ask for a secret by email or chat, remove before proving, delete to tidy, ask to disable security, rotate a healthy secret
  • Prevent: register, owners and shared mailbox, 90/60/30/14/7 reminders, verified notifications, product that accepts two credentials and reads federation metadata
  • Make it yours: answer the four open questions, file three real errors, fill ten register rows