All lessons Leer en español

Security in depth · Unit 29

Incident response: decisions under uncertainty

Connect preparation, analysis, containment, recovery, and learning.

11 minready

Helpful before thisDefenses and detectionLearn with care: permission, people, and AI

After this lesson you can

  • Apply stated incident criteria while keeping cause and data-disclosure claims separate.
  • Choose an authorized containment action using evidence, continuity and reversibility.
  • Write recovery gates and a useful handoff that distinguish partial restoration from closure.

Lessons in this unit

Browse 9 lessons in this topic
  1. Digital evidence: the story and its limitsUnderstand integrity, provenance, timelines, and confidence without overclaiming.11 min
  2. Tabletop: practice a security incident togetherRun a fictional scenario that reveals decision, communication, and recovery gaps.12 min
  3. Triage is a decision under uncertaintyPrioritize an unresolved account alert without declaring travel or compromise proven.4 min
  4. Containment changes the situationChoose a proportionate response using a supplied service-impact comparison.4 min
  5. Share the facts each audience needsPrepare a leadership update that supports a decision without inventing a cause.4 min
  6. A timestamp needs clock contextCompare event-time ranges before claiming which event happened first.3 min
  7. Some evidence disappears as systems changeBalance a temporary memory record against an explicitly documented harm deadline.4 min
  8. Define what makes recovery completeDecide whether a reachable service has met its agreed return-to-operation checks.4 min
  9. Case: the booking service is downSeparate service impact from compromise claims, then decide what counts as a verified recovery.10 min

A service shows unusual account activity. Nobody yet knows whether it is a mistake, a compromised account, or a broader incident. A response plan helps the team reduce uncertainty while protecting people and services.

Containment: Actions to limit continuing harm during an incident.

Prepare and understand → Contain and investigate → Recover and improve1Prepare and understand2Contain and investigate3Recover and improve
Response activities overlap and inform each other; they are not a rigid one-way checklist.

Preparation is part of response capability

Know the systems, owners, important data, and dependencies before a crisis. Establish authority to isolate systems, revoke access, contact providers, and approve restoration. Keep a communication path available when normal identity or messaging systems fail.

An event is an observation. Declaring an incident is a decision based on the organization’s criteria and available evidence. Record what is known, what is suspected, and what would change the assessment.

Contain with an explicit tradeoff

Containment limits continuing harm. Disabling an account or isolating a device may protect other systems, but may also interrupt essential work or change evidence. Urgency, safety, business impact, and investigative needs must inform the authorized decision.

Assign an incident lead and keep a decision log with time, owner, reason, and result. Work can proceed in parallel, but conflicting changes need coordination. Avoid claiming a root cause merely because an early clue looks persuasive.

Recovery needs verification and follow-through

Removing a visible symptom is not proof that the underlying cause is addressed. Validate the restoration plan, access changes, system integrity, and monitoring before normal operation resumes. Continue watching for recurrence.

Communicate verified facts, uncertainty, and the next update time to the people who need them. Notification obligations depend on the situation and require appropriate legal or contractual review. Afterwards, turn lessons into owned improvements with dates rather than blame.

Worked response: a service can be open and still wrong

A fictional community booking service treats confirmed record changes outside approved business rules as an integrity incident, whether caused by misuse, error or a defect. The response lead may pause its import job; normal service restoration requires the owner’s approval. All times below use one agreed clock.

  • I1, 14:00: The owner confirms twenty-four booking records changed outside the approved rules. The responsible session and cause are unknown. No evidence of data disclosure is supplied.
  • I2, 14:03: The same import path continues creating incorrect changes. A pause would interrupt automatic intake; an approved manual queue can safely handle thirty minutes of demand.
  • I3: The lead is authorized to pause that job. Independent audit records and the pre-change snapshot are already retained. The packet gives no evidence that a whole-service shutdown would preserve more relevant material.
  • I4, 14:07: After the authorized pause, the observation record shows no further unwanted writes. The public page remains available.
  • I5, 14:12: Twenty affected records have been reconciled; four remain unverified. Normal operation requires all twenty-four reconciled, corrected import behavior verified, fresh monitoring and owner approval. The next stakeholder update is due at 14:15.
PredictAt 14:12 the page loads and harmful writes have stopped. Is that enough to declare recovery complete?

No. It supports observed containment and availability, not I5’s full recovery conditions. Four records remain unverified and the corrected import path has not been demonstrated. The team can communicate progress without granting normal-operation approval prematurely.

Run investigation and continuity together

Open the incident under the stated integrity criterion. Keep competing causes visible: a legitimate account can run a defective job, and a familiar process name does not establish authorization for every resulting change. Investigators should preserve which records support each hypothesis and what additional context would distinguish them.

The lead’s pause addresses continuing harm while using the approved manual route. Assign someone to monitor the queue’s thirty-minute limit; continuity capacity is a condition to track, not an unlimited promise. Record the pause’s time, authority, reason, expected consequence and observed result so another responder can continue.

The 14:15 update should state confirmed integrity impact, successful observed containment, the four outstanding reconciliations, unknown cause and the next decision. Keep restricted evidence with its custodians and tailor the communication to the audience’s need.

Model recovery gate: reopen the import only after the stated evidence and approval exist. A separately authorized degraded mode could be considered, but this packet grants none. After recovery, assign follow-up work for the failed control and the response lessons without turning an unproven causal theory into the final explanation.

EXPLORE THE CONCEPT

Choose the next decision

A fictional account performs an unexpected sensitive action.

Document the event and assess current harm

Establish evidence and urgency. Escalate according to defined criteria instead of assuming either harmlessness or catastrophe.

Immediately announce a confirmed data breach

An unusual event alone does not prove a breach. Communicate what is verified and what remains uncertain.

Restore service without checking the cause

Availability matters, but an unresolved cause may recreate the incident. Validate recovery conditions and monitoring.

A simplified learning model. It connects to no systems and uses no real data.

Turn the idea into a decision

Good response reduces harm and uncertainty together. Record decisions so another responder can understand and continue the work.

Terms you met

Containment

Check yourself

No timer. No penalties. Read the explanation and try again whenever you like.

  1. What incident statement follows from I1 and the stated policy?

    Show the answer

    Correct answer: Open an integrity incident because confirmed unauthorized record changes meet the policy criterion; cause and disclosure remain unknown. The policy defines the trigger and I1 supplies it. Declaring an incident does not require inventing an attacker or claiming data left the service.

  2. Which immediate decision best fits I2-I3?

    Show the answer

    Correct answer: Have the authorized lead pause the import job, retain the independent evidence, and use the approved manual queue while monitoring its limits. This acts on continuing harm within the supplied authority and preserves the stated continuity option. Its effectiveness still needs observation.

  3. What does I4’s successful pause establish?

    Show the answer

    Correct answer: The observed unwanted writes stopped after the pause; it does not prove the cause is understood or the affected records repaired. A containment result answers whether continuing harm stopped under observation. Cause and data integrity are separate recovery questions.

  4. Which recovery decision is supported at 14:12?

    Show the answer

    Correct answer: Keep normal-operation approval pending until all twenty-four affected records and the corrected import behavior meet I5’s criteria. The page and twenty checked records are partial evidence. Four records and the corrected path remain unverified, so the normal gate is not met.

Try it

  • WriteWrite a handoff at 14:12 from I1-I5: confirmed impact, unknown cause, incident lead, containment decision, observed result, recovery gate and next update. Model the status as an integrity incident under investigation, not a confirmed disclosure; do not close on the restored page alone.
References