In the previous lesson, analysing the consultancy's breach and responding to the e-mail provider's, you twice needed something Nimbus still does not have: an incident log, a team with roles, a decision on who is told and within what deadline, and a sequence agreed in advance. Go back for a moment to the first-72-hours table in 02-06: day 20, 09:00, Marta calls the team together and nobody knows what to do first. Disconnect? Warn the customers? Ring whom? That gap is not filled by the best risk register or the most complete control catalogue. It is filled by a document written beforehand, rehearsed and reachable when everything else is encrypted. This lesson is that document: the team, the severity matrix, preserving evidence, the real trade-offs of containment, the communication templates, the runbooks and the post-mortem.
Contents
- Why the plan is written beforehand: event, incident and crisis
- The NIST SP 800-61 response cycle
- Preparation: 80 % of the outcome
- Classification and severity
- Detection and analysis, and preserving evidence
- Containment: the real trade-offs
- Eradication and recovery
- Communication: half the job
- The incident log
- Runbooks by scenario
- The blameless post-mortem
- Tabletop exercises
- Why the plan is written beforehand: event, incident and crisis
At three in the morning, with the diaries of 40 clinics at a standstill and a phone that will not stop ringing, nobody improvises well. The ability to reason calmly collapses exactly when it is most needed, and the decisions of those first minutes — shutting down a server, deleting a suspicious file, answering a journalist — condition everything that follows, including the possibility of finding out afterwards what happened. The plan is not there to give you brilliant answers: it is there so that the important decisions have already been taken by the time you have to execute them.
| Concept | Definition | Nimbus example | Who decides |
|---|---|---|---|
| Event / alert | An observable occurrence; an alert if it meets a criterion and deserves review | A spike in 500 errors; 200 failed sign-in attempts in 5 minutes | It is logged; the alert is looked at by whoever is on call |
| Incident | An event that compromises, or may compromise, confidentiality, integrity or availability | An administrative account signs in from an unknown IP at 04:00 | The incident manager |
| Crisis | An incident that threatens business continuity, reputation or legal position | The ransomware in 02-06 | Management |
The difference between an incident and a crisis is not one of technical size, it is one of who has to be in the room. An incident is handled by the technical team with a manager; a crisis calls for management, probably legal and communications advice, and decisions that are not technical — whether to notify, whether or not to pay, whether or not to speak publicly. Escalating late from incident to crisis is one of the most expensive and most frequent failures.
- The NIST SP 800-61 response cycle
flowchart LR
P["1. PREPARATION\nTeam, contacts, runbooks,\ntools, training.\nDone BEFOREHAND"]
D["2. DETECTION AND ANALYSIS\nTriage, scope, severity.\nWhat we know vs what we assume"]
C["3. CONTAINMENT, ERADICATION\nAND RECOVERY\nStop, clean up, come back"]
A["4. POST-INCIDENT ACTIVITY\nBlameless post-mortem,\ncorrective actions into the\nrisk register (04-01)"]
P --> D --> C --> A
A -->|"improves preparation"| P
C -.->|"new finding:\nscope is reassessed"| D
Two observations about the diagram. First: preparation is 80 % of the outcome. Everything that can be done in the cold light of day — deciding who is in charge, having the phone numbers, knowing where the logs are, having tested a restore — multiplies the effectiveness of the other three phases; what is not done beforehand will not be improvised during. Second: the arrow back from containment to detection is not decorative. In a real incident new things are discovered constantly — a second compromised server, one more account — and every finding forces a reassessment of the scope. A team that treats the analysis as closed too early contains half the problem and leaves the attacker inside.
- Preparation: 80 % of the outcome
3.1 Nimbus's response team
Four functions, always with a deputy, because the incident will happen in August:
| Function | Primary | Deputy | Responsibility |
|---|---|---|---|
| Incident manager (decides) | Marta (CTO) | Iván | Declares the incident, sets the severity, authorises disruptive actions, decides whether to escalate to a crisis |
| Technical lead (executes) | Lucía | Iván | Contains, investigates, eradicates and recovers |
| Communication | Rubén | Marta | Talks to customers; channels enquiries; nobody else replies |
| Documentation | Sara | Rubén | Keeps the incident log in real time |
| External support | Forensic retainer · Legal advisers · Insurer · Cloud provider | Activated according to severity |
Three team rules. Whoever executes does not decide: Lucía should not be weighing up whether to warn customers while she is isolating a server; separating the two functions avoids the classic error of the busiest person taking the most important decisions. One single voice outwards: if three people reply to customers, there will be three versions and one of them will be wrong. And the documenter does nothing else: it looks like a luxury in a company of 38 people, but without them nobody will remember 48 hours later what was done at 04:12.
3.2 Out-of-band contacts
The plan cannot live only in the system that may be encrypted. In 02-06 the attacker reached the cloud account, the buckets and the backups; if the plan, the phone numbers and the runbooks had been solely in the corporate office suite, the team would have been left with no plan exactly when it needed one.
| Item | Where it lives as well as in the main system |
|---|---|
| Response plan and runbooks | Printed PDF in the office + a copy on Marta's and Lucía's phones |
| Team and external phone numbers, and the cloud provider's with the account number | Printed card in the wallet; without console access you will see none of that |
| Emergency communication channel | A group on a messaging app different from the corporate one |
| Emergency (break-glass) credentials | Sealed envelope in the safe, with a usage procedure and a log |
3.3 Forensic retainer, legal contact and tools
Nimbus has no in-house forensic capability and should not pretend it has. What it can have, and cheaply, is a retainer agreement signed in advance with a forensic provider: negotiating price and terms on the day of the incident is slow and expensive. The same goes for legal advisers and for the insurer's contact, whose phone number must be on the printed card because, as you saw in 04-01, many policies require notification within a short deadline and the use of their own experts. As for technical preparation, the minimum viable: centralised logs off the machines that generate them, with 90 days of hot retention and a year cold; a read-only user for investigating without modifying; a clean analysis machine; and break-glass credentials tested at least once a year. And the minimum training: that the whole team knows two things — how an incident is declared and who to call — and that Lucía and Iván have executed at least one runbook in a tabletop exercise.
- Classification and severity
| Level | Objective criteria | Response time | Who is activated |
|---|---|---|---|
| S1 Critical | Customer data compromised or exfiltrated; service down for everyone; ransomware; compromise of the cloud account (A-05) | 15 min to start, 24/7 | Full team + management + externals. It is a crisis |
| S2 High | Compromise of a privileged account; improper access to a customer's data; prolonged partial outage; confirmed malware on an endpoint with access | 1 h in working hours, 4 h outside | Manager + technical lead + communication |
| S3 Medium | Phishing with credentials handed over but no evidence of use; exploitable critical vulnerability not exploited; loss of an encrypted laptop | 4 h working hours | Technical lead |
| S4 Low | Phishing reported and blocked; failed access attempt; external scanning | 1 working day | Whoever detects it |
Three classification rules that prevent pointless arguments. When in doubt, escalate: lowering the severity afterwards costs one e-mail, raising it late costs days of attacker dwell time, and nobody will be reprimanded for having declared an S2 that turned out to be an S3. Severity is reviewed continuously: the incident in 02-06 started as "a performance issue" that Rubén was handling as an S4, and the initial severity is almost never the final one. And uncertainty about scope raises the severity, it does not lower it: "we do not know whether data left" is treated as S1 until proven otherwise, precisely because not being able to determine the scope was the central problem of day 22 in 02-06.
- Detection and analysis, and preserving evidence
5.1 Where incidents come from
Five sources, ordered from best to worst. Your own alert is the ideal one and the one Nimbus barely has (04-03). An employee often raises the alarm, and how quickly depends entirely on reporting not being punished (POL-04 5.3.1). A customer, which was the detection route in 02-06, twenty days late. An external notice from a researcher, the national CERT or a supplier. And the attacker themselves, with the ransom note or the extortion e-mail: the worst possible way to find out.
5.2 Triage: what we know and what we are assuming
The most valuable discipline of the first few hours is separating facts from hypotheses in the incident log, explicitly, because in a real incident assumptions turn into facts through repetition: somebody says "it looks like they came in through the VPN", half an hour later it is "they came in through the VPN" and two hours later the VPN is being rebuilt while the attacker is still inside by another route.
FACTS (verified, with evidence and time)
- 04:12 Account consultora-01 signs in from 203.0.113.45 (cloud log)
- 04:31 Access key AKIA...7Q created for backup-svc (cloud audit)
- 05:02 Outbound transfer of 42 GB from the attachments bucket (metrics)
ASSUMPTIONS (to be confirmed: who checks it and by when)
- Entry with leaked credentials, not by exploitation [Lucia, 12:00]
- The production database has not been accessed [Ivan, 12:00]
- Only tenant 118 is affected [Ivan, 14:00]
RULED OUT (with the reason)
- VPN access: there is no VPN session in the window (checked 09:40)5.3 Preserve evidence before touching anything
Every containment action destroys information: shutting a machine down wipes the memory, restarting a container removes the ephemeral file system and rotating a credential prevents you seeing what else it was doing. Hence the rule — capture first, act afterwards — following the order of volatility, in which whatever disappears soonest is collected first.
| Order | Evidence | Disappears on |
|---|---|---|
| 1 | RAM, processes and active connections | Shutting down or restarting |
| 2 | Routes, ARP, sessions and connected users | Restarting the network or the machine |
| 3 | Temporary files and the ephemeral file system | Restarting the container |
| 4 | Disk and local logs | Reinstalling |
| 5 | Centralised logs, backups and cloud audit | Rotating (7 days in 02-06) or deletion |
#!/usr/bin/env bash
# initial-collection.sh - Basic collection on a Nimbus Linux server.
# Run BEFORE containing and WITHOUT restarting or shutting down. This is NOT
# forensic analysis: in an S1 or S2 this is the only thing that gets touched
# before the forensic team arrives. No -e: if one command fails, we carry on
# collecting the rest.
set -uo pipefail
CASE="${1:?Usage: $0 <case-id>}"
DEST="/mnt/evidence/${CASE}/$(hostname)-$(date -u +%Y%m%dT%H%M%SZ)"
mkdir -p "${DEST}" # /mnt/evidence: an EXTERNAL volume, not the local disk
# 1. Time context: without the exact time and the offset, the timeline is worthless.
{ date -u; timedatectl 2>/dev/null; uptime; } > "${DEST}/00-time.txt"
# 2. VOLATILE FIRST: processes, connections and deleted files still open
# (that last one is a classic of malware that removes itself from disk).
ps auxwwf > "${DEST}/01-processes.txt"
ss -tunapo > "${DEST}/02-connections.txt" # includes the PID of each socket
lsof -nP +L1 > "${DEST}/03-deleted-open.txt"
# 3. Sessions and users: who is here and who has been.
{ who -a; last -Fw | head -50; } > "${DEST}/04-sessions.txt"
# 4. Persistence: where an attacker hooks in to survive a reboot.
{ crontab -l 2>/dev/null; ls -la /etc/cron.*/ /var/spool/cron/ 2>/dev/null
systemctl list-units --type=service --state=running
ls -la /etc/systemd/system/; cat /root/.ssh/authorized_keys 2>/dev/null
} > "${DEST}/05-persistence.txt"
# 5. Copy of the local logs BEFORE they rotate.
tar czf "${DEST}/06-logs.tar.gz" /var/log 2>/dev/null || true
# 6. INTEGRITY: hash everything collected and the listing itself. Without this,
# the evidence is arguable before an insurer or in a later claim.
( cd "${DEST}" && sha256sum ./* > SHA256SUMS ) > "${DEST}/07-integrity.txt"
# 7. CHAIN OF CUSTODY: who, when, by what method and on which machine.
printf 'Case: %s\nMachine: %s\nCollected by: %s\nDate UTC: %s\nMethod: initial-collection.sh v1 (no shutdown)\n' \
"${CASE}" "$(hostname)" "${SUDO_USER:-$USER}" "$(date -u +%FT%TZ)" \
> "${DEST}/08-custody.txt"
echo "Evidence in ${DEST}. Do NOT modify anything else on this machine."Four decisions in the script that are matters of method, not of programming. It does not shut the machine down, because shutting it down destroys the most valuable evidence and, with some ransomware, can make the state of the encryption worse. It collects onto an external volume, not onto the compromised disk, so as not to overwrite unallocated space. It calculates hashes, because evidence without verifiable integrity is arguable before an insurer or a court. And it documents the chain of custody: who collected it, when, by what method and where it is. When to stop and call a forensic specialist: stop any further handling and call the retainer if there are signs of personal data being exfiltrated; the attacker has had administrative access; there is ransomware; the incident may end in a legal claim, a police report or an insurance claim; or you simply do not know what you are looking at. The instinct to "investigate just a bit more" is understandable and is exactly what destroys the evidence you later need to determine the scope — and without the scope you cannot notify accurately, which was the deadlock of day 22 in 02-06.
- Containment: the real trade-offs
Containment aims to stop the harm without destroying the information you need in order to understand the scope, and it splits into two: short-term containment, with immediate measures to stop the bleeding — isolating the machine from the network while keeping it powered on, blocking an IP, disabling an account, cutting an integration — and long-term containment, which lets you keep operating while you eradicate: standing up a clean environment in parallel, adding filtering rules, restricting non-essential functionality. The three trade-offs you have to have thought through in the cold light of day:
| Trade-off | In favour of acting now | In favour of waiting | Nimbus's criterion |
|---|---|---|---|
| Isolate now versus observe | Stops exfiltration in progress | Observing reveals the full scope and other entry routes | If there is active exfiltration or encryption in progress, isolate immediately. In any other case, up to 60 min of observation with the manager's authorisation |
| Cutting access tips off the attacker | You throw them out | It may make them accelerate and destroy, like the deletion of the backups on day 20 | Cut everything at once in a coordinated action, never piecemeal |
| Shutting down versus keeping powered on | Stops the encryption | Shutting down destroys the memory and may make decryption impossible | Do not shut down: isolate from the network while keeping it powered on |
And the catalogue of typical containment decisions at Nimbus, with what you have to know about each one:
| Decision | How | Watch out for |
|---|---|---|
| Isolate a machine | A network rule leaving it with no inbound or outbound traffic, except the analyst's | Do not shut it down or restart it |
| Revoke sessions and tokens | Rotate the signing key and withdraw the compromised kid (03-06) |
It invalidates every user's session: people have to be warned |
| Rotate credentials | Secrets manager; rotate the ones the attacker could have seen too | Rotate in the right order so you do not cut off your own access |
| Block an IP or disable an integration | Blocklist at the edge; revoke the third party's API key | The IP is ephemeral; cutting the integration may stop functionality and is the manager's decision |
| Disable an account | Disable, never delete | Deleting it destroys the evidence and its history |
- Eradication and recovery
Eradicating means removing the attacker's presence and the cause that let them in: closing the vulnerability, removing the persistence detected in step 4 of the script, withdrawing the accounts and keys they created and correcting the configuration that made it possible. The most important rule of the section applies here: when in doubt, rebuild rather than clean up. If a server has had an attacker with administrative privileges on it, you cannot prove you have cleaned it completely; you can only prove you have built a new one. At Nimbus this is especially cheap because the infrastructure is containers and automated deployment: rebuilding from the verified image and code is faster than auditing a suspect system, and it ends in certainty rather than in hope. Naturally, rebuilding only helps if you do not reintroduce the vulnerability: first you eradicate the cause, then you rebuild. And going back to production requires five conditions, written down so that commercial pressure cannot cut them short: root cause identified and corrected; no signs of persistence; potentially exposed credentials rotated; restored data verified (04-06); and enhanced monitoring for at least two weeks, because attackers come back, and they come back the same way if nobody closed it.
- Communication: half the job
| Recipient | When | Who | What is said |
|---|---|---|---|
| Internal team and management | Immediately (S1/S2) | Incident manager | What is known, what is not known, what each person is doing, what must NOT be done; to management, business impact and the decisions needed |
| Affected customers | As soon as there is a useful fact to communicate; without waiting to know everything | Rubén, with approved wording | What has happened, what affects them, what they should do, when there will be more information |
| All customers | If there is an impact on the service | Rubén | Status and outlook |
| Supervisory authority | If there is a personal data breach: 72 h from becoming aware | Marta with advice | Notification in accordance with the applicable procedure |
| Police and insurer | A criminal offence (ransomware, extortion, fraud); the policy, often 24-72 h | Marta with advice | Police report; opening the claim |
| Suppliers involved | Immediately if they are part of it | Lucía | Request for information and for action |
| Media | Only if they ask | Marta, nobody else | Prepared statement |
Validation note. The 72-hour deadline for notifying a personal data breach to the supervisory authority, the obligation to communicate to those affected, the report to law enforcement and the implications of paying a ransom — which may have legal consequences depending on the jurisdiction, the identity of the attacker and international sanctions law — are legal questions. Nimbus does not decide them in the crisis room: it consults legal advisers and the compliance lead from the first hour. The detail of the obligations is covered in 06-03.
8.1 Template: notice to affected customers
Subject: Important information about the security of your Nimbus Reservas account
Dear [NAME OF CENTRE] team,
We are writing to inform you of a security incident that may affect your centre's
data on our platform. We would rather tell you now, even though the investigation
is still under way, than wait until we have every answer.
WHAT HAS HAPPENED
On [DATE] at [TIME] we detected [ONE-SENTENCE DESCRIPTION, NO JARGON].
Our team acted immediately and [CONTAINMENT MEASURE ALREADY APPLIED].
WHAT INFORMATION IS INVOLVED
Based on what we know at this point, [CONFIRMED DATA CATEGORIES].
[If applicable:] We have no record of [CATEGORY] having been accessed.
We are continuing to verify the exact scope and will inform you once confirmed.
WHAT WE ARE DOING
- [MEASURE 1: containment applied] · [MEASURE 2: investigation with external support]
- [MEASURE 3: notification to the relevant authorities, if applicable]
WHAT WE RECOMMEND YOU DO
- [ACTION 1: review your team's access] · [ACTION 2: be alert to e-mails
requesting data]. Nimbus will never ask for your password by phone or e-mail.
NEXT COMMUNICATION
We will update you again before [DATE AND TIME], whether or not there is news.
For any query: [SINGLE CHANNEL] · [PHONE] · [CONTACT PERSON]
We sincerely regret this situation and thank you for your trust.
[NAME], [JOB TITLE] — Nimbus Reservas, S.L.Five drafting rules this template applies: do not promise what you do not know — "your data is safe" without having verified it is the sentence that later destroys your credibility; say explicitly what is not yet known, because communicated uncertainty generates more trust than false certainty; give the customer concrete actions, which reduces calls and gives them back a little control; commit to a date for the next communication and honour it even if there is no news; and a single contact channel, so that support does not collapse and different versions do not emerge.
8.2 Template: internal note
[INTERNAL - DO NOT FORWARD] Incident INC-2026-014 - Update 3 - 14:00
SITUATION: Contained. Investigation under way. Severity S2 (revised down from S1).
WHAT WE KNOW: unauthorised access to an employee's account between 04:12 and
09:40. They reached the ticketing tool. No evidence of access to production.
WHAT WE DO NOT KNOW: whether ticket attachments were downloaded. Confirmation at 18:00.
ACTIONS UNDER WAY: review of the ticketing tool logs (Ivan) · rotation of the
user's and their device's credentials (Lucia) · draft communication (Ruben).
WHAT THE TEAM MUST DO
- Do not discuss the incident outside the company, including on social media,
and redirect ANY customer or press enquiry to Ruben.
- Report immediately any strange e-mail or call: there may be attempts to take
advantage of the situation to impersonate us.
- Carry on working normally unless told otherwise.
NEXT UPDATE: 18:00, whether or not there is news.
- The incident log
A chronological record, written at the time and not reconstructed afterwards, of everything observed and decided. It matters for three reasons: it allows the later analysis to work from real data instead of memories; it underpins the legal defence and the insurance claim, which will ask when you knew and when you acted; and it avoids duplicated work when new people join the incident twelve hours in.
# Incident log INC-YYYY-NNN
Severity: __ · Manager: __ · Documenter: __ · Opened: YYYY-MM-DD HH:MM UTC
| Time (UTC) | Who | Type | Detail | Evidence |
|---|---|---|---|---|
| 08:05 | Rubén | FACT | Third clinic reports the diary will not load | Tickets #4412-14 |
| 08:41 | Marta | DECISION | S1 incident declared. Team activated | — |
| 08:50 | Lucía | ACTION | Ran `initial-collection.sh`; then isolated without shutting down | SHA256SUMS hash |
| 09:10 | Marta | ASSUMPTION | Entry may have been via a leaked credential. To confirm (Lucía, 12:00) | — |
| 09:30 | Marta | DECISION | Forensic retainer and legal advisers contacted | E-mail |
## Types: FACT · ASSUMPTION · DECISION · ACTION · COMMUNICATION
## Rules
1. It is written at the time, with the UTC time. It is never deleted: it is
corrected with a new entry that supersedes the previous one.
2. Every DECISION carries who took it and why; every ASSUMPTION, who is
responsible for confirming it and by when.
3. No opinions about people and no attributions of blame are written down.That last rule is not politeness: the log may end up being read by a lawyer, an expert witness or a customer, and a sentence like "this happened because Lucía never checks anything" turns a technical document into an additional problem.
- Runbooks by scenario
A runbook is the step-by-step procedure for a specific scenario, written to be executed by somebody under stress with no time to think. Minimum structure: scenario, activation indicators, initial severity, roles, numbered steps by phase, closure criteria and known errors.
10.1 Runbook written out in full: compromise of an employee's credentials
The most likely scenario at Nimbus, because it requires no technical vulnerability at all.
# runbook-RB-01.yaml
id: RB-01
scenario: "Compromise of an employee's credentials"
initial_severity: S2 # rises to S1 if the account is privileged
triggers: ["The employee reports having entered their password on a suspicious
site", "Access from an unusual country or IP", "Forwarding rules created
without the user's knowledge", "Mass sending from an internal account"]
roles: {decides: Marta, executes: Lucia, communicates: Ruben, documents: Sara}
steps:
detection_and_analysis:
- "1. Open the incident log: time, source of the detection and account affected."
- "2. Do NOT change the password yet: capture the state first."
- "3. Export the sign-ins of the last 30 days (IP, country, agent, result)
and keep them as evidence."
- "4. Review forwarding rules, delegations and OAuth applications: it is the
most common persistence and it survives a password change."
- "5. Determine the scope using the access inventory (POL-02 5.1.3).
If the account has privileged access -> S1."
containment:
- "6. Revoke ALL active sessions, not just the suspicious one."
- "7. Reset the password and force a new second factor."
- "8. Remove unrecognised forwarding rules, delegations and OAuth tokens."
- "9. Rotate that account's API keys and personal tokens."
- "10. If it had access to production: rotate the secrets it could have seen
and review the audit table (C-07)."
eradication:
- "11. Analyse the user's endpoint (session theft or malware) and identify
the vector, recording it for the post-mortem."
recovery:
- "12. Return access with new credentials and verified MFA."
- "13. Enhanced monitoring of the account for 14 days."
communication:
- "14. Internal: the note from section 8.2."
- "15. If customer data was accessed: assess notification WITH LEGAL ADVICE."
- "16. To the user: thank them for reporting it. Never reprimand (POL-04 5.3.1)."
closure_criteria: ["No anomalous activity for 14 days", "Persistence removed and
verified", "Vector identified and a corrective action opened in the risk register"]
known_errors:
- "Changing the password before exporting the logs: the trail is lost"
- "Forgetting the forwarding rules: the attacker carries on reading the e-mail"
- "Not revoking sessions: the stolen one stays alive despite the new password"
- "Blaming the user: guarantees the next person will not tell you"10.2 Outline of three other runbooks
| Runbook | Triggers | First three actions | Critical peculiarity |
|---|---|---|---|
| RB-02 Ransomware (S1) | Encrypted files, ransom note, mass encryption processes | Isolate from the network without shutting down · Verify the state of the immutable backups before touching anything · Activate the forensic retainer, legal advisers and insurer | Do not shut down the machines; the decision on whether to pay belongs to management with legal advice and is never technical (see the note in section 8) |
| RB-03 Data leak (S1) | Your own data published, external notice, extortion, a pattern of mass downloads in the audit log | Determine what data and from which customers · Preserve the access logs before they rotate · Start the 72-hour clock with legal advice | The main work is bounding the scope: without the scope there is no accurate notification (the deadlock of day 22 in 02-06) |
| RB-04 Outage of a critical supplier (S2/S1) | Unavailability of the cloud, the gateway or e-mail | Confirm that it is the supplier's and not your own · Activate the manual emergency procedures (04-06) · Communicate status to customers with an outlook | It is not an attack, but the business impact may be greater; it links into the continuity plan in 04-06 |
- The blameless post-mortem
The blameless post-mortem is the follow-up analysis meeting whose explicit purpose is to understand the system, not to assess the people. Its premise: if a human error was able to cause a serious incident, the problem is not the person, it is the system that allowed an individual error to have that consequence. How it is run, in four rules: it is convened between 3 and 7 days afterwards, no earlier — information is missing — and not much later — the detail is lost; the person most involved in the failure tells the story, without interruptions or judgements; sentences containing "should have" are banned and replaced with "what information was missing to take a different decision"; and nobody with disciplinary authority over the participants attends, or there will be no honesty.
11.1 The five whys on the 02-06 incident
(1) Why were the data and the backups encrypted? Because an attacker gained administrative control of the cloud account (A-05). (2) Why did they gain that control? Because they found an A-05 credential in a .env file on the application server. (3) Why did they reach that server? Because they came in with the consultancy's shared account, with no MFA and with permanent access. (4) Why did that permanent access with no MFA exist? Because it was granted in 2023 as an operational exception and nobody ever reviewed it again: there was no expiry, no periodic review and no contractual requirement. (5) Why was there no review and no requirement? Because no third-party risk management process existed: no prior assessment, no clauses, no access review.
Root cause: the absence of a third-party access management process with review and expiry. Notice two things. First: the root cause is not "the attacker" nor "Lucía left a .env lying about"; it is a process that did not exist, which is precisely what can be fixed. Second: the five whys produce one chain, but a real incident has several. There are at least two more here that deserve their own analysis: why it took 20 days to detect, and why the backups were destructible. The method is applied once per chain, not once per incident.
11.2 Post-mortem report template
# Post-mortem INC-YYYY-NNN — [Descriptive title, no people's names]
Date of the incident: __ · Date of the analysis: __ · Facilitator: __
Severity: __ · Duration: detection __ · containment __ · resolution __
## 1. Executive summary (5 lines, for management)
What happened, who it affected, what was done and what is going to change.
## 2. Impact
Customers affected · data involved · downtime · estimated cost ·
notification obligations triggered.
## 3. Timeline and metrics
Taken from the incident log, marking the first real sign, detection, declaration,
containment and resolution. Metrics: time to detection (from the first sign),
to containment (from detection) and to recovery, each against its target.
## 4. Root cause analysis
"Five whys" chains (one per causal line). Root cause(s) identified.
## 5. What went well · 6. What was missing
The first is mandatory: the controls and decisions that did help must be kept.
The second: absent controls, information that was not available and decisions
taken blind.
## 7. Corrective actions
| # | Action | Type (control/policy/process) | Owner | Date | Related risk | Status |
## 8. Resulting updates
Risk register · policies · control catalogue · third-party register ·
runbooks · continuity plan.The corrective actions section is the only one that matters six months from now, and that is why each action carries an owner and a date, and is tracked to closure in the monthly risk review from 04-01. A post-mortem whose actions are not closed is a writing exercise: the incident will come back under another name.
- Tabletop exercises
A tabletop exercise is a talked-through simulation: you get the team together for two hours, put a scenario to them and ask "so what do you do now?", without touching any system. It is the cheapest way to test the plan, and it always finds the same gaps: nobody knows who decides, the forensic specialist's number is not there, the runbook mentions a system that no longer exists, or everyone assumes somebody else was warning the customers.
Three formats with different cost and value: the plan walkthrough in a meeting (30 min, no cost, quarterly), the tabletop exercise proper (2 h, very low cost, every six months) and the technical drill of a runbook — actually restoring — which takes half a day, has a low-to-medium cost and is done once a year (04-06).
A minimum script for a two-hour one: the scenario is presented in three successive injects — the initial notice, a twist after 30 minutes ("a customer has posted about the incident on social media") and a complication at 60 ("the person who decides is on a plane"); every decision is recorded in a real incident log; and it ends with the list of gaps found, each with an owner and a date, just like a post-mortem. If the exercise does not generate actions, it was not done properly.
Common Mistakes and Tips
- Having no plan, or keeping it only where it may be encrypted. On day 20 at 09:00 nobody knew what to do first. Tip: a one-page plan with roles and phone numbers is worth more than a 60-page manual nobody has read, and it must exist on paper and on two people's phones, with an emergency channel separate from the corporate one.
- Shutting the machine down on instinct, or changing the password before exporting the logs. The first destroys the memory and may make the encryption worse; the second loses the trail and leaves the stolen session alive. Tip: isolate from the network while keeping it powered on, capture before containing and always revoke sessions.
- Communicating too much or too little. Promising the data is safe without knowing destroys your credibility; silence destroys it just as surely. Tip: communicate what you know, say what you do not know and commit to the next update.
- Not escalating for fear of overreacting, or investigating "just a bit more" before calling the forensic specialist. Tip: write down that over-escalating is never held against anyone, define the five stop criteria from section 5.3 in writing and stick to them.
- A post-mortem with culprits. It guarantees the next incident gets hidden. Tip: nobody with disciplinary authority in the room, and actions on the system, not on people.
Exercises
Exercise 1 — Classify and decide the first steps
Classify each situation as S1-S4, justify the level and state the first three actions in order:
- Sara reports having entered her corporate credentials on a page imitating the payroll portal. It happened 20 minutes ago.
- The monthly external scan finds that port 5432 is open to the internet again after last week's migration. There is no evidence of access.
- A clinic reports that, on opening one of its patients' records, it can see the name of a patient from another centre.
- Rubén forwards a generic phishing e-mail that was blocked by the filter and that no user opened.
Exercise 2 — Correct a badly executed response
At 03:40, Lucía spots an unknown process consuming CPU on the application server. She acts like this:
03:42 Kills the process with kill -9
03:44 Deletes the suspicious binary from /tmp
03:47 Restarts the server "just in case"
03:55 Checks that everything is fine and goes back to sleep
09:00 Mentions what happened in the daily team meetingIdentify every mistake, state what information has been irreversibly lost and rewrite the correct sequence with plausible times.
Exercise 3 — Post-mortem and corrective actions
Apply the five whys to this chain from the 02-06 incident that was not analysed in section 11.1: "On day 20, when a restore was attempted, there was no usable backup". Identify the root cause and draft three corrective actions in the format of section 8 of the template, stating for each one which risk from the 04-01 register it links to.
Solutions
Exercise 1
1. Phishing with credentials handed over: S3, rising to S2 depending on the reach of Sara's account. Sara handles administration and HR, so she has access to payroll and employee data (A-16) and probably to invoicing: that pushes it to S2. First actions: (a) open the incident log and export her account's sign-ins before touching anything; (b) revoke all active sessions, reset the password and force a new second factor; (c) review forwarding rules, delegations and OAuth applications on her mailbox. It is literally RB-01. And a fourth action, non-technical but mandatory: thank her for reporting it, because she raised the alarm within 20 minutes and that is what turns a disaster into a minor incident.
2. PostgreSQL re-exposed: S3. There is no evidence of access, but it is a critical service exposed to permanent automated scanning, so it cannot wait until the next day. Actions: (a) close the security group immediately — containment is trivial and destroys no evidence; (b) review the PostgreSQL connection logs over the exposure window to rule out access, and if there was any, escalate to S1; (c) open the root cause: why the migration reintroduced the rule, which is a case of control drift from 04-03 and must be corrected in the deployment process, not just in the security group.
3. One customer sees another's data: S1. It is a tenant isolation failure, that is, a confidentiality breach of data that reveals health information, and even though the origin is a programming error and not an attack, the impact is the same. Actions: (a) preserve evidence: capture the exact request, the user, the time and the associated audit record; (b) determine the scope — is it an isolated case or have they been seeing other people's data for months? — which is done by querying the C-07 audit table; (c) contain, disabling the affected functionality if the scope is not bounded within the first hour, and in parallel start the consultation with legal advisers because of the 72-hour clock. The typical mistake here is treating it as a bug on the work queue instead of as a security incident.
4. Blocked phishing: S4. There was no impact: it is logged, you check whether other users received it and you use it as input for training (06-05), without consuming the response team.
Exercise 2
Mistakes, and what each one destroyed:
| Action | Mistake | Information irreversibly lost |
|---|---|---|
Immediate kill -9 |
Contains before capturing | Process memory, active connections, parent and child processes, full arguments |
| Deleting the binary | Destroys the main evidence | Malware sample, hash, timestamps, the chance to identify the family |
| Restarting the server | Destructive containment with no need for it | All the memory, temporary files, sockets, and the persistence fires again: if the attacker left a start-up mechanism, the restart relaunches it |
| "Checks that everything is fine" | Confuses the absence of a symptom with the absence of compromise | — |
| Waiting until 09:00 | Does not declare the incident or tell anyone for 5 hours | A window in which the attacker could carry on or clean up their traces |
Correct sequence:
03:42 Opens the incident log: time, symptom and how she detected it.
03:45 Notifies Marta. A provisional S2 incident is declared.
03:50 WITHOUT killing the process: ps auxwwf, ss -tunapo, lsof of the PID,
/proc/<pid>/ (cmdline, environ, exe, maps) and a copy of the binary to
an external volume. Runs initial-collection.sh.
04:05 Hashes, closing of the evidence and chain of custody.
04:10 Marta authorises containment: it is ISOLATED from the network, powered
on. It is not shut down.
04:20 Reviews persistence (cron, systemd, authorized_keys) and recent access.
04:40 Rotates the credentials that server could expose.
05:00 Assesses whether to call the forensic retainer (criteria from 5.3).
08:30 Internal note. Rebuild from a clean image once the cause is corrected.Exercise 3
Five whys. (1) Why was there no usable backup? Because the backups were deleted by the attacker on day 20 at 02:10 and the only alternative was an external disk five weeks old. (2) Why were they able to delete them? Because they lived in the same cloud account as production and were reachable with the same administrative credentials. (3) Why were they there? Because they were configured for operational convenience and cost, without modelling the scenario of an attacker with administration credentials. (4) Why was that weakness not spotted earlier? Because there was no risk assessment and no control catalogue asking "what happens if the attacker has the administration credentials?". (5) Why did nobody know that the five-week-old copy was no good either? Because a restore had never been tested: neither its integrity nor how long it would take was known.
Root cause: the backup strategy was designed against technical failure — a disk that breaks — and not against an adversary, and its effectiveness was never verified.
| # | Action | Type | Owner | Date | Risk | Status |
|---|---|---|---|---|---|---|
| 1 | Immutable backups with locked retention in a separate account, with credentials production does not know | Control (C-14) | Lucía | +30 d | R-02 | Open |
| 2 | Quarterly restore test to an isolated environment, measuring the real elapsed time against the RTO and recording the result | Process | Lucía | +45 d | R-02, R-10 | Open |
| 3 | Alert on deletion of backups or snapshots sent to Marta as well as Lucía, so that a mass deletion does not depend on being seen by the person who might be compromised | Control (detective) | Lucía | +15 d | R-02, R-01 | Open |
Notice that the three actions are one preventive, one verification and one detective, applying the rule from 04-03 about covering several functions. And notice too that number 2 adds no new protection at all: it only checks that number 1 works. Without it, a year from now Nimbus would once again have backups it merely believes in.
Conclusion
You have written the plan before the incident, which is the only way it is any use, because at three in the morning nobody improvises well and the important decisions have to have been taken in advance. You distinguish event, alert, incident and crisis, knowing that the difference between an incident and a crisis is not one of technical size but of who has to be in the room, and that escalating late is one of the most expensive failures. You know the NIST SP 800-61 cycle with its two lessons: preparation is 80 % of the outcome and the analysis is reopened every time a new finding appears, because closing the scope too early leaves the attacker inside. You have Nimbus's response team with primaries and deputies and its three rules — whoever executes does not decide, one single voice outwards, and the documenter does nothing else — the out-of-band contacts on paper and on a separate channel because the plan cannot live only in the system that may be encrypted, the forensic retainer and the legal contact signed in the cold light of day, and the four-level severity matrix with the rule that when in doubt you escalate and that uncertainty about scope raises the severity. You know how to triage separating facts from assumptions in writing, so that a hypothesis does not become a certainty through repetition. And you have mastered what ruins the most incidents: preserving evidence before touching anything, with the order of volatility, a collection script that does not shut the machine down, writes to an external volume, calculates hashes and documents the chain of custody, and the five criteria for stopping and calling a forensic specialist. You know the real trade-offs of containment — isolate now versus observe, that cutting access tips off the attacker and can accelerate destruction as it did on day 20, and why you do not shut down — and in eradication, the rule of rebuilding rather than cleaning up when there were administrative privileges, with the five conditions for going back to production with enhanced monitoring.
You take away the templates: the notice to customers with its five drafting rules — do not promise what you do not know, say what you do not know, give concrete actions, commit to the next communication and a single channel — the internal note, the incident log with its writing rules, the complete runbook for a credential compromise with its known errors and the outline of three others, and the post-mortem report whose corrective actions section is the only one that matters six months later. And you take away the method of the blameless post-mortem applied to 02-06, with the root cause that is not "the attacker" nor a person but the absence of a third-party access management process with review and expiry, plus the reminder that an incident has several causal chains and each deserves its own analysis. The lesson closes with tabletop exercises, the cheapest way there is of discovering that the forensic specialist's number is missing and that everyone assumed somebody else was warning the customers. But go back over exercise 3 and you will see what is missing. The three corrective actions from the backup root cause — immutability in a separate account, quarterly restore test, deletion alert — are not incident response: they are another discipline. The response plan tells you how to act when the diaries of 40 clinics are at a standstill, but not how long they can be at a standstill before the business cannot take it, nor how much data you can afford to lose, nor how Rubén and the clinics carry on serving people while the system is down, nor how long a restore that nobody has ever timed actually takes.
In the last lesson of the module, Disaster Recovery and Business Continuity (04-06), you will answer that: the difference between BCP and DRP, the business impact analysis, RTO and RPO set from the BIA and not from wishful thinking, the backup strategies with the 3-2-1-1-0 rule developed digit by digit, the restore test as the heart of the lesson — because an unverified backup is not a backup — the ransomware scenario that breaks the classic plans, high availability that is not a backup, the start-up order by dependencies and the manual emergency procedures.
Fundamentals of Information Security Course
Module 1: Introduction to Information Security
- Basic Concepts of Information Security
- Types of Threats and Vulnerabilities
- Principles of Information Security
- Assets, Attack Surface and Threat Actors
Module 2: Cybersecurity
- Definition and Scope of Cybersecurity
- Types of Cyber Attacks
- Social Engineering and Phishing
- Protection Measures in Cybersecurity
- Identity, Authentication and Access Control
- Cybersecurity Incident Case Studies
Module 3: Cryptography
- Introduction to Cryptography
- Symmetric Cryptography
- Asymmetric Cryptography
- Hash Functions, HMAC and Password Storage
- Cryptographic Protocols
- Key Management, Certificates and PKI
- Applications of Cryptography
Module 4: Risk Management and Protection Measures
- Risk Assessment
- Security Policies
- Security Controls
- Third-Party and Supply Chain Risk
- Incident Response Plan
- Disaster Recovery and Business Continuity
Module 5: Security Tools and Techniques
- Vulnerability Analysis Tools
- Monitoring and Detection Techniques
- Penetration Testing
- Network Security
- Application Security
- System Hardening and Endpoint Security
- Cloud and Container Security
Module 6: Best Practices and Regulations
- Best Practices in Information Security
- Security Regulations and Standards
- Personal Data Protection and GDPR in Practice
- Compliance and Auditing
- Training and Awareness
- Ethics, Legal Aspects and Responsible Disclosure
