Resilience is more than reacting well to an incident
When something goes wrong, the immediate focus is usually response:
contain the problem, communicate, investigate and restore operations.
That is necessary, but it is not the full picture.
Operational resilience asks a broader question:
Can the organisation continue delivering what matters, recover effectively and learn from the disruption?
Current NIST guidance treats incident response as part of wider cybersecurity risk management rather than an isolated technical process. Its 2025 revision of SP 800-61 specifically integrates incident response across the Cybersecurity Framework 2.0 so organisations can improve preparation, detection, response and recovery.
Start by identifying what must keep working
Not every system, process or asset has the same operational importance.
Before an incident happens, organisations should understand:
- which services are critical;
- which systems support those services;
- which data is essential;
- who owns each critical process;
- which external suppliers are involved;
- how long disruption can reasonably be tolerated;
- what minimum level of service must be maintained.
This is where business continuity and incident response intersect.
ISO 22301 frames business continuity around preparing for, responding to and recovering from disruptions while maintaining the ability to deliver products and services at an acceptable predefined capacity.
The practical implication is simple:
Do not wait for an incident to decide what matters most.
Clear roles reduce chaos
Incidents create pressure.
Under pressure, ambiguity becomes expensive.
A practical incident structure should answer questions such as:
- Who declares the incident?
- Who coordinates the response?
- Who communicates internally?
- Who speaks to customers or external parties?
- Who preserves evidence?
- Who makes operational decisions?
- Who approves recovery actions?
- Who records what happened?
These responsibilities should exist before the incident.
NIST’s incident-response guidance emphasises preparation across organisational activities rather than treating response as a purely reactive cybersecurity task.
A technically strong team can still perform badly if nobody knows who is authorised to make decisions.
Detection without context creates noise
Modern organisations can generate large numbers of alerts.
The challenge is not merely detecting something unusual.
It is determining:
- what happened;
- what is affected;
- whether the event is credible;
- whether operations are at risk;
- whether personal or sensitive information is involved;
- whether escalation is required;
- what evidence must be preserved.
This requires context.
A security alert connected to a non-critical test system is not the same as an event affecting customer operations or core business data.
The objective should be useful detection, not simply more alerts.
Containment must protect the business as well as the system
Containment decisions often involve trade-offs.
Disconnecting a system may reduce technical risk but could also interrupt critical operations.
Keeping it online may preserve service but allow an incident to spread.
The correct action therefore depends on:
- severity;
- affected assets;
- business impact;
- evidence requirements;
- available alternatives;
- recovery capability.
This is why incident response must connect technical teams with operational decision-makers.
The response should minimise both:
incident impact and response-induced disruption.
Evidence and documentation matter
During an incident, teams often concentrate entirely on fixing the immediate problem.
But afterwards, the organisation may need to know:
- what happened;
- when it happened;
- which systems were involved;
- what decisions were made;
- who approved those decisions;
- what communications occurred;
- what data was affected;
- what remediation was completed.
A documented timeline becomes valuable for:
- internal review;
- forensic investigation;
- regulatory or contractual requirements;
- insurance;
- legal matters;
- customer communication;
- future prevention.
NIST’s recovery guidance specifically includes planning, playbooks, testing and continuous improvement after cybersecurity events.
Good incident management therefore creates an audit trail, not just a resolution.
Recovery should be planned before it is needed
“Restore from backup” is not a complete recovery strategy.
Recovery planning should consider:
- which systems are restored first;
- how backups are validated;
- how compromised credentials are handled;
- how restored environments are verified;
- who confirms normal operations;
- how customers are informed;
- what temporary controls remain in place.
NIST defines recovery as restoring capabilities or services impaired by a cybersecurity incident and improving existing strategies based on lessons learned.
The strongest recovery plans are tested.
An untested recovery plan is still partly an assumption.
Backups are necessary but not sufficient
Backups are central to resilience, but simply having backups does not guarantee recovery.
Organisations should know:
- whether backups are complete;
- whether they are protected from the same incident affecting production;
- how long restoration takes;
- whether restored data is usable;
- who has authority to initiate recovery.
A backup that has never been successfully restored in a controlled test should not automatically be treated as a proven recovery capability.
Communication is an operational control
Poor communication can make a manageable incident worse.
Employees need to know:
- what happened;
- what they should do;
- what they should not do;
- who they should contact;
- whether normal systems can be trusted.
Customers and external stakeholders may require different information.
Communication must therefore be:
- accurate;
- appropriately timed;
- authorised;
- consistent;
- documented.
NIST explicitly includes internal and external communications within response and recovery activities.
Silence creates uncertainty.
Speculation creates risk.
Controlled communication creates confidence.
Resilience requires learning
An incident should produce more than a closed ticket.
After recovery, ask:
- What allowed this to happen?
- Why was it not prevented?
- Why was it not detected earlier?
- Did escalation work?
- Were roles clear?
- Was evidence preserved?
- Did recovery work as expected?
- What created unnecessary delay?
- Which control needs changing?
NIST recovery guidance explicitly recommends using lessons from previous events to improve recovery planning and continuity of important functions.
This is one of the most important distinctions between organisations that merely survive incidents and those that become more resilient after them.
Operational resilience extends beyond cybersecurity
Cyber incidents are only one form of disruption.
The same resilience principles apply to:
- technology outages;
- supplier failures;
- fraud;
- physical-security incidents;
- infrastructure failures;
- communication breakdowns;
- key-person dependency;
- operational errors.
The architecture is broadly similar:
Prepare → Detect → Assess → Respond → Recover → Learn
This is why risk, security, business continuity and operations should not operate as completely separate disciplines.
They converge during disruption.
Build simple playbooks
Large policies have value, but incidents require actionable instructions.
A practical playbook can include:
Trigger
What causes the playbook to activate?
Owner
Who leads?
Immediate actions
What happens first?
Escalation
Who must be informed?
Evidence
What must be preserved?
Communication
What needs to be said, by whom?
Recovery
How is normal service restored?
Closure
Who confirms the incident is complete?
Review
What happens afterwards?
The objective is not to script every possible scenario.
It is to remove avoidable uncertainty.
The Juchepi perspective
At Juchepi Group, we view secure operations as the combination of:
people, process, technology, evidence and recovery.
Security is not simply preventing incidents.
It is building an operating environment that can:
- recognise risk;
- respond with control;
- protect sensitive information;
- preserve evidence;
- maintain critical services;
- recover effectively;
- improve after disruption.
The goal is not an organisation that never experiences problems.
That is unrealistic.
The goal is an organisation that is prepared when problems occur.
Strengthening operational resilience?
Juchepi Group helps organisations design practical incident workflows, secure communication processes, operational controls, continuity structures and recovery procedures around real operating risks.
Start a risk and secure-operations conversation with Juchepi Group.