Resilience & Readiness
When an Incident Hits, Execution Should Start — Not Planning
The first minutes after disruption are the wrong time to begin sourcing capacity, negotiating terms, provisioning access, or deciding who is in charge. Good readiness moves that work upstream.

The first question is not "How fast can we respond?"
When something breaks, everybody suddenly becomes a student of preparedness. The conference bridge fills up. Someone asks who has the latest runbook. Someone else asks whether the backup team actually has access. Procurement gets pulled into a conversation it has never seen before. A manager discovers that the 'alternate provider' is really a company we once met at a conference.
And then we start measuring response time.
That is useful, but I think it begins one question too late. The more important question is: how much of the response had already been completed before anything broke?
A good incident response should feel less like improvisational theater and more like executing a play we have already rehearsed. Not because incidents are predictable. They are not. But many of the constraints that make response slow are completely predictable: capacity, contracts, access, training, routing, decision rights, communications and the simple problem of who is allowed to do what.
If those questions are still open when the incident begins, the organization is not responding yet. It is planning under pressure.
Move the work to the left
Software teams often talk about 'shifting left': moving testing, security or quality work earlier in the lifecycle, when problems are cheaper and easier to solve. Customer operations can borrow the same idea.
Move the work left of the incident.
Qualify the capacity before you need it. Understand the commercial terms before urgency destroys your negotiating leverage. Provision access before the primary environment is impaired. Train people before the queue is overflowing. Define routing rules before managers are making them up in Slack. Decide who can activate the plan before five executives are debating authority on a conference call.
The incident itself should consume readiness, not create it.
Ready.gov makes a similar point in the broader emergency-planning context: the actions taken in the first minutes matter, and establishing the response plan in advance saves valuable time. NIST contingency-planning guidance likewise treats recovery strategies, plans, testing, training, exercises and maintenance as work that belongs in the preparedness cycle rather than as activities invented after disruption.
That is not a customer-service-specific rule. It is a systems rule. The more decisions we can safely make before uncertainty arrives, the more scarce human attention remains available for the decisions that truly require judgment in the moment.

What should already be true before the phone rings
Readiness is not a binder. It is a collection of operating conditions that already exist.
Capacity. There is a real source of additional or alternate capacity, not merely a list of vendors that might be interested.
Commercial terms. The parties understand how activation works, what the economics are, and what commitments can be made quickly.
Access. Required systems, data, credentials, privacy controls and permissions are provisioned or can be enabled through a tested process.
Training. People know the work, the policies, the tools and the boundaries of their authority.
Routing. The organization knows what work can move, where it can move, and what should happen when the preferred route is unavailable.
Governance. Decision rights, escalation paths, communications responsibilities and operating command are explicit.
Exercises. The plan has been tested enough to expose the things that look obvious on paper and fail in practice.
FEMA's continuity guidance emphasizes that exercises validate plans and capabilities, clarify roles and responsibilities, and surface resource gaps. That last part matters. The point of an exercise is not to prove that the plan is brilliant. It is to discover where reality disagrees with the document while there is still time to fix it.
Prepared does not mean rigid
There is an obvious objection here: real incidents never follow the script. Correct.
Preparation is not an attempt to predict every event. It is an attempt to remove avoidable uncertainty so the team can spend its energy on the uncertainty that remains.
Think about a soccer team. You do not know exactly where the ball will be in the 73rd minute. You do know the formation, the roles, the substitutions, the signals and what kinds of movements your teammates are likely to make. Structure creates freedom. It gives players a platform from which to improvise intelligently.
Incident readiness works the same way. If the alternate team is trained, the data permissions are known and the activation authority is clear, managers can adapt the response to the actual situation instead of burning the first hour trying to assemble the machinery of response.
Good preparation does not eliminate judgment. It makes judgment more valuable.
The operating model after the incident should use different verbs
There is a simple test I like for readiness: look at the verbs in your incident plan.
If the first page says source, evaluate, negotiate, contract, provision, train and design, you are looking at a preparation plan that happens to begin after the incident.
After the event, the verbs should look more like activate, reroute, communicate, monitor, adjust and learn.
That distinction sounds semantic. Operationally, it is enormous.
It is also why a spreadsheet full of backup suppliers does not necessarily create resilience. Availability is a possibility. Readiness is an operating state. A company can have plenty of theoretical capacity in the market and still have nowhere useful to send a customer interaction at 2:00 a.m. on the day something fails.
The point is not that every company needs multiple BPOs, multiple sites or some elaborate command structure. One provider may have genuinely independent capacity. An internal team may be the best failover path. Automation may absorb part of the workload. Different programs need different architectures.
The design principle is simpler: there should be another usable path, and the work required to make that path usable should be completed before the primary path is impaired.
Readiness is a product you build before you need the outcome
Most organizations are very good at buying outcomes they can see. Staffing. Software. Service levels. Projects. We are less naturally inclined to invest in something whose value is mostly invisible when nothing is going wrong.
That is the strange economics of readiness. On a quiet Tuesday, it can look like unused optionality. During a consequential event, that optionality becomes the difference between executing a response and beginning a search for one.
This is one reason I have come to think about readiness as something we build and maintain, not merely a document we possess. Capacity changes. People leave. Permissions expire. Systems change. Commercial assumptions drift. A plan that was executable six months ago can quietly become theoretical.
So readiness needs a lifecycle: curate the options, prepare the capacity, test the paths, activate when required, monitor performance, learn from the event, and repair whatever the event consumed.
The best time to discover that your backup path does not work is yesterday. The second-best time is during an exercise.
The worst time is when the customer is already waiting.
Sources & References
- Ready.gov (FEMA), Last updated September 7, 2023. Emergency PlansReference 1
- NIST Special Publication 800-34 Rev. 1, May 2010, updated November 2010. Contingency Planning Guide for Federal Information SystemsReference 2
- FEMA. Continuity Guidance Circular - 2024 Update — Continuity exercise and improvement guidance.Reference 3