Skip to main content

Seasonal & Surge Operations

You’re Gonna Need a Bigger Boat: Improve Scalability in Customer Service

When demand suddenly outruns staffing, the answer is not always more people. Better customer-operations systems create multiple capacity levers — and the ability to route work where the network can absorb it.

Alan Pendleton | Originally published on Medium March 4, 2021Article · Open Article · 5 min read
Customer-operations leaders and frontline staff coordinating through a live demand spike in a contemporary operations environment.

When wide-eyed Chief Brody tells Captain Quint, “You’re gonna need a bigger boat,” he perfectly channels customer-service leaders staring at a sudden onslaught of demand. The queue is climbing, the forecast is wrong, and the operation that felt perfectly adequate yesterday suddenly looks very small.

Our reflex is usually the same as Quint’s: attack the problem with brute force. Add overtime. Pull people from somewhere else. Launch an all-hands swarm. Hire faster. Ask the outsourcer for more seats. In other words, get a bigger boat.

Sometimes that is exactly what the situation requires. But it is a very expensive way to discover that we have been defining capacity too narrowly.

A bigger team is only one answer

Customer-service organizations tend to think about capacity in terms of agent population. If I have 100 people, I can do roughly what 100 people can do. When demand suddenly looks like work for 200 people, the answer appears obvious: find another 100 people.

That headcount view is understandable, but it turns scalability into a staffing emergency instead of an operating-design problem. It also treats every unit of demand and every unit of supply as though they were identical. They are not.

Different contacts have different urgency, complexity, skills, channel requirements, and automation potential. Different teams have different schedules, proficiencies, locations, technology access, and available slack. External partners have different business cycles, labor markets, capabilities, and ramp profiles. There is usually more capacity in the system than a simple headcount number reveals — and more than one way to create additional capacity when it is genuinely needed.

Capacity is a system, not a headcount number

Like the marine biologist on board the Orca, Matt Hooper, we can approach the problem with a beard, glasses, and a dose of science. The beard is optional.

Computer systems do not respond to every surge by installing twice as much hardware. They use a portfolio of techniques: distribute work, run more things concurrently, route around constrained paths, preserve redundancy, prioritize what matters, and add resources when the cheaper levers are exhausted.

Customer operations can be designed the same way. Scalability is the ability to change effective capacity as demand changes — by reducing work, unlocking capacity that already exists, moving work to a different ready resource, or adding supply. Headcount is one lever inside that system, not the system itself.

The practical question is no longer simply, “How many agents do we have?” It is, “What can this operating network do when demand changes quickly?”

Diagram showing available customer-service capacity supported by demand shaping, scheduling and intraday management, cross-training, workflow and automation, external flex capacity, and routing and orchestration.
Capacity can be created or unlocked through multiple operating levers; headcount is only one of them.

Build more than one capacity lever

Reduce or reshape the work. Self-service, automation, better knowledge, simpler processes, asynchronous channels, backlog triage, and deliberate prioritization can all change how much human capacity a demand spike actually consumes. This does not mean pushing customers away. It means being intentional about which work needs a person, which work needs to happen now, and which work can be made easier.

Unlock capacity already on the payroll. Schedules, intraday management, adherence, shrinkage planning, cross-training, channel balancing, and shared support pools determine how much theoretical staffing becomes usable capacity. A person who is technically employed but cannot handle the queue that is overflowing is not useful capacity for that moment.

Move work before adding people. Routing is a capacity lever. Work can sometimes move across teams, queues, geographies, sites, business units, or service partners when another part of the network has room. This requires common enough processes, knowledge, access, and quality controls that moving the work does not create a second problem somewhere else.

Add supply when you truly need it. Overtime, seasonal hiring, flexible labor, temporary resources, external partners, specialist providers, and additional sites all have a place. The point is not to eliminate brute-force capacity. It is to avoid making brute force the only move available.

Diversification creates optionality

One of the most powerful capacity levers is diversification. If all customer work depends on one team, one location, or one outsourcing partner, then the capacity of that one operating path becomes the capacity of the whole system.

A portfolio changes the math. Different teams and providers have different strengths, busy periods, staffing pools, geographies, languages, and constraints. A demand spike that overwhelms one resource may arrive at a moment when another has room. When work can move intelligently among them, the network can absorb variability that would break a single-threaded model.

The investment analogy from the original version of this essay still works: a diversified portfolio gives you more options than a single stock. But there is an important operational caveat. More providers are not automatically better. A portfolio that cannot share work, data, knowledge, standards, or decisions is just a larger collection of silos.

Optionality becomes valuable when the organization can orchestrate it. That means knowing what each resource can do, maintaining the prerequisites to use it, and having routing and governance mechanisms that can shift work without turning the customer experience into chaos.

Comparison diagram showing a single operating path creating a bottleneck versus an elastic network routing demand across a core team, flex team, BPO or specialist provider, and automation.
Diversification creates useful capacity only when demand can be routed across resources that are ready to absorb it.

Elasticity has to exist before the peak

The worst time to invent a capacity strategy is after the queue has already exploded.

A backup provider that still needs six weeks of training is not surge capacity. A second site without credentials is not redundancy. Cross-trained people who cannot access the right systems are not a flex pool. An automation workflow that has never been tested under real operating conditions is not a dependable relief valve.

Elasticity is built in advance. It comes from defined activation thresholds, current knowledge, ready access, trained resources, commercial mechanisms, tested routing, clear decision rights, and a management cadence that can see pressure building before customers feel the full effect.

This is why forecasting still matters even in an elastic model. Forecasting tells you what you expect. Elasticity gives you a way to respond when reality refuses to cooperate.

Bigger boat or better system?

Customer operations will always face moments when the answer really is more capacity. Sometimes you need more people, more hours, more technology, or more external support. Sometimes you really do need a bigger boat.

But a scalable operation should not need to build that boat from scratch every time demand changes.

The better design is an operating network with multiple capacity levers: reduce work where it makes sense, unlock the capacity already available, move demand across prepared resources, and add supply when the economics justify it. Diversification provides options. Routing makes those options usable. Governance keeps the response coordinated.

That is the difference between reacting to a surge and being designed for one.