When the Grid Goes Dark: A Communications Continuity Plan
Almost every failure your agency drills for is local and brief. A long, wide-area power or infrastructure failure is neither, and it arrives with a cruel timing problem: the moment your radios, phones, and dispatch consoles are most likely to fall silent is the exact moment the public needs you to be reachable. A communications continuity plan is the written answer to a simple question you should be able to answer before it happens rather than during it. When each layer of your system goes away, what is the next layer, and does everyone already know how to use it?
- Why a grid-down event breaks the most at once
- The dependency chain almost nobody maps
- Designing a system that degrades gracefully
- Backup power realities at the station and the site
- Planning for days, not minutes
- Exercising the degraded modes
- Keeping the plan written and current
- A practical continuity checklist
Why a grid-down event breaks the most at once
Most communications interruptions are narrow. A repeater fails and you shift to a spare. A tower loses a feed line and coverage drops in one district. A carrier has a bad afternoon and cellphones get flaky for a few hours. These are annoying, but they are single points of failure with obvious workarounds, and they tend to fix themselves or get fixed quickly.
A long, wide-area power or infrastructure failure is different in kind, not just in degree. It does not take out one link. It removes the shared foundation that dozens of independent-looking systems quietly stand on. The dispatch center, the carrier towers, the internet path your computer-aided dispatch rides on, the microwave and fiber that backhaul your radio sites, and the office phones you would use to coordinate around all of it are drawing from the same well. When that well runs dry across a whole region, the failures do not happen one at a time in a way you can chase. They happen together, in the first hour, while call volume is climbing.
That is the part worth sitting with. The scenario that disables the most of your communications at once is also the scenario that generates the most need for communications. People trapped by weather, downed lines, and failed heat or cooling. Welfare checks. Medical calls that cannot reach a dispatcher. Mutual aid partners trying to find each other. The demand curve and the capability curve cross in the worst possible direction, and they do it precisely when you have the least slack to improvise.
The dependency chain almost nobody maps
Ask most agencies to name their communications systems and you will get a clean list: the radio system, cellphones, the dispatch center, station phones, maybe a paging system. Ask them to draw what each of those depends on to keep working, and the exercise usually stalls in about five minutes, because the dependencies are invisible until they are gone.
Here is the chain that a grid-down event exposes, roughly in the order it tends to bite:
- Cellular service depends on tower sites that have their own batteries and generators. Those hold for a while, then they do not. Coverage does not vanish all at once. It thins as individual sites exhaust their backup power, and the ones nearest the hardest-hit areas often go first.
- Internet and data depend on powered equipment at every hop between you and wherever your data lives. A backup power source at your building does you no good if a node three miles away is dark.
- Dispatch operations depend on the center itself having power, on the phone and data trunks reaching it, and on the consoles, logging recorders, and mapping tools those dispatchers rely on to do the job.
- Repeater and tower sites depend on power for the repeater, the backhaul link that ties the site back to the rest of the system, and the site infrastructure like cooling. A site can have a fully charged battery and still be useless if the microwave hop feeding it lost power somewhere in the middle.
- The coordination tools you would reach for to manage all of the above, including email, messaging platforms, shared documents, and staff notification systems, mostly depend on the same internet and cellular layers that are already failing.
The uncomfortable insight is at the bottom of that list. The tools you would use to run the response often depend on the very things the event is taking away. If your plan for a communications failure lives only in an online document or a notification app, your plan has the same single point of failure as the problem it is meant to solve.
Sit down with your radio administrator and draw every communications system as a box, then draw a line from each box to everything it needs to function: power, backhaul, a carrier, an internet path, a physical site. Where many lines converge on one thing, that thing is your real vulnerability. A grid-down event is just the case where the single converging line is power to an entire region.
Designing a system that degrades gracefully
You cannot make a wide-area failure impossible. You can make it survivable by designing your communications so that when a layer disappears, there is a clearly defined next layer and everyone already knows the drop is coming. The goal is graceful degradation: not a single system that never fails, but a stack of fallbacks that each still let you do the essential job of moving information between people.
A workable stack, from most capable to most primitive, usually looks like this:
- Primary system. Your normal radio system with full infrastructure. Wide-area coverage, dispatch integration, the way you work every day.
- Backup channels on the same system. Alternate talkgroups or channels, a spare repeater, or an alternate site that keeps you on infrastructure even when part of it is degraded.
- Simplex, or direct radio-to-radio. When the infrastructure is gone, radios can still talk directly to each other without any repeater in between. Range shrinks, and you lose the wide-area reach, but two units within line of sight of each other can still coordinate. This is the layer most crews underuse because they rarely practice it. It should be a named, understood fallback, not a rumor.
- Amateur and volunteer backup. Trained volunteer radio operators can move messages across distances your simplex cannot reach, using their own equipment and their own power. Set this relationship up in advance, in writing, with agreed meeting points and check-in schedules, because you cannot recruit and coordinate volunteers in the middle of the failure that would need them.
- Runners and hard-wired methods. The oldest layer still works. A person carrying a written message between two fixed points. A hardwired phone line that does not depend on the local internet. Pre-agreed physical rally points where a representative from each station or agency shows up at set times. It is slow, and it is unglamorous, and it has never once failed for lack of power.
The value of writing the stack down in this order is that it turns panic into procedure. Nobody has to invent the next move under pressure. When the primary system drops, the plan already says what the crew tries next, on what channel, at what interval, and what to do if that fails too.
A fallback layer is only useful if people know when to switch to it. For each drop in your stack, write the trigger in plain language: how a crew confirms the primary is actually down rather than momentarily quiet, how long they wait before moving to the next layer, and how they announce the change so nobody is left transmitting into dead air. Ambiguity here is where good plans still fail.
Backup power realities at the station and the site
Every layer above the runner depends on power somewhere, so backup power is where a continuity plan either holds or quietly falls apart. The honest work here is arithmetic and maintenance, not equipment shopping.
Think in terms of runtime rather than presence. A generator or a battery bank is not a yes-or-no capability. It is a number of hours at a given load, and that number is smaller than people assume. Batteries carry the gap for minutes to hours depending on how much they are asked to power. Generators carry longer, but only as long as they have fuel and only if they actually start. The questions that matter are conceptual and you can answer them without an engineering degree:
- What has to stay powered, and what does not? Keeping the essential radio and dispatch functions alive is a much smaller load than keeping the whole building comfortable. Know which is which before the event, and know how to shed the rest.
- How long does each backup source last at that essential load? Battery runtime shrinks fast as load grows. Generator runtime is a function of tank size and burn rate. Do the rough math ahead of time and write the answer down in hours.
- Where does the fuel come from after the tank runs low? This is the question that separates a short-outage plan from an extended-duration plan. Fuel deliveries themselves depend on the same failing infrastructure, and demand for fuel spikes across the whole region at once. Onsite reserve, a resupply agreement, and a realistic delivery expectation all belong in writing.
- Does the backup power start and carry the load when tested? A generator that has not been exercised under load is a hope, not a plan. Cold starts fail, transfer switches stick, and fuel goes stale. The only way to trust it is to run it.
The same math applies at remote radio sites, and it is easy to forget because those sites are out of sight. A repeater site with a modest battery and no generator may give you a few hours after the grid drops and then go dark on its own, quietly, while your station generator hums along. That is why the dependency map from earlier matters so much. Your wide-area coverage can fail even when your building has power, because a link in the chain you do not own or visit ran out of runtime.
Planning for days, not minutes
The single most common flaw in communications continuity planning is scoping to the wrong duration. Plans get built around the brief outage everyone has actually experienced: power out for an hour, cellphones patchy for an afternoon, back to normal by dinner. A genuine grid-down scenario can run for days, and every assumption that was reasonable for one hour becomes dangerous over that span.
Extend the timeline and new problems appear that a short-outage plan never had to consider:
- Batteries that carried the first hours are exhausted, and the question becomes whether they recharge and how.
- Generators need refueling on a schedule, which means someone has to track burn rate, watch the tank, and trigger resupply before it is urgent rather than after.
- People need relief. A dispatcher, a radio operator, or a runner cannot work indefinitely. A multi-day communications posture is a staffing plan as much as a technical one, with shifts, rest, and food built in.
- Volunteer and mutual aid partners have their own limits. The amateur radio operator helping you also has a family, a home losing power, and finite endurance. Plan for rotations there too.
- Information decays. Over days, the picture changes constantly. Without a discipline for logging and handing off what is known, each shift starts blind.
Planning for extended duration does not mean buying more of everything. It means asking the day-two and day-three questions during calm weather, and writing down answers that acknowledge scarcity honestly. A plan that assumes fuel arrives on request, that assumes people can work around the clock, or that assumes coverage that depends on a site nobody has checked in a year, is a plan written for the outage you have already survived, not the one that will actually test you.
Exercising the degraded modes
A continuity plan that has never been run is a document, not a capability. The layers most likely to save you in a grid-down event are the ones your crews use least in normal operations, which means they are also the ones your crews are least fluent in when it counts. Simplex operation, volunteer radio coordination, and runner routes all feel obvious on paper and turn out to be full of friction the first time they are used for real.
Exercise the fallbacks, not just the primary system, and exercise them the way they would actually be used:
- Run a shift on simplex. Have crews operate direct radio-to-radio for a defined window and discover for themselves where the range dies, who cannot hear whom, and how the traffic discipline has to change without a repeater tying everyone together.
- Activate the volunteer backup as a drill. Confirm the phone tree, the meeting points, and the check-in schedule work when they are exercised rather than assumed. A relationship that only exists on paper tends to stay on paper.
- Walk the runner routes. Time them. Confirm the rally points are reachable, that the message forms are usable, and that the people who would carry them know where to go.
- Test the backup power under load, and test the switch from one layer of the stack to the next, so the trigger points you wrote down get pressure-tested by people who were not the ones who wrote them.
Every exercise produces a list of small, specific things that did not work as written. That list is the entire point. Capture it, fix the plan, and run it again. A degraded mode that has been rehearsed a few times becomes muscle memory. One that lives only in a binder becomes a scramble at the worst possible moment.
Keeping the plan written and current
A communications continuity plan has two failure modes over time, and both are quiet. The first is that it lives only in a form that the event itself destroys. If your plan is a file that needs power and internet to open, you have built a lock whose key is inside the locked room. Keep a current printed copy at every station and in every dispatch position, and keep it somewhere a person can physically reach in the dark.
The second failure mode is drift. Systems change. Channels get renamed, a repeater site moves, a phone number for a mutual aid partner stops working, a volunteer coordinator retires, a generator gets replaced with one that has a different runtime. Every one of those changes silently invalidates a line in the plan, and nobody notices until the plan is being used in earnest and a step points to something that no longer exists.
The defense is unglamorous and it works: a named owner, a review on a fixed schedule, and a discipline that any change to a communications system or a partner relationship triggers a matching update to the plan. Tie the review to something you already do, so it does not depend on anyone remembering. A plan that is reviewed on a calendar and updated whenever the underlying reality changes stays trustworthy. A plan that was written once and admired is worse than no plan, because it gives false confidence in steps that quietly rotted.
A practical continuity checklist
Use this as a starting frame. Adapt every line to your own systems, partners, and geography, and keep the finished version short enough that a crew can actually work through it under stress.
- Dependency map. Every communications system drawn to the power, backhaul, carrier, and internet paths it depends on, with the shared choke points marked.
- The fallback stack, in order. Primary system, backup channels, simplex direct, volunteer and amateur backup, runners and hardwired methods, each with the trigger that moves you to it and the channel or route it uses.
- Backup power by location. For every station and remote site, what is powered, the runtime in hours at essential load, the fuel reserve, and the resupply plan.
- Extended-duration provisions. Refueling schedule and responsibility, staffing rotations for dispatch and radio operators, volunteer relief, and a logging and handoff discipline for multi-day events.
- Partner contacts. Mutual aid and volunteer radio contacts with meeting points and check-in schedules, kept current and available on paper.
- Exercise record. When each degraded mode was last drilled, what broke, and whether the fix landed back in the plan.
- Physical availability. A current printed copy at every station and dispatch position, plus the named owner and the next scheduled review date.
None of this requires new equipment to begin. It requires sitting down before the grid goes dark and answering, on paper, the questions that a wide-area failure will otherwise ask you in the worst hour of the worst day. The agencies that come through a long outage with their communications intact are almost never the ones with the most gear. They are the ones who mapped the dependencies, wrote the fallbacks in order, did the runtime math honestly, drilled the degraded modes, and kept the whole thing current and printed where a person could reach it.
A continuity plan is only as good as your ability to reach it when everything else is down, and only as trustworthy as the last time it was updated. RunBoard keeps your continuity plans, SOPs, and standard operating guidelines organized, versioned, and easy to review on a schedule, so the copy your crews rely on is the current one, and so the annual review that keeps it honest actually happens.