When the Grid Goes Dark: A Communications Continuity Plan

Almost every failure your agency drills for is local and brief. A long, wide-area power or infrastructure failure is neither, and it arrives with a cruel timing problem: the moment your radios, phones, and dispatch consoles are most likely to fall silent is the exact moment the public needs you to be reachable. A communications continuity plan is the written answer to a simple question you should be able to answer before it happens rather than during it. When each layer of your system goes away, what is the next layer, and does everyone already know how to use it?

In this guide
  1. Why a grid-down event breaks the most at once
  2. The dependency chain almost nobody maps
  3. Designing a system that degrades gracefully
  4. Backup power realities at the station and the site
  5. Planning for days, not minutes
  6. Exercising the degraded modes
  7. Keeping the plan written and current
  8. A practical continuity checklist

Why a grid-down event breaks the most at once

Most communications interruptions are narrow. A repeater fails and you shift to a spare. A tower loses a feed line and coverage drops in one district. A carrier has a bad afternoon and cellphones get flaky for a few hours. These are annoying, but they are single points of failure with obvious workarounds, and they tend to fix themselves or get fixed quickly.

A long, wide-area power or infrastructure failure is different in kind, not just in degree. It does not take out one link. It removes the shared foundation that dozens of independent-looking systems quietly stand on. The dispatch center, the carrier towers, the internet path your computer-aided dispatch rides on, the microwave and fiber that backhaul your radio sites, and the office phones you would use to coordinate around all of it are drawing from the same well. When that well runs dry across a whole region, the failures do not happen one at a time in a way you can chase. They happen together, in the first hour, while call volume is climbing.

That is the part worth sitting with. The scenario that disables the most of your communications at once is also the scenario that generates the most need for communications. People trapped by weather, downed lines, and failed heat or cooling. Welfare checks. Medical calls that cannot reach a dispatcher. Mutual aid partners trying to find each other. The demand curve and the capability curve cross in the worst possible direction, and they do it precisely when you have the least slack to improvise.

The dependency chain almost nobody maps

Ask most agencies to name their communications systems and you will get a clean list: the radio system, cellphones, the dispatch center, station phones, maybe a paging system. Ask them to draw what each of those depends on to keep working, and the exercise usually stalls in about five minutes, because the dependencies are invisible until they are gone.

Here is the chain that a grid-down event exposes, roughly in the order it tends to bite:

The uncomfortable insight is at the bottom of that list. The tools you would use to run the response often depend on the very things the event is taking away. If your plan for a communications failure lives only in an online document or a notification app, your plan has the same single point of failure as the problem it is meant to solve.

Map it before you need it

Sit down with your radio administrator and draw every communications system as a box, then draw a line from each box to everything it needs to function: power, backhaul, a carrier, an internet path, a physical site. Where many lines converge on one thing, that thing is your real vulnerability. A grid-down event is just the case where the single converging line is power to an entire region.

Designing a system that degrades gracefully

You cannot make a wide-area failure impossible. You can make it survivable by designing your communications so that when a layer disappears, there is a clearly defined next layer and everyone already knows the drop is coming. The goal is graceful degradation: not a single system that never fails, but a stack of fallbacks that each still let you do the essential job of moving information between people.

A workable stack, from most capable to most primitive, usually looks like this:

The value of writing the stack down in this order is that it turns panic into procedure. Nobody has to invent the next move under pressure. When the primary system drops, the plan already says what the crew tries next, on what channel, at what interval, and what to do if that fails too.

Name the trigger, not just the fallback

A fallback layer is only useful if people know when to switch to it. For each drop in your stack, write the trigger in plain language: how a crew confirms the primary is actually down rather than momentarily quiet, how long they wait before moving to the next layer, and how they announce the change so nobody is left transmitting into dead air. Ambiguity here is where good plans still fail.

Backup power realities at the station and the site

Every layer above the runner depends on power somewhere, so backup power is where a continuity plan either holds or quietly falls apart. The honest work here is arithmetic and maintenance, not equipment shopping.

Think in terms of runtime rather than presence. A generator or a battery bank is not a yes-or-no capability. It is a number of hours at a given load, and that number is smaller than people assume. Batteries carry the gap for minutes to hours depending on how much they are asked to power. Generators carry longer, but only as long as they have fuel and only if they actually start. The questions that matter are conceptual and you can answer them without an engineering degree:

The same math applies at remote radio sites, and it is easy to forget because those sites are out of sight. A repeater site with a modest battery and no generator may give you a few hours after the grid drops and then go dark on its own, quietly, while your station generator hums along. That is why the dependency map from earlier matters so much. Your wide-area coverage can fail even when your building has power, because a link in the chain you do not own or visit ran out of runtime.

Planning for days, not minutes

The single most common flaw in communications continuity planning is scoping to the wrong duration. Plans get built around the brief outage everyone has actually experienced: power out for an hour, cellphones patchy for an afternoon, back to normal by dinner. A genuine grid-down scenario can run for days, and every assumption that was reasonable for one hour becomes dangerous over that span.

Extend the timeline and new problems appear that a short-outage plan never had to consider:

Planning for extended duration does not mean buying more of everything. It means asking the day-two and day-three questions during calm weather, and writing down answers that acknowledge scarcity honestly. A plan that assumes fuel arrives on request, that assumes people can work around the clock, or that assumes coverage that depends on a site nobody has checked in a year, is a plan written for the outage you have already survived, not the one that will actually test you.

Exercising the degraded modes

A continuity plan that has never been run is a document, not a capability. The layers most likely to save you in a grid-down event are the ones your crews use least in normal operations, which means they are also the ones your crews are least fluent in when it counts. Simplex operation, volunteer radio coordination, and runner routes all feel obvious on paper and turn out to be full of friction the first time they are used for real.

Exercise the fallbacks, not just the primary system, and exercise them the way they would actually be used:

Every exercise produces a list of small, specific things that did not work as written. That list is the entire point. Capture it, fix the plan, and run it again. A degraded mode that has been rehearsed a few times becomes muscle memory. One that lives only in a binder becomes a scramble at the worst possible moment.

Keeping the plan written and current

A communications continuity plan has two failure modes over time, and both are quiet. The first is that it lives only in a form that the event itself destroys. If your plan is a file that needs power and internet to open, you have built a lock whose key is inside the locked room. Keep a current printed copy at every station and in every dispatch position, and keep it somewhere a person can physically reach in the dark.

The second failure mode is drift. Systems change. Channels get renamed, a repeater site moves, a phone number for a mutual aid partner stops working, a volunteer coordinator retires, a generator gets replaced with one that has a different runtime. Every one of those changes silently invalidates a line in the plan, and nobody notices until the plan is being used in earnest and a step points to something that no longer exists.

The defense is unglamorous and it works: a named owner, a review on a fixed schedule, and a discipline that any change to a communications system or a partner relationship triggers a matching update to the plan. Tie the review to something you already do, so it does not depend on anyone remembering. A plan that is reviewed on a calendar and updated whenever the underlying reality changes stays trustworthy. A plan that was written once and admired is worse than no plan, because it gives false confidence in steps that quietly rotted.

A practical continuity checklist

Use this as a starting frame. Adapt every line to your own systems, partners, and geography, and keep the finished version short enough that a crew can actually work through it under stress.

None of this requires new equipment to begin. It requires sitting down before the grid goes dark and answering, on paper, the questions that a wide-area failure will otherwise ask you in the worst hour of the worst day. The agencies that come through a long outage with their communications intact are almost never the ones with the most gear. They are the ones who mapped the dependencies, wrote the fallbacks in order, did the runtime math honestly, drilled the degraded modes, and kept the whole thing current and printed where a person could reach it.

Keep the plan where a crew can actually find it

A continuity plan is only as good as your ability to reach it when everything else is down, and only as trustworthy as the last time it was updated. RunBoard keeps your continuity plans, SOPs, and standard operating guidelines organized, versioned, and easy to review on a schedule, so the copy your crews rely on is the current one, and so the annual review that keeps it honest actually happens.