When public safety agencies evaluate cloud-hosted software, resiliency is often discussed in broad terms.

Words like high availability, disaster recovery, multi-region, and redundancy sound reassuring. But from an operational standpoint, resiliency is not just about architecture diagrams or product positioning. It is about what actually happens when something fails, how quickly service is restored, what functionality returns first, and whether expectations between the vendor and the agency are truly aligned.

This week, I shared a number of posts on topics like RPO, RTO, multi-AZ vs. multi-region, and single-tenant vs. multi-tenant infrastructure. Taken together, they all point back to the same core issue: resiliency is not one feature. It is the combination of architecture, operations, contractual clarity, and shared understanding.

Resiliency is not one feature. It is the combination of architecture, operations, contractual clarity, and shared understanding.

Resiliency Starts with Clear Definitions

A lot of confusion begins when agencies and vendors use the same words but mean different things.

Two of the most important examples are RPO and RTO. Recovery Point Objective addresses data loss. Recovery Time Objective addresses restoration time. Those seem simple on the surface, but operationally they raise much bigger questions.

Without clarity on those questions, an agency may believe it is buying one level of resiliency while the vendor is committing to another.

Architecture Matters, But Topology Alone Does Not Tell the Full Story

Cloud resiliency is often described through topology. Is the application hosted in a single Availability Zone? Is it distributed across multiple Availability Zones? Is it capable of failing over across regions?

These are important distinctions, but they do not answer everything. A single-AZ deployment can still be hosted in the cloud, but it carries obvious availability risk if that zone experiences a disruption. A multi-AZ design typically improves resiliency within a region by reducing dependency on one data center location. A multi-region design can add another layer of resilience, especially for larger disaster scenarios, but it also introduces more complexity around replication, failover, cost, testing, and operations.

That is why the better question is not simply whether a system is in the cloud. The better question is how resiliency is actually designed and how that design affects operations, recovery expectations, and service commitments.

Operations Are Where Resiliency Becomes Real

The real test of resiliency is not the sales slide. It is the operational response.

When something fails, agencies should understand who is alerted first, what monitoring detects the issue, whether failover is automatic or manual, whether there is a documented runbook, how often recovery is tested, what dependencies must be restored in sequence, and what communication the customer will receive during the event.

A resilient environment is not just one with redundant infrastructure. It is one where the operational processes surrounding that infrastructure are mature, rehearsed, and understood.

SLA, RTO, and DR Are Related — But Not the Same

Availability SLAs, disaster recovery objectives, and the broader DR process are often grouped together, but they serve different purposes.

An SLA may define service availability over a month or year, often with exclusions and narrow measurement criteria. An RTO may describe the target time to recover after a disaster is declared. A disaster recovery process includes the actual operational steps, authority, escalation path, and sequencing required to restore service.

The gap between these can be significant. From an operational perspective, software may become unavailable long before a disaster is formally declared, and non-core functions may lag well behind core service restoration. That is why agencies should look beyond the headline number and ask more operationally specific questions.

Single-Tenant and Multi-Tenant Models Each Have Tradeoffs

This topic is often framed too simplistically, as if one model is always better than the other. In reality, both can support strong operations if they are designed and managed well.

A single-tenant model may offer more isolation and customer-specific control, but it can also introduce consistency challenges if each environment is deployed differently. A multi-tenant model may improve standardization and allow vendors to invest more efficiently in resiliency and automation, but it also changes how customers experience the platform, releases, and shared-service incidents.

The right question is not which model sounds better in theory. The better questions are how standardized the environment is, how failures are isolated, how updates are managed, how customer impact is measured and communicated, and how backup, recovery, and failover are handled.

Resiliency Is Also About Customer Understanding

One of the biggest risks in any cloud migration is misaligned expectations. A vendor may believe it has delivered a resilient system because the environment includes backups, redundancy, and documented recovery targets. An agency may believe it purchased a resilient system because it expects minimal downtime, rapid failover, and full restoration of operations in a crisis.

Those are not always the same thing.

Agencies do not need to become cloud architects, but they do need enough understanding to ask better questions, interpret the answers correctly, and connect technical design to operational outcomes. That includes understanding the difference between uptime and recovery, between redundancy and recoverability, between an architecture pattern and a contractual commitment, and between a cloud label and a cloud operating reality.

Read the original LinkedIn article

This blog is adapted from the original LinkedIn article in the resiliency series.

View the LinkedIn article →

Final Thought

Resiliency in public safety is not just about where software is hosted. It is about whether the architecture, operational processes, recovery expectations, and customer understanding all line up.

A system can be in the cloud and still leave important questions unanswered. A system can have redundant components and still create confusion during an incident. A contract can include targets and still fail to define what recovery truly means in practice.

The goal is not to be for or against any particular model. The goal is clarity. Because when it comes to public safety operations, resiliency should not be assumed. It should be understood.