Public safety agencies rely on software that cannot be evaluated like an ordinary office application. When a financial or administrative system is unavailable, work may be delayed. When CAD, RMS, mobile, corrections, 911, mapping, or a critical interface is unavailable, personnel may have to continue operating without information they normally depend on to protect the public.
That is why software resiliency should be evaluated as an operational capability, not simply as a technical feature. The central question is not whether the system is described as redundant, highly available, cloud-based, or disaster-recovery capable. The central question is whether the agency can continue operating when part of the service fails—and whether the claimed recovery capability has been demonstrated.
Resiliency Applies On Premises and in the Cloud
Resiliency conversations often become debates about hosting location. Cloud platforms may provide multiple availability zones, automated infrastructure, managed services, and geographic scale. On-premises systems may provide direct control over hardware, networks, personnel, and local operating procedures.
Neither hosting model is automatically resilient.
An on-premises environment can be highly resilient when it has redundant compute, storage, networking, power, geographic recovery, trained staff, strong monitoring, and regularly tested procedures. It can also appear redundant while still depending on one storage array, one building, one network path, one electrical feed, or one person who knows how to restore it.
A cloud-hosted environment can be highly resilient when the application, database, identity, connectivity, interfaces, backups, monitoring, and support model are designed around independent failure domains and tested recovery processes. It can also contain redundant cloud infrastructure while the agency still depends on one internet circuit, one firewall, one VPN, one identity provider, or an untested interface path.
The location of the software matters. The complete service path matters more.
High Availability and Disaster Recovery Solve Different Problems
High availability is intended to keep a service operating through expected component failures. If an application server fails, another server may continue processing. If a cloud availability zone is disrupted, workloads may continue elsewhere. If a network path fails, traffic may move to another path.
Disaster recovery addresses larger failures that overwhelm or invalidate the primary environment. Examples include loss of a data center, regional disruption, ransomware, database corruption, widespread configuration errors, destructive administrative mistakes, or an event that affects both primary and redundant components.
A system can have strong high availability and weak disaster recovery. It can also have a documented disaster recovery plan but little meaningful high availability. Agencies should understand both capabilities, the events each one covers, and what operations will look like while recovery is underway.
Availability Does Not Automatically Equal Operational Availability
Availability percentages are useful, but they do not explain how an outage will affect dispatchers, field users, records staff, corrections personnel, supervisors, or agency partners.
A contract may state 99.9% availability without clearly explaining:
- Which application components are covered
- Whether degraded performance counts
- Whether mobile users and remote facilities are included
- Whether critical interfaces are included
- Whether planned and emergency maintenance are excluded
- Whether agency connectivity, VPN, firewall, identity, or third-party failures are excluded
- How partial outages are measured
- When an incident begins and ends for SLA purposes
- Whether service credits are the agency's only contractual remedy
A CAD application may technically be running while an interface is unavailable. An RMS may be healthy while field users cannot connect. A vendor's platform may be online while an agency's network path is down. A system may remain accessible but perform too slowly to support normal operations.
Application availability asks whether the software is running. Operational availability asks whether public safety personnel can actually perform their work.
Resiliency Must Cover the Complete Service Path
Public safety software usually depends on more than the visible application. The operational service may include application servers, databases, storage, identity, internet service providers, agency networks, firewalls, VPNs, domain services, dispatch workstations, mobile computers, mapping, telephony, radio interfaces, state and federal systems, third-party integrations, and vendor and agency support teams.
Each dependency can become a failure point. That is why an architecture review should not stop at a vendor's hosting environment. The agency side, the vendor side, and the connections between them must be evaluated together.
A highly available application connected through a single agency circuit is not an end-to-end resilient service. A geographically separate disaster recovery environment that omits essential interfaces may restore software without restoring operations. A backup that has never been restored may provide reassurance without evidence.
RTO and RPO Must Be Translated Into Agency Impact
Recovery time objective, or RTO, describes the targeted period for restoring service after a qualifying disruption. Recovery point objective, or RPO, describes how much recent data may be unavailable or lost when service is restored.
Those terms should be translated into operational questions:
- What will dispatch, field, records, or corrections personnel do if the system remains unavailable for the full RTO?
- How will the agency capture work performed during the outage?
- How will missing or duplicate information be reconciled after restoration?
- Does the RTO apply to the complete production service or only the core application?
- Does the RPO cover attachments, interfaces, configuration, audit records, mobile data, reporting, and agency-specific content?
- Who decides the system is sufficiently restored for operational use?
A recovery objective is meaningful only when the agency understands what will be restored, what may be missing, and how personnel will operate before normal service returns.
Backups Are Essential, but They Are Not Disaster Recovery
A backup may contain data without providing the application servers, security configuration, network access, authentication, interfaces, documentation, staffing, or replacement infrastructure required to operate the system.
The existence of a completed backup does not prove that it is complete, uncorrupted, compatible with the current software version, protected from the same event affecting production, or restorable within the required timeframe.
The important evidence is not only that backups are created. It is that the organization can restore the service from them and validate that the restored environment is usable.
Resiliency Must Be Tested
A diagram is not a test. A vendor presentation is not a test. A disaster recovery document is not a test. A successful backup notification is not a test.
Testing may include:
- Failing an application or database component
- Losing a network path or internet circuit
- Failing over between sites, regions, or availability zones
- Restoring the system from backup
- Activating a disaster recovery environment
- Validating identity and user access
- Testing mobile, remote-site, and field connectivity
- Confirming that critical interfaces operate after failover
- Measuring actual recovery time and recovered data
- Operating temporarily in a degraded or fallback mode
- Returning safely to the primary environment
- Testing incident communication and escalation
Not every test needs to create a production outage. Controlled environments, tabletop exercises, component tests, partial failovers, and planned maintenance windows can all produce useful evidence. The important point is that the test should match the claim being made.
Agencies should understand what was tested, when it was tested, which systems and interfaces were included, who participated, what failed, how long recovery took, whether data was lost, what improvements were identified, and whether corrective actions were completed.
Planned Maintenance Is Also an Availability Event
Unexpected failures are not the only source of disruption. Software upgrades, operating system patches, database maintenance, infrastructure changes, security updates, and emergency changes can also affect operations.
Agencies should understand how often maintenance occurs, how much notice is provided, whether downtime is required, whether redundant components are maintained separately, whether interfaces remain available, how rollback works, and whether maintenance is excluded from availability calculations.
Planned does not mean operationally insignificant. A scheduled interruption still requires preparation when the service supports emergency communications and response.
Shared Responsibility Must Be Explicit
Resiliency is rarely owned by only the software vendor or only the agency. The vendor may operate the application, hosting environment, backups, monitoring, and recovery systems. The agency may own local networks, internet circuits, endpoints, identity, firewall configuration, fallback procedures, and user communications. Other providers may support telephony, interfaces, mapping, connectivity, or regional services.
Before an incident occurs, the parties should know who detects the problem, who declares a disaster, who initiates failover, who contacts connectivity providers, who validates interfaces, who communicates with users, who approves the return to service, and who reconciles data afterward.
What an Independent Resilience Architect for Public Safety Software Actually Does
The phrase resilience architect for public safety software can sound as though the consultant develops or operates the software. That is not necessarily the role.
Independent resiliency assessment and assurance
Public Safety Cloud Standards does not build a vendor's CAD, RMS, mobile, corrections, or 911 software and does not control the vendor's production environment.
PSCS independently analyzes the architecture, vendor claims, contracts, recovery commitments, dependencies, operating procedures, and testing evidence to help an agency determine whether the promised resiliency is supported.
When gaps are identified, PSCS can define or recommend a more resilient architecture and operating model. The vendor, implementation provider, or responsible agency team must then implement, operate, and test the required changes.
An independent review may examine:
- High-availability design and actual failure domains
- Disaster recovery architecture and activation procedures
- RTO, RPO, backup, and restoration commitments
- Application, database, identity, connectivity, and interface dependencies
- Availability definitions, exclusions, remedies, and contract language
- Monitoring, alerting, escalation, and incident communication
- Evidence from failover, restoration, and recovery testing
- Agency fallback procedures and operational continuity
- Shared responsibilities across the vendor, agency, and third parties
- Recommended architecture, contract, process, and testing improvements
This creates a practical distinction between accepting a claim and validating the systems, processes, commitments, and evidence behind it.
Questions Agencies Should Be Able to Answer
- What failures can high availability absorb automatically?
- What events require disaster recovery?
- What exactly is included in the availability calculation?
- How long can the complete service remain unavailable?
- How much recent data could be lost or unavailable?
- Do critical interfaces work in the recovery environment?
- Can the agency operate if cloud or local connectivity fails?
- When was restoration last tested successfully?
- Who owns each part of detection, failover, recovery, and communication?
- What evidence supports the vendor's resiliency claims?
Resiliency Is an Ongoing Operational Discipline
Resiliency is not a feature that can be purchased once and forgotten. Systems, interfaces, networks, personnel, cloud services, software versions, operating procedures, and vendor responsibilities change over time.
A design that was resilient at go-live may develop new single points of failure years later. A recovery procedure that once worked may no longer match the current environment. A contract commitment may not reflect the agency's present operational needs.
The goal is not to claim that public safety software will never be unavailable. No architecture can make that promise. The goal is to anticipate failures, understand operational impact, establish clear responsibilities, prove recovery capabilities, and prepare personnel to continue serving the public.
Whether software is on premises, in a vendor data center, or in a public cloud, the standard should remain the same: when something fails, can the agency continue operating—and has that capability actually been tested?
Related Public Safety Resiliency Resources
Independent Public Safety Software Resiliency Review
PSCS helps agencies evaluate whether vendor resiliency claims are supported by the architecture, operational processes, contractual commitments, dependencies, and testing evidence behind them.
Reviews can support procurement, contract negotiation, pre-go-live validation, outage follow-up, disaster recovery planning, and ongoing vendor accountability.