Cloud outages tend to begin with something boring.
A dashboard turns yellow. Database latency creeps upward. Someone posts a screenshot in Slack and asks, “Is anyone else seeing this?” Five minutes later, appointment updates are failing and a support ticket says a patient can’t open a document.
Now it’s an incident.
For a healthcare app, recovery gets complicated fast. Bringing the servers back is only part of the job. Patient data has to survive the failure, access controls still need to work, and the team needs some confidence that the restored system actually represents what happened before everything went sideways.
The dangerous assumption is that having backups means you’re ready.
Your backup job can be healthy while your recovery plan is broken
Green check marks are comforting. I wouldn’t trust them on their own.
Take a fairly ordinary healthcare SaaS stack. PostgreSQL holds patient and appointment data. Files live in object storage. Redis handles temporary state. Authentication comes from another service. Audit events stream somewhere else. A queue handles notifications and background work.
Now imagine the database gets corrupted during a deployment.
The team restores a recent snapshot and gets PostgreSQL running again. The admin dashboard loads. Everyone feels the tension drop a little.
Then the strange problems start.
A consent form uploaded that morning won’t open. A cancelled appointment reappears. Several background jobs fire twice. Audit records stop a few minutes before the failure, so nobody can easily tell which actions happened before the restore point and which happened after it.
The database recovered. The product didn’t.
That’s the distinction worth caring about.
The infrastructure underneath the application also needs to survive recovery without turning into a security free-for-all. HIPAA web hosting can cover pieces such as encrypted backups, monitoring, disaster recovery, and a Business Associate Agreement, while the application team still has to account for files, keys, queues, APIs, and the other dependencies sitting above that layer. If half the stack is missing, a running server doesn’t help much.
A useful exercise is to pretend the primary environment no longer exists.
Not “the database is unavailable.” Gone.
Could someone build the system somewhere else? Do they know which infrastructure is created from code and which pieces were configured by hand? Are the certificates documented? What happens to DNS? Where are the secrets? Can the restored application decrypt yesterday’s files?
There’s usually one answer that makes me nervous: “Dave knows how that works.”
Dave might be on vacation.
RTO looks very different once you start a stopwatch
Recovery time objectives have a habit of sounding better in planning meetings than they do during real recovery work.
An RTO of one hour feels responsible. So does a 15-minute recovery point objective. Put them in a spreadsheet and everything looks under control.
Then try restoring the system.
The database takes 22 minutes. The standby environment can’t access the encryption key. Someone fixes that and discovers a security group is blocking traffic. Authentication works next, but the TLS certificate on an internal endpoint expired because nobody had used that environment in eight months.
Forty-five minutes are gone and patients still can’t use the application.
This happens because teams often set recovery targets around the most obvious component instead of the full user workflow.
A better question is: how long until a receptionist can book an appointment again?
Or until a clinician can retrieve the record needed for the next patient?
Follow that workflow backward and the dependency list gets interesting. There may be a database, authentication provider, storage bucket, API gateway, certificate, queue, search service, DNS record, encryption key, and several bits of infrastructure nobody mentioned during the RTO discussion.
The same problem appears during cloud migrations. PlainEnglish’s piece on deciding what belongs in the cloud and what should remain on-prem gets at an important point: applications make more sense when you look at dependencies rather than boxes on an architecture diagram.
Recovery planning benefits from the same treatment.
Not every part of the system needs the same target either. A reporting index that can be rebuilt overnight doesn’t belong in the same category as current appointment data. Cached recommendations can disappear completely. Audit records are another story.
NIST’s contingency planning guidance connects recovery decisions with business impact and system criticality. In practice, that means spending your fastest recovery budget on the things people genuinely can’t work without.
And yes, recovery speed costs money.
Keeping another environment warm is more expensive than restoring one from scratch. Cross-region replication costs more than a nightly backup. Better recovery usually means accepting some amount of idle infrastructure, additional storage, engineering time, or operational complexity.
The important part is making that trade intentionally.
The first person to test the runbook shouldn’t be the person who wrote it
Recovery documentation has a funny failure mode: it makes perfect sense to its author.
Of course step six says “restore application secrets.” The engineer who wrote it knows which secrets, where they live, and which role has permission to retrieve them.
Hand the same document to somebody else and step six suddenly becomes a 40-minute investigation.
That’s useful.
A recovery test should create a little discomfort. You want the awkward questions while production is healthy and everyone has coffee nearby.
Give another engineer the runbook and a clean environment. Tell them the primary database is gone at 10:17 a.m. Their job is to recover the service without asking the runbook’s author what any of the steps “really mean.”
Watch where they stop.
Maybe they can’t find the correct snapshot. Maybe the infrastructure code creates most of the environment but leaves DNS out. Maybe a backup restores successfully but points the app toward an old storage location.
These failures are much cheaper during a drill.
There’s another question people forget to ask after the app loads: is the data right?
Open a few records. Check recent transactions. Pull up an uploaded file. Compare counts. Look at the audit trail. Trigger a normal workflow from beginning to end.
A 200 OK response proves the web server answered. That’s about it.
Ransomware adds another unpleasant wrinkle. The latest backup isn’t always the backup you want. If an attacker had access for three days before anyone noticed, yesterday’s snapshot could contain compromised files or unwanted changes.
HHS discusses backups and recovery as part of its ransomware guidance. The operational question is simple: can your team identify a known-good recovery point, or will they automatically restore whatever has the newest timestamp?
Encryption can slow things down too.
A well-protected production environment may have strict permissions around its keys. That’s good until someone builds a fresh recovery environment and realizes the new roles can’t decrypt anything. PlainEnglish’s guide to AWS KMS and encryption controls shows why key management needs to be treated as part of infrastructure, not an accessory bolted onto it.
The tempting emergency fix is to broaden permissions.
That’s exactly when mistakes happen.
Your recovery process should already know which identities need which keys. An outage is a terrible time to start improvising IAM policies.
The awkward failures are the ones worth planning for
I like clean failover diagrams as much as anyone. Primary region on the left, standby region on the right, neat arrow in the middle.
Reality has more smudges.
A system might be half-working. Reads succeed, writes fail intermittently. One API is healthy while another returns old data. Retries pile up in a queue. A developer wants to fail over immediately, while someone else is worried that doing so will replay transactions twice.
There may not be a single moment when the system is clearly “down.”
That’s why high availability and recovery aren’t interchangeable.
Replication helps when hardware fails. It can also replicate corruption beautifully.
A bad deployment can damage both sides. So can a compromised administrator account. An application bug might delete data through perfectly legitimate API calls while every monitoring dashboard says the infrastructure is healthy.
Different failures need different escape routes.
Versioned storage helps with overwritten files. Isolated backups give you somewhere to go when production credentials are compromised. Infrastructure as code means fewer settings need to be reconstructed from memory. Strict access controls reduce the amount of damage one credential can cause.
PlainEnglish’s discussion of defense in depth applies just as well to recovery. You want several imperfect controls catching different mistakes rather than one supposedly perfect safety net.
The most useful recovery plans also describe what happens before full service returns.
Maybe clinicians can temporarily view existing records but not edit them. Perhaps appointment requests can enter a holding queue. A billing workflow might stay offline because processing transactions against uncertain data would make cleanup worse.
Those decisions are easier at 2 p.m. during planning than at 2 a.m. with phones buzzing on the desk.
“Back online” shouldn’t simply mean the health check turned green.
It should mean people can use the application again without creating another incident.
Wrap-up takeaway
Most teams don’t discover a missing backup. They discover that the backup covered less of the application than they thought. The database comes back, then somebody notices the files are missing, the keys won’t work, or one forgotten service still points at the dead environment. That’s the kind of problem a restore test is supposed to uncover. Pick one workflow that actually matters to users, such as opening a patient record or booking an appointment, and trace everything it needs from login to completion. Then ask the uncomfortable question: if the primary environment disappeared this afternoon, could the team rebuild all of it without calling the one person who remembers how it was set up?
Further Reading
Discover more articles on similar topics across our network
Why Manufacturing Control Matters for Stable Wax Performance in Salons
Venture
How You Can Build a Career Teaching the Next Generation of Nurses
Venture
Comments
Loading comments…