FC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 QFC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 Q
Tuesday, 29 September 2026 · Oslo · London · New York

Enterprise · Long-form Report

Why Backup Systems Fail When They Are Finally Needed

Backups run nightly for years and appear healthy on every dashboard. Then a restore is actually required, and the organisation discovers the archive is incomplete, unencrypted in the wrong places, or untestable at scale.

Tall storage racks filled with hard drives in a dim data hall, cabling running in loose bundles between cabinets, shot in high-contrast black and white.
Tall storage racks filled with hard drives in a dim data hall, cabling running in loose bundles between cabinets, shot in high-contrast black and white.

Independent coverage

M

By Michelle Greenlee

Contributing Writer — Enterprise Tech / Cybersecurity · Freelance

Edited by Dr. Elin Lindqvist, MD

Published 28 September 2026

10 min read

Evidence: Analysis

Backup is one of the oldest disciplines in enterprise IT, and one of the least examined. Most organisations run nightly jobs, retain copies for months or years, and report a green status in every audit. The assumption is that a green backup report implies a working recovery capability. The gap between those two things is where data loss actually happens.

The failure pattern is consistent across industries. An organisation suffers a ransomware event, a storage array failure, or a corrupted application upgrade. The response team opens the backup console, selects a recovery point, and begins the restore. The job that has run successfully every night for years now fails, or succeeds in a way that recovers far less than the business expects. The discrepancy is not bad luck. It is the predictable result of how backup systems are operated, funded, and audited.

Backup jobs measure writes, not recoverability

Backup software reports on what it can observe: whether the job ran, how much data was read from the source, and whether the copy landed in the target. A successful job means the write happened. It says nothing about whether the written data can be read back, reassembled, and mounted by an application under time pressure.

The difference matters because modern environments are full of conditions that break restores while leaving backups green. An application's configuration database is excluded from the job by a filter someone added three years ago. A virtual machine grew beyond the size the backup policy assumed, and the job now processes it with incremental-forever logic that has never been consolidated. A SaaS platform changed its API, and the connector that backs up the tenant silently captures only metadata. Every one of these produces a healthy report and an unusable recovery point.

The only test that measures recoverability is a restore. Organisations that restore regularly into an isolated environment find these failures in controlled conditions. Organisations that do not, find them during the incident, when the cost of the discovery is highest.

Restore performance is never engineered

Backup infrastructure is sized for the nightly window. The job has eight hours to copy the day's changes to storage, and the network, media servers, and deduplication appliances are dimensioned for that throughput. Restore requirements are different in kind, not just direction. A full recovery of a large database may need to move terabytes in hours, to a different set of hosts, with the application waiting and the business watching.

Because restore performance is rarely tested, the bottlenecks are invisible until they matter. The deduplication appliance that reads back at a fraction of its write speed. The catalog server that becomes a single point of serialisation for thousands of concurrent restore sessions. The staging disk that is large enough for one recovery, not ten running in parallel. Each of these turns a documented recovery time objective into an aspiration. A recovery that was estimated at four hours on a whiteboard takes two days when it runs through infrastructure that was built for a different workload.

Application consistency is the hidden dependency

A file-level backup captures bytes. A usable recovery requires that the bytes form a consistent state for the application that will consume them. This is the hardest part of backup engineering, and the part most often deferred.

Databases need their transaction logs captured in sequence with the data files. Virtualised workloads need quiesced snapshots at the hypervisor layer. Containers need the stateful components identified and backed up separately from the stateless ones. SaaS platforms need their export APIs called with the right scopes, on schedule, with the output stored in a format the platform can reimport. When any of these consistency mechanisms is missing, the restore completes and the application does not start, or starts with silent corruption that surfaces days later.

The failure is especially common in hybrid estates. A recovery plan written for the on-premises data centre does not cover the SaaS applications that now hold customer records, and the SaaS vendors' own retention windows are shorter than most administrators assume. Deleted data in a cloud collaboration platform is often unrecoverable after thirty or ninety days, regardless of what the corporate backup policy says.

Backup credentials are the first thing attackers take

Ransomware has changed the threat model for backup. Modern operators do not encrypt first and negotiate later. They map the environment, elevate privileges, and delete or corrupt the backup infrastructure before triggering the encryption event. Backup servers are frequently domain-joined, administered with shared credentials, and reachable from the same networks as the workloads they protect.

An attacker with domain administrator rights can disable backup agents, delete retention copies, and clear job history in minutes. The nightly reports stop, but nobody investigates, because backup alerts route to a low-priority queue. Days later, the encryption event runs against an organisation that no longer has a recovery path.

The defences are understood but unevenly applied. Backup infrastructure should hold separate credentials, in a separate identity boundary, with immutable or object-locked retention that even administrators cannot delete before the retention period expires. Many organisations buy these capabilities and never enable them, because enabling them changes operational workflows and requires a level of coordination that never reaches the top of the backlog.

Retention policy drifts away from reality

Retention policies are written once and rarely revisited. The policy says critical systems are retained for twelve months and tested quarterly. The reality is a mosaic of legacy jobs, each configured by a different team during a different project. Some systems are retained for thirty days because a job was cloned from a template. Others are retained for seven years because a compliance requirement from a closed project was never reviewed. Some critical applications have no job at all, because their owners assumed the platform team covered them.

This drift is invisible in normal operations and decisive during recovery. The team restoring a compromised system discovers that the last clean recovery point is older than the business can tolerate, or that a tier of the application was never in scope. The retention audit that would have caught this is a paperwork exercise in most organisations, performed against the policy document rather than the actual job configuration.

Restores are tested on too small a scale

The organisations that do test usually test small. A single file is restored to verify the pipeline works. This test proves almost nothing about a disaster recovery event, which involves restoring dozens of interdependent systems in a coherent order, into an environment that may not exist yet, with a recovery team working from documentation that may not match current architecture.

A meaningful test recreates the problem at scale. It restores a complete application stack, including its dependencies, identity integration, and network paths, into an isolated segment, and measures how long it takes and what breaks. The first such test in a mature organisation routinely fails, and the failures are instructive: missing configuration, undocumented dependencies, expired credentials, and recovery documentation that describes a previous version of the environment. Each failure found in a test is one that does not have to be discovered during the incident.

What resilient backup operations actually look like

The organisations that recover well share a few practices. They separate the backup infrastructure from the production identity domain, so that a compromise of production credentials does not reach the copies. They enable immutability on at least one retention tier. They run scheduled restores of full application stacks, not single files, and they track the measured recovery time against the stated objective. They include SaaS platforms in the same discipline as on-premises systems, with export jobs and restore drills. They treat the retention configuration, not the policy document, as the source of truth, and audit it automatically.

Above all, they accept that backup is a recovery capability, not a storage discipline. The nightly job is only the raw material. The capability is proven by how often restorations are performed, how long they take, and how much of the business they actually recover. An organisation that has never restored a critical system under time pressure does not know whether it has a backup system. It only knows that it has a write schedule.

"A backup that has never been restored is not a backup. It is a scheduled write with an unproven recovery attached to it."

Sources

Published 28 September 2026