§ 01Continuity and Disaster Recovery

Timing a real failover: the steps that never get rehearsed

Recovery plans are written to be read by auditors and executed by nobody. This note has two sources: a failover declared against our own estate on 21 March, and years of recovery exercises the author ran on much larger estates for previous employers, described here with those employers unnamed. It does not report a client failover.

· 4 min read · Tobias Reinhardt, Principal Engineer, Continuity and Disaster Recovery

What this estate can and cannot test

Identity in one cloud tenant, a code repository, a case management system, a detection laboratory, and the encrypted storage that holds engagement evidence. The whole of it is a few terabytes, not hundreds. Nothing we did on 21 March proves we can move an imaging archive or a data warehouse between sites inside a clinical change window.

What this estate does test is the sequence, and the sequence is where plans fail. On the larger estates I worked on earlier in my career — a utility group's data centre team, later a hospital group's IT department — storage promotion was almost never the step that decided the clock. It is the most rehearsed step because it is the one the vendor documents. The steps that ran long were the ones nobody owned.

The exercise

We declared a simulated total loss of the primary region at 09:00 on a Saturday. The declaration was announced internally and to nobody else, and no client service was affected. One rule made it worth doing: the engineer at the keyboard was not the engineer who wrote the run book, and was not allowed to ask them anything. A run book that only works in the hands of its author is a personal memory with a document wrapped round it.

The steps that decided the clock

  • Authority to declare. The run book named one person. That person is sometimes asleep, on a plane, or the reason you are declaring. It now names three and states what happens when none of them answers within 15 minutes.
  • Credentials. The break-glass credentials for the recovery environment sat in a vault that authenticates through the identity tenant we had just declared lost. A circular dependency, written by us, reviewed by us, and invisible until someone tried to use it under the assumption that the tenant was gone. This cost us more than an hour and it is the most embarrassing finding in the exercise.
  • The run book's own location. Ours lived in a wiki hosted in the failed region. The offline copy existed and was 5 weeks out of date, so the engineer worked from a document that no longer matched the environment for the first two steps.
  • DNS. Records carrying a time-to-live of 3,600 seconds meant a change made at minute 40 did not reach every resolver for another hour. The records in the failover set now sit at 300 seconds permanently, and a scripted step lowers the rest 24 hours before any planned exercise.
  • Dependencies that do not move with you. Licence servers pinned to the old addresses, webhooks from external services still pointed at the failed endpoint, egress addresses allowlisted by a third party, and certificate issuance rate limits that punish you for reissuing under pressure. Each is trivial in isolation and each adds minutes.
  • Out-of-band communication. Our coordination channel authenticates through the same identity provider. If that is what failed, the people running the recovery cannot reach each other through the tool they run everything else in. There is now a channel agreed in advance that does not depend on it.
  • Failback. The plan documented the failover and stopped. Failback is where data written in the recovery environment either merges cleanly or is quietly lost, and it is the half that never gets rehearsed because the exercise is declared a success once service returns.

What the clock said

Service restored to the point where the team could work normally: 3 hours 41 minutes, against a target of 2 hours that we had written ourselves. Roughly two-thirds of the elapsed time went on the credential dependency and on waiting for DNS. Neither is a capacity problem, neither would have been found by another document review, and both had been reviewed on paper twice.

A plan that has never been executed is a document about recovery, not a recovery capability

What a timed failover has to produce

  • A recorded time for each step, taken from a clock rather than reconstructed afterwards, with the person who performed it named against it.
  • An abort point fixed before the exercise starts, and the authority to use it held by the organisation whose service is at risk, not by the engineers running the test.
  • The run book rewritten in the order the steps were actually performed, and signed by whoever performed them.
  • Evidence in a form an auditor can read without a narrative from the people who ran the exercise, including the failures.

What we have not proved

Volume. Nothing in our estate approaches the scale at which replication bandwidth, array behaviour or the sheer count of virtual machines becomes the constraint, and on the larger estates I ran exercises for, that scale is exactly what decided the recorded time. So the honest position is the one at the top of this note. We have timed our own estate and found six things wrong with it, and I have timed considerably larger ones under a previous employer. When we time a client failover, we will publish the number it produces.

Scoping is done by the director who will sign the report

The first scoping call is free. Where an assessment follows, it is a fixed fee agreed in writing before it starts.