Lumy Labs
← All insights

Cloud

Zero downtime is a property of your rollback

Teams rehearse the deploy and improvise the rollback. That is backwards: the deploy is the step you can retry, and the rollback is the one you cannot.

Lumy Labs7 min read
A cockpit control panel photographed close up

Ask a team how they achieve zero-downtime deploys and you will hear about blue-green, canaries, health checks and connection draining. All useful. All about going forward.

Ask the same team what happens if the new version is subtly wrong, and the answer is usually a pause. That pause is where the outage lives.

Schema changes are one-way unless you make them otherwise

Code rolls back cleanly because the old artefact still exists. Data does not. A migration that drops a column, narrows a type or rewrites values in place has destroyed the state the previous version needed, and no amount of deployment tooling brings it back.

The discipline that fixes this is expand and contract: add the new shape, write to both, move reads across, and only remove the old shape once nothing has read it for long enough that you would have noticed. Each step is independently reversible, which is the entire point.

  • Expand: add the new column or table, nullable and unused
  • Backfill in batches, with the old path still authoritative
  • Dual-write, then move reads, one caller at a time
  • Contract only after a full observation window with no reads

The window between the two deploys

During any rolling deploy, two versions of your code are live at once. That is not an edge case, it is the normal state for several minutes, and every change has to be correct in that window.

This is why a change that renames a field and updates every reader in the same release is unsafe even though it looks atomic in the diff. For the length of the rollout, old readers are meeting new writes. The fix is boring and reliable: make the writer tolerant before you make the reader depend on it.

Rehearse the rollback, not just the deploy

Most teams have a documented rollback procedure that has never been executed against production-shaped data. It is a paragraph, not a capability.

Rehearsing it changes what you build. Teams that have actually rolled back a migration stop writing destructive ones, because they have felt what it costs. The rehearsal is cheaper as a scheduled exercise than as a discovery at two in the morning.

Feature flags are not a substitute

A flag turns behaviour off. It does not restore data that the new behaviour has already written in a shape the old behaviour cannot read.

Flags are excellent for controlling exposure and terrible as a data-safety mechanism. Use them to decide who sees a feature, and use expand-and-contract to decide whether you can go back.

Let's talk

Living with this problem?

If this one landed close to home, tell us where you are stuck and we will tell you honestly whether we can help.

Or email start@lumylabs.co

What happens next

  1. A real replyWithin 48 hours
  2. Before specificsNDA first
  3. Not a sales pitchA real proposal