Ask a team how they achieve zero-downtime deploys and you will hear about blue-green, canaries, health checks and connection draining. All useful. All about going forward.
Ask the same team what happens if the new version is subtly wrong, and the answer is usually a pause. That pause is where the outage lives.
Schema changes are one-way unless you make them otherwise
Code rolls back cleanly because the old artefact still exists. Data does not. A migration that drops a column, narrows a type or rewrites values in place has destroyed the state the previous version needed, and no amount of deployment tooling brings it back.
The discipline that fixes this is expand and contract: add the new shape, write to both, move reads across, and only remove the old shape once nothing has read it for long enough that you would have noticed. Each step is independently reversible, which is the entire point.
- Expand: add the new column or table, nullable and unused
- Backfill in batches, with the old path still authoritative
- Dual-write, then move reads, one caller at a time
- Contract only after a full observation window with no reads
The window between the two deploys
During any rolling deploy, two versions of your code are live at once. That is not an edge case, it is the normal state for several minutes, and every change has to be correct in that window.
This is why a change that renames a field and updates every reader in the same release is unsafe even though it looks atomic in the diff. For the length of the rollout, old readers are meeting new writes. The fix is boring and reliable: make the writer tolerant before you make the reader depend on it.
Rehearse the rollback, not just the deploy
Most teams have a documented rollback procedure that has never been executed against production-shaped data. It is a paragraph, not a capability.
Rehearsing it changes what you build. Teams that have actually rolled back a migration stop writing destructive ones, because they have felt what it costs. The rehearsal is cheaper as a scheduled exercise than as a discovery at two in the morning.
Feature flags are not a substitute
A flag turns behaviour off. It does not restore data that the new behaviour has already written in a shape the old behaviour cannot read.
Flags are excellent for controlling exposure and terrible as a data-safety mechanism. Use them to decide who sees a feature, and use expand-and-contract to decide whether you can go back.



