A system is ready for production only when the team can restore a known-good state within a measured, acceptable window. Before the ops lead signs off, require a rollback procedure that names the trigger, owner, recovery point, commands, dependencies, verification checks, and maximum recovery time.
The uncomfortable moment usually comes near the end of a deployment review. The new path has passed its tests. Dashboards exist. Owners have approved the change. Then someone asks what happens if customer records are altered before the fault becomes visible.
“Deploy the previous version” is rarely a complete answer. Old application code may expect an old schema. Queued jobs may keep processing under new assumptions. A third-party service may have received irreversible instructions. An AI agent may have performed a business action correctly according to a flawed prompt or policy.
At that point, the team does not have a rollback plan. It has a hope.
A previous release may not be a known-good state
Rollback discussions often focus on binaries, containers, or feature flags because those are easy to picture. Production state is harder.
A release can change database schemas, cached values, permissions, generated files, search indexes, message formats, and external records. Returning the application to yesterday’s build does not automatically return those systems to yesterday’s condition. It can make the failure harder to diagnose by placing old code on top of newly changed data.
Define “known-good” before deployment. It should include a specific software version, configuration set, schema version, data recovery point, permission model, and set of dependent services. If any component cannot be restored, say so plainly. The response may need to be containment and forward repair rather than rollback.
That distinction matters more as autonomous software gains permission to execute business functions. Snowflake’s launch of an autonomous AI platform, designed to give enterprise users access to company data and the ability to perform business functions, illustrates the operational question. Access can be revoked quickly. Completed actions may be much harder to reverse.
Teams evaluating systems like this should ask which actions are transactional, which are compensatable, and which are irreversible. A generated summary can be discarded. A modified customer record may require reconstruction. A message sent to a customer cannot be unsent in any meaningful operational sense.
Write the recovery path as executable work
A useful rollback plan starts with a trigger. “If there are problems” leaves the decision open precisely when evidence is incomplete and pressure is rising.
Choose observable thresholds instead. These might include a defined error rate, failed reconciliation check, unexpected permission change, data mismatch, or increase in a named queue. Record who can call the rollback and who performs it. If six people must agree during an incident, approval latency belongs in the recovery estimate.
The plan should then identify the exact recovery sequence:
- Stop new work from entering the affected path.
- Preserve logs, request identifiers, configuration, and relevant state for investigation.
- Disable the release or restore the named prior version.
- Reverse compatible schema and configuration changes.
- Replay, repair, or quarantine work created during the affected window.
- Verify recovery with customer-facing checks, not deployment status alone.
Each instruction needs enough detail for someone other than its author to execute it. “Restore the database” hides the important questions: Which snapshot? Where is it stored? How long does restoration take? What writes will be lost? Can recovery happen in place, or does it require a new environment?
The same discipline applies to incident materials. The 2 A.M. Alert Packet explains why critical context must exist before the alert arrives. A rollback document should live beside that packet, use the same service names as the dashboards, and point to commands that have been tested recently.
Measure recovery before promising it
Recovery time objectives written in a planning document can become fictional through neglect. The only credible number comes from rehearsal.
Run the rollback in an environment that resembles production closely enough to expose dependency order, data volume, permissions, and timing. Record the duration from the rollback decision to verified service recovery. Include waiting time for approvals, credentials, artifact downloads, database restoration, cache rebuilding, and validation.
Then test the edge cases most likely to turn a five-minute estimate into an hour:
- The previous artifact is unavailable or fails its integrity check.
- The operator with the required permission is offline.
- The old release cannot read the current schema.
- Jobs created by the failed version remain in a queue.
- External actions completed before containment.
- Monitoring reports healthy while a core user task still fails.
A failed rehearsal is useful evidence. It reveals that the recovery claim exceeded the recovery capability while the stakes are controlled.
Make one person prove the return trip
Before the ops lead leaves the review, assign one person to run the documented procedure without coaching from its author. Give them the same access an on-call operator would have. Start the clock at the decision to roll back and stop it only after the defined customer-facing checks pass.
If they need an undocumented credential, a command from someone’s shell history, or an engineer who happens to remember the old schema, record the gap and block the release until it is resolved. If an action cannot be reversed, add a containment step, a repair owner, and a clear estimate of the affected window.
Finish the review with four facts written where the on-call team can find them: the last known-good state, the rollback trigger, the person authorized to act, and the measured recovery time. The final line should be a timestamp from the latest successful rehearsal.
Comments
No comments yet.