A successful request can leave unfinished work
A Virtual Machine is more than its entry in a list. It has storage, network connections and permissions. Deleting it means cleaning up resources across several systems.
A service can accept a delete request while some of that work is still pending. Kubernetes makes this visible through finalizers, which can keep an object present until cleanup conditions are met.
Our revised lifecycle carries that distinction through to the administrator. Deletion starts a tracked operation; failures remain visible and can be retried. Removing a resource from the normal inventory no longer has to remove the information needed to finish its cleanup.
Do not lose the record of unfinished work
Reviewing our delete paths brought us back to a recurring problem: the record of a resource could disappear before all of its cleanup was complete. That made a later failure harder to recover from.
Permissions were part of the same problem. If an administrator can no longer reach the record, retaining it offers little help. An access error must not be interpreted as proof that the resource is gone.
A failed operation must preserve both its cleanup record and the administrator’s access to that record.
Keep intent separate from completion
We retain a record that deletion has been requested while cleanup proceeds. This is often called a tombstone: a record of a resource that is no longer in normal use.
Once deletion is requested, the resource leaves normal lists and quotas and rejects new changes. Its retained record remains available for inspection and retries. Administrators can tell that the resource is out of service without losing track of its unfinished teardown.
Cleanup steps tolerate being repeated. If an earlier attempt removed some dependencies and then failed, a retry can continue toward the same outcome. This avoids making an administrator reconstruct the operation from scattered records.
Retained recordOut of normal lists and quotas
Dependencies determine what can go next
Resources depend on one another. A workload may still be using a volume. A network address may still be attached to an interface. Cleanup has to respect those relationships.
Our teardown workflows respect those dependencies and wait for required cleanup to finish. Tenant offboarding extends the same approach across infrastructure, identity and supporting services. Network addresses return to their pools only after the objects using them are confirmed absent.
An unavailable service leaves a question unanswered. If the platform cannot inspect a subsystem, it cannot conclude that the subsystem is empty. Releasing an address too early, for example, could assign it to a new workload while the old one still uses it.
A retained record is not a backup
The record of a deleted resource may outlive the infrastructure itself. That is useful for auditing and explaining what happened, but it does not make a deleted disk recoverable.
Removing those retained records is a separate operation with its own retention policy and checks. It should not be confused with the teardown the user originally requested.
Secrets and encryption keys are handled separately. They keep their own recovery and destruction rules instead of following the general resource-retention policy.
Cleanup finds resources by their ID
People rename resources and reuse familiar names. The platform finds a resource’s infrastructure by its ID and a backing name it assigns, never by the display name. A rename, or a new resource that reuses a deleted one’s name, cannot point a delete at the wrong object.
That is one reason we review deletion as a complete lifecycle. Adding a resource type also means defining who owns its cleanup, how failure is reported and what establishes completion.
Model checking helps us examine ordering and retry failures. Cluster testing then checks the real infrastructure behavior that the models simplify.
Retryable teardown now, an offboarding report next
The retained-record lifecycle, dependency-aware cleanup, tenant offboarding checks and administrator-triggered purge are implemented. Administrators can follow and retry unfinished work across systems.
Next, we are validating the newer file-storage teardown end to end and building an offboarding report. The report will list what was removed when a tenant leaves, based on checked outcomes, and name any cleanup that could not be confirmed.
Established tools, broader responsibilities
Kubernetes finalizers and garbage collection handle dependencies represented in the cluster. A private cloud also manages systems outside it, so those mechanisms cover only part of the job.
Soft deletion, project cleanup tools and Crossplane offer useful precedents. Our concern is how the complete resource lifecycle behaves when one of its systems fails halfway through.
References
- Kubernetes finalizers
Keeping an object present until its cleanup conditions are met.
- Kubernetes garbage collection
Owner references and deletion of dependent objects.
- OpenStack ospurge
The project resource cleanup tool.
- Crossplane documentation
Managed Resources.
