Search the docs

Troubleshooting: Removal and Decommissioning

Data Not Transferred Before a Site Is Removed

Removing a participant with gfs-participant-remove does not move that site’s data ownership anywhere first — this is a known limitation of the supported removal flow today (5.2/5.3), not a bug with a pending fix. If the site being removed solely owned some data, that data becomes unavailable once the site is gone. See the warning in Removing a participant from a share.

OSV Stuck in "cleaning" / Decommission Stuck

Symptom: shared-OSV decommission stalls, often near a high percentage with few chunks remaining, or stays in a cleaning state. One specific cause is confirmed (below); other causes are not yet confirmed by engineering.

Known case: stuck at a fixed low percentage (5.3 regression, fixed in 5.3.1)

One confirmed cause (HS-42342): object-volume-remove could get stuck reporting a fixed non-zero instance count — for example EXECUTING at 38% with a Decommissioning: 1 instances remaining status message — even after the object actually responsible had already been cleaned up. This happened if the volume held a corrupt/PDL (partial-data-loss) object when removal started: the internal cleanup process could give up and never re-check once that object was later removed, leaving the task looking permanently stuck rather than slowly progressing.

This was a 5.3 regression, introduced earlier in the 5.3 line and fixed in 5.3.1 — it did not affect 5.2. Our 5.3 test build (5.3.1-1201) already includes the fix, so this specific case could not be reproduced live during this pass. If a decommission is reported stuck at an unchanging percentage with a nonzero remaining-instance count on an early 5.3.0/5.3.1 build, and a corrupt/PDL object was involved, this is a likely explanation — upgrading resolves it without needing to cancel or restart the removal. This is one specific, confirmed cause of this symptom; other causes are still pending engineering confirmation per the CAUTION above.