Safety, recovery, and honest limitations

Beta (v0.9.2 ). DedupCommando deletes and relinks real files, so it is built around layered safeguards. This page is a summary: the authoritative reference is Safety, recovery and limitations, and chapter 3 of the manual walks through every command. Read one of them before your first apply, and keep backups.

Is it safe to delete duplicate files?

Only if a mistake can be undone, and DedupCommando is designed around that. A scan never modifies your files: the walk, hash and group phases change nothing on the filesystem. Nothing is deleted or relinked until you review the plan and confirm it with Y, and even then a "delete" is a move into a quarantine, made under a fresh ZFS snapshot.

Layers of protection

Every destructive batch (delete, hardlink or reflink) runs behind these guardrails:

Recovering deleted duplicates: one file or a whole batch

The Summary screen at the end of an apply lists the snapshots it created and the quarantine directory it used. Note both.

Restore a single file without a rollback

You do not need a rollback to bring back one file. Find it in the quarantine and move it to its original path:

find /tank/.dedcom-quarantine -type f -name 'photo.jpg'
mv /tank/.dedcom-quarantine/20260527-143215-512874000-4821-0/media/photo.jpg /tank/media/photo.jpg

If a hardlink now occupies that path, remove it first with rm /tank/media/photo.jpg (this removes only the link, not the original in the group), then run the mv. See bringing back one file.

Roll back a whole batch

If the whole batch was wrong, roll back each affected dataset to its snapshot:

zfs rollback tank@dedcom-<ts>
zfs rollback tank/media@dedcom-<ts>

Warning: a rollback returns the entire dataset to the moment of the snapshot, so everything written to it since, by any program, is lost too. If anything else wrote to the dataset after the batch, restore the files you need from the quarantine instead; ls /tank/.dedcom-quarantine/<ts>/ shows what was evacuated (details).

If an apply stopped early

Esc stops an apply only after the current action finishes; the process can also be cut off between actions. Either way the snapshots already exist, the originals of finished actions are in quarantine, and the actions not yet reached keep their marks. Press F11 again to finish, or roll back as shown above (details).

Purging the quarantine safely

Quarantined files and dedcom snapshots keep using pool space until you remove them, and removing them cannot be undone.

  1. Wait. The manual recommends one to two weeks of normal use, until you are sure the result is stable.
  2. Preview. dedcom --purge-quarantine lists the .dedcom-quarantine directory of every detected dataset with its file count and total size. It deletes nothing.
  3. Purge. dedcom --purge-quarantine --yes deletes those directories: a final rm -rf of every batch in every dataset, with no selective mode.
  4. Destroy the snapshots separately. The purge does not touch them. List them with zfs list -t snapshot | grep dedcom- and remove each with zfs destroy. Space is freed only after both steps.

After a hardlink, the target's own owner, permissions, ACLs and xattrs survive only on the original in quarantine, so a purge removes them for good (details).

What the protections do not cover

Other limits: DedupCommando is Linux-only (x86_64 or aarch64, kernel 3.15 or newer), there is no headless apply, and it typically runs as root. Hardlink and reflink both work within a single dataset, and reflink also needs ZFS 2.3 or newer with block_cloning. See the limitations and what the guardrails do not cover.

Before your first apply: a checklist

  1. Read chapter 3; the manual calls it mandatory before the first real apply.
  2. Rehearse on a test pool: scripts/make-test-pool.sh from the release bundle creates /testpool on a file image without touching real disks, and teardown-test-pool.sh removes it (details).
  3. Run as root with zfs in PATH. Without snapshot privileges, applying is not recommended.
  4. Check that each root lies entirely on ZFS: df -T <path> should report zfs (details).
  5. Scan live data with the Idle profile (G in the scan configuration cycles to it).
  6. Make sure no other dedcom is running.
  7. Pick a maintenance window: writing workloads stopped, no zfs send of the same datasets.
  8. In the F11 confirmation, read the By type line and the listed paths, which reveal a batch that marked the wrong side of a group, and press S to save the plan as a .sh audit trail before Y (details).
  9. Start dedcom --strict-verify if something else might be editing the files.
  10. Afterward, re-scan and compare the two scans with ScanDiff; anything under Modified is a red flag (details).

Next steps