DedupCommando FAQ
Short answers for ZFS and Proxmox VE administrators, each linked to the manual section with the detail. DedupCommando is beta software (v0.9.2 ): read the safety guide before applying anything, and keep backups.
Why not just turn on ZFS dedup (zfs set dedup=on)?
zfs set dedup=on is block-level deduplication inside the pool; the manual calls it usually unnecessary and notes it eats RAM and CPU. DedupCommando works at the file level instead: it finds byte-for-byte identical files and reclaims their space with ordinary filesystem operations (delete to quarantine, hardlink or reflink). There is no always-on dedup table and no background daemon, and nothing changes until you confirm a plan. See what DedupCommando does not do.
Is it safe to use on a production host?
A scan changes nothing on the filesystem, and the Idle profile (one thread, nice 19, ionice idle), mandatory on live data, keeps it from starving VMs and backups. Duplicates are deleted or relinked only when you apply, and every batch runs under a per-dataset ZFS snapshot, with deleted files moved to a quarantine and each action revalidated first. DedupCommando is still beta software, so rehearse on a test pool, apply in a maintenance window, and keep backups. See the Idle profile on production data and safety and recovery.
What happens if a file changes between the scan and the apply?
Right before each action, dedcom re-checks the target and the keeper: a symlink check and a size check every time, then the content (in the default Hybrid mode each distinct file is hashed once per batch; --strict-verify re-hashes before every action). If anything differs, that one action is canceled with a reason such as changed after the scan (content), and the rest of the batch continues. Canceled items keep their marks, so you can re-scan, check them and press F11 again. Revalidation does not warn you before F11, so on large datasets apply when writing workloads are stopped. See revalidation.
What exactly is kept when a duplicate becomes a hardlink or a reflink?
A hardlink makes the duplicate's path point to the keeper's inode, so owner, permissions, ACLs and xattrs become the keeper's, and a write through one path is seen at every path. A reflink is a new inode that shares the keeper's blocks: dedcom writes the replaced file's owner, mode, ACLs, extended attributes and timestamps onto it before publishing, or cancels the action if it cannot. Editing a reflinked copy leaves the others unchanged. Both work only within one dataset, and in both cases the original goes to quarantine first, where it keeps its own metadata until you purge. See hardlink and reflink.
Why does it need root?
Taking ZFS snapshots and scanning outside your home directory need privileges, so dedcom typically runs as root, which is the default on Proxmox VE. As an unprivileged user the snapshots are unavailable unless you set up zfs allow or sudo, and without snapshot safety, applying actions is not recommended. The tool may also fail to see your pools when zfs runs unprivileged. See installation and dedcom does not see my pools.
Does it work without ZFS?
Scanning does: on non-ZFS filesystems the walk and hash run fine. There is no snapshot insurance there, so applying actions is not recommended, and a delete is canceled when dedcom cannot determine the file's dataset. Reflink additionally needs ZFS 2.3 or newer with block_cloning active. Before applying, check with df -T <path> that every root lies entirely on ZFS. See what DedupCommando does not do.
What was it tested on?
It was developed and tested against ZFS pools including Proxmox VE, and is tested on Proxmox VE 9.1 with OpenZFS 2.3. It runs on Linux x86_64 and aarch64 with kernel 3.15 or newer. The pre-built packages and binaries need glibc 2.39 or newer (Debian 13, Ubuntu 24.04, Proxmox VE 9 or newer), so on Proxmox VE 8 or Debian 12 you build from source. See requirements and installation.
How do I undo a delete?
A delete is a move into the dataset's quarantine, so undoing it for one file is a plain mv back to the original path:
find /tank/.dedcom-quarantine -type f -name 'photo.jpg'
mv /tank/.dedcom-quarantine/<ts>/media/photo.jpg /tank/media/photo.jpg
To undo a whole batch, run zfs rollback <dataset>@dedcom-<ts> for each affected dataset, but this reverts the entire dataset, including everything else written since. After dedcom --purge-quarantine --yes, the quarantine copy is gone for good. See bringing back one file.
Can I run it from cron, without the UI?
Yes, except for applying: --scan, --stats, --compact-db, --export-csv and --purge-quarantine run without the TUI, exit non-zero on error and never prompt, so a writing mode just fails if another instance holds the lock. Applying is interactive by design: it happens in the F11 confirmation, where you can also save the plan as a .sh script to review or run by hand. For a nightly scan on a live host, set the Idle profile once in the TUI; later --scan runs reuse it. See the cron example and why there is no headless apply.
How much memory does a big scan need?
The walk and hash phases need only tens of MB. The grouping phase peaks at about 2.5 KiB per hashed file with the default algorithm (roughly 2.4 GiB for 1 million files, 4.8 GiB for 2 million, 12 GiB for 5 million), then returns it. Before that phase, dedcom compares its forecast with free RAM and warns you if it does not fit. --merkle-dirs produces the same groups with memory proportional to tree depth, typically tens to hundreds of MB, and the manual calls it required at 5 million files. See estimating the memory peak.
Should I run it on backup storage or virtual machine disks?
No: leave other programs' stores out of the scan. That means a Proxmox Backup Server datastore; a restic, borg, kopia, Arq or Duplicacy repository; a Time Machine sparse bundle; a virtual machine's disk. Such a store needs every one of its files at its own path with its own content, and restic, borg, kopia, Duplicacy, Arq and Proxmox Backup Server already store each piece of data only once. A delete inside a store breaks it at once, and a hardlink makes two copies one file, so a program that later writes in place changes both. A scan root takes in every dataset mounted below it and there is no way to exclude a path, so choose roots that do not reach any store. A folder of plain copies you made yourself is not a store. The manual covers this in section 8.9.
How do I verify a download?
Each release on GitHub Releases ships a SHA-256 checksum, a minisign signature, a CycloneDX SBOM and a SLSA build-provenance attestation. Check the tarball with sha256sum -c dedcom-<version>-<triple>.tar.gz.sha256, its signature with minisign -Vm dedcom-<version>-<triple>.tar.gz -p minisign.pub, and its provenance with gh attestation verify. If you install from the APT repository, apt checks the GPG-signed repository metadata for you. See verifying releases.
Is it free, and what is the license?
Yes. DedupCommando is free, open-source software released under the Apache-2.0 license. Each release tarball includes the LICENSE, NOTICE and THIRD-PARTY-NOTICES files, and the SBOM lists the licenses of its dependencies. See who wrote it and for whom.
Next steps
- Documentation home: install, first scan and guides.
- User manual: every chapter, from installation to troubleshooting.
- Safety, recovery and limitations: the authoritative safety reference.