Offline, file-level deduplication on ZFS — without a dedup table

Beta — v0.9.2 . DedupCommando changes real files. Read the safety model before you apply anything, and keep backups.

ZFS can deduplicate on its own: set dedup=on and the pool drops duplicate blocks as they are written, at the cost of a dedup table that needs plenty of RAM to stay fast. DedupCommando takes the other route. It scans data already on your pools, finds byte-for-byte identical files, and turns the extra copies into hardlinks, block-cloned reflinks or quarantined deletes — one reviewed batch at a time, under a ZFS snapshot.

Built-in dedup and file-level dedup are different tools

The dedup property works on blocks, inline: where it is enabled, duplicate blocks are removed as data is written (zfsconcepts(7)). Each pool keeps one dedup table (DDT) with an entry per unique block, consulted for every dedup-able block written or freed; when it does not fit in memory, each miss costs a random disk read (workload tuning). OpenZFS recommends at least 1.25 GiB of RAM per 1 TiB of storage with dedup on, and advises against enabling it unless necessary.

OpenZFS 2.3 added fast dedup (the fast_dedup feature, for new dedup tables — zpool-features(7)), and the table can now be capped on disk (dedup_table_quota) and pruned of old unique entries (zpool ddtprune). The table gets cheaper but does not go away: dedup stays inline, block by block, for data written while it is on.

DedupCommando works one level up, with file operations — unlink, link, clone — not pool blocks (what it does not do).

dedup=onDedupCommando
When it runsInline, on every writeWhen you scan, review and apply a plan
What it matchesIdentical blocksWhole files, byte for byte
Data already on the poolLeft as it is; only new writes are deduplicatedWhat it works on
MemoryA pool-wide table, used on every write and freeOnly during a scan: about 2.5 KiB per hashed file while grouping, then released
Partly identical filesIdentical blocks inside them are foundNot found
Review before changesNone; it is automaticEvery group, then the batch plan

What offline, file-level means

Offline. No daemon, watch mode or inotify hook runs in the background. You start a scan yourself, or from cron with --scan, and no file is touched until you confirm a batch. Scans resume after an interruption, and a hash cache lets repeat scans skip unchanged files.

File-level. Each candidate is hashed whole with BLAKE3 (--verify adds a byte-by-byte comparison), identical files form a group, and one file per group is the keeper. There is no fuzzy matching. Files under 4096 bytes are skipped by default, and .zfs snapshot directories and the quarantine are never scanned (scanning). Groups are ranked by the space they would free; duplicate folders ("twin folders") show up too.

Three ways to reclaim space

Mark one keeper per group with F7, then an action on each copy you want to reclaim. Without a keeper, the actions are ignored.

For both link types, the replacement is built under a temporary name, the original moves into the quarantine, and the replacement takes over the path; if that step fails, the original is put back (atomic publication). In the project's acceptance test on scratch pools with OpenZFS 2.3.4 and 2.4.3, a reflinked file kept mode 0600, a non-root owner and a user xattr on its own inode while sharing the keeper's blocks.

No action frees space at once: the originals wait in the quarantine, and the batch's snapshot still holds their blocks. Hardlink vs reflink goes deeper on the two link types.

How each batch is protected

Requirements

On other filesystems a scan runs, but without snapshots applying actions is not recommended.

Checking that space came back

Space returns only after both safety nets are cleared. The Summary after an apply lists the snapshots, the quarantine folder and the exact commands to run; the manual suggests a week or two of normal use first (after apply):

zfs list -t snapshot | grep dedcom-    # the batches' safety snapshots
dedcom --purge-quarantine              # lists quarantined files and bytes, deletes nothing
dedcom --purge-quarantine --yes        # irreversible: deletes every quarantine, prints "Reclaimed: …"
zfs destroy tank@dedcom-<timestamp>    # irreversible: one per dataset in the batch

Until then, zfs rollback to the batch's snapshot undoes it for the whole dataset, including anything written since, and a single file comes back with an mv out of the quarantine (recovery). Space saved by reflinks shows in the pool's allocation rather than in a dataset's used space, and in the project's tests it took over a minute to appear.

FAQ

Does it find files that are only partly identical? No. It matches whole files, byte for byte, by BLAKE3 hash. Blocks shared between otherwise different files are what dedup=on catches and a file-level tool does not.

Can a reflink join files in two datasets of the same pool? No. dedcom clones within one dataset, as with a hardlink, because every ZFS dataset is a separate filesystem. For copies in different datasets, delete to quarantine is the way to reclaim the space.

Do reflink savings survive zfs send | zfs recv? Not in the project's test with OpenZFS 2.4.3: after a reflink pass the source pool held 926.1 MiB, while the receiving pool took 1285.7 MiB, about the size before the pass. Plan the receiving side for the full, undeduplicated size.

Next steps