Find and delete duplicate files on Linux, safely

DedupCommando (command dedcom) is a duplicate file finder for Linux, with a multi-panel terminal interface and a headless mode for scripts and cron. It finds files that are byte-for-byte identical, whatever their names, and helps you reclaim the space: move the copies to a quarantine, or replace them with hardlinks or reflinks. It is built for ZFS, where every batch of changes runs under a snapshot.

DedupCommando v0.9.2 is beta software that changes real files. Read the safety model before applying anything, and keep backups.

How DedupCommando finds duplicate files

To find duplicate files on Linux, a scan runs in three phases and keeps its results in an SQLite database, dedcom.db:

  1. Walk. Traverse the roots you chose, lstat every entry and record the file list: minutes per million files.
  2. Hash. Read the files and compute a BLAKE3 hash of their content. On hard disks, files are read in (device, inode) order, which cut cold-scan time by 21% on a two-HDD pool. This is the long phase: hours for two million files on hard disks.
  3. Group. Collect files with the same hash into duplicate groups, and build directory signatures for twin folders. Seconds to minutes, with a temporary memory peak of about 2.5 KiB per hashed file. Before this phase dedcom compares its forecast with free RAM; --merkle-dirs brings the peak down to tens or hundreds of MB.

Matching is exact: the same BLAKE3 hash, with no fuzzy or similar-image matching. For an extra check, --verify compares each group byte by byte after hashing, which reads every file twice. A file that fails to read is not counted as a duplicate.

By default, files smaller than 4096 bytes are skipped, and --include-ext jpg,heic limits a scan to some extensions. ZFS snapshot directories (.zfs) and DedupCommando's own quarantine are always skipped. Results are sorted by the space each group would free, and files that are already hardlinks of one another count once. More in the scanning chapter.

Resumable scans and the hash cache

Walking and hashing save progress to dedcom.db as they go, hashes after every 64 files. After Esc, a reboot or a power cut, the next start offers to resume where the scan stopped. Only grouping starts over, from the saved hashes. A resume needs the same set of roots and keeps the settings stored with the scan (resume).

On a repeat scan, a file whose device, inode, size and mtime are unchanged takes its hash from the cache instead of being read again: tens of seconds per million files. If some program changes file content without updating mtime, --no-hash-reuse re-hashes everything (hash cache).

The database lives in ~/.local/state/dedcom/, and --state-dir moves it. One scan of 2.2 million files takes roughly 200–400 MiB.

Two interfaces and an observer mode

Twin folders

While grouping, DedupCommando builds a signature of each directory's entire contents. Directories with the same signature have identical content, recursively, and appear as twin folders in the Commando's directory-group views, with the full path of each copy (twin folder views).

Headless mode for scripts and cron

CommandWhat it does
dedcom --scan /tankScan one or more roots and print the first 50 groups
dedcom --statsShow statistics for all scans and for the database
dedcom --export-csv dups.csvExport the newest scan to CSV; refuses if it has not finished
dedcom --compact-dbEmpty the session trash and compact the database
dedcom --purge-quarantineReport the size of the quarantine; with --yes, delete it

Exit codes: 0 success, 1 runtime error, 2 bad arguments. Headless mode never prompts: if an interactive session holds the lock, a writing command exits with an error. It never applies actions either: that needs a person to confirm keepers and marks in the interface (headless mode).

A nightly scan from cron:

# /etc/cron.d/dedcom — every night at 02:00
0 2 * * * root flock -n /var/lock/dedcom.scan /usr/local/bin/dedcom --scan /tank >> /var/log/dedcom-scan.log 2>&1

Headless --scan uses the intensity profile of the last scan configuration, Balanced by default. Set Idle (one thread, nice 19, ionice idle) once in the interface and start a scan with it, so nightly scans do not slow down VMs or backups (cron example).

Delete duplicate files safely

Finding duplicates only reads. Files change only when you mark them and confirm a plan. In each group you mark one keeper, and each copy gets an action: delete to quarantine, hardlink or reflink.

"Delete" does not unlink: the file moves to .dedcom-quarantine/<timestamp>/ at the root of its dataset and can be moved back until you purge it. Before the first change, dedcom snapshots every dataset in the batch and aborts the batch if any snapshot fails. Right before each action it checks that the target and the keeper still match the scan, and a mismatch cancels that action.

On ext4, XFS and Btrfs

On ext4, XFS or Btrfs the walk and hash run as usual, so you can find duplicates and export a report there. Applying changes outside ZFS is not recommended, because there is no snapshot to roll back to. Delete can also cancel with target file's dataset could not be determined when it cannot tell which ZFS dataset a file is on. A script of your own built from the CSV skips DedupCommando's checks and snapshot.

Install

DedupCommando runs on Linux (x86_64 or aarch64, kernel 3.15 or newer). The pre-built packages need glibc 2.39 or newer: Debian 13, Ubuntu 24.04, Proxmox VE 9 or later. ZFS with zfs in PATH is strongly recommended, and dedcom typically runs as root. On Debian-family systems, install from the signed APT repository:

# as root (Proxmox default); on non-root Debian run: sudo -i
curl -fsSL https://dedupcommando.github.io/apt/dedcom-archive-keyring.gpg \
  -o /usr/share/keyrings/dedcom-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/dedcom-archive-keyring.gpg] https://dedupcommando.github.io/apt stable main" \
  | tee /etc/apt/sources.list.d/dedcom.list
apt update && apt install dedcom

On other distributions, use the release tarball and verify it first.

Your first scan in five minutes

These steps follow the quickstart chapter. The five minutes are your part; the scan itself takes as long as your disks need.

  1. Run dedcom. On the first start, tick the notice with Space and press Enter.
  2. Press F9 and choose "Configure and start a scan…", or press Shift+F9.
  3. Mark the roots with Space. If the host runs VMs or backups, press G until the profile reads Idle. Press S to start.
  4. Wait for the three phases. Esc stops after the current chunk, and the next start offers to resume. On two hard disks, hashing in Idle runs at about 50–100 MiB/s.
  5. In the active panel, press v until it shows "groups", largest savings first. Press Tab, then v until the next panel shows "group files".
  6. Put the cursor on the file to keep, press o to open it in a files panel, and mark it as the keeper with F7. Open each copy the same way and mark it F5 (hardlink), F6 (reflink) or F8 (delete).
  7. Press F11 (or x). Check the "By type" line and the listed paths; S saves the plan as a shell script for your records. Press Y to apply: Y re-checks each file's content right before acting, while a saved script checks only inode, size and times.
  8. The summary names the snapshot and the quarantine. After a week or two of normal use, reclaim the space:
zfs list -t snapshot | grep dedcom-      # the batch snapshots
zfs destroy tank@dedcom-<timestamp>      # one per dataset
dedcom --purge-quarantine                # shows what would be deleted
dedcom --purge-quarantine --yes          # deletes it; this cannot be undone

If something looks wrong before then, move single files back out of the quarantine, or zfs rollback to the batch snapshot, which also discards anything else written to that dataset since.

Next steps