Duplicate files on Proxmox VE: file-level deduplication for ZFS storage

Beta — v0.9.2 . DedupCommando changes real files. Try the first batch on non-critical data, keep backups, and read the safety model before you apply anything.

Proxmox VE hosts usually keep their data on ZFS, and over the years the same ISO images, container templates, media and shared files pile up in more than one place. DedupCommando finds those byte-for-byte duplicates on the host's datasets and reclaims the space by hardlink, block-cloned reflink or delete to quarantine — each batch under a ZFS snapshot.

It is an independent tool, not a Proxmox plugin: it does not modify Proxmox VE and works only on the files in your datasets. It is tested on Proxmox VE 9.1 with OpenZFS 2.3.

Not the same as backup deduplication

Proxmox Backup Server deduplicates backups: it splits them into chunks identified by their SHA-256 checksum and reuses those chunks across the backup snapshots in a datastore (PBS technical overview). That happens inside the backup server.

This page is about something else: duplicate files on the host's own ZFS storage — ISO and template stores, media libraries, and datasets you share with VMs and containers. It is not ZFS's block-level dedup property either; file-level deduplication on ZFS explains the difference.

Install on Proxmox VE 9

Install dedcom from the project's signed APT repository:

# as root (Proxmox default); on non-root Debian run: sudo -i
curl -fsSL https://dedupcommando.github.io/apt/dedcom-archive-keyring.gpg \
  -o /usr/share/keyrings/dedcom-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/dedcom-archive-keyring.gpg] https://dedupcommando.github.io/apt stable main" \
  | tee /etc/apt/sources.list.d/dedcom.list
apt update && apt install dedcom

After that, apt upgrade keeps it current. The packages need glibc 2.39 or newer, which means Proxmox VE 9 or later: on Proxmox VE 8 (Debian 12) the install stops on an unmet libc6 dependency, and building from source is the way there. The installation chapter also covers the release tarball and how to verify it.

Run it as root, over SSH or in the web shell

Run dedcom as root on the host, the Proxmox default. As an unprivileged user it cannot take ZFS snapshots without zfs allow or sudo, and without snapshots applying actions is not recommended. It is a terminal program that needs a UTF-8, 256-color terminal, so an SSH session works as well as the Proxmox web shell.

The web shell (xterm.js) does not pass Shift with F-keys, so press the backtick first: ` then F9 does what Shift+F9 does. If F11 toggles fullscreen instead of reaching dedcom, press x to execute the marked actions (troubleshooting).

Only one instance writes at a time; a second window can watch with dedcom --read-only.

Scan gently: the Idle profile

A scan reads every candidate file, and on a pool with running VMs or backups the manual makes the Idle profile mandatory. It caps the scan at one read thread, at the lowest CPU and I/O priority — the manual's way to keep a scan from starving your guests:

ProfileRead threadsPriority
Turboall cores (nproc)default
Balanced (default)2default
Idle1nice 19 + ionice idle

Press G in the scan configuration to cycle Turbo → Balanced → Idle. The profile is saved with the scan, so a resumed scan keeps it, and headless scans reuse the last one you set. If guests still slow down, press Esc: the scan stops after the current chunk, keeps its progress and can resume later (intensity profiles).

What to leave out of a scan

Leave other programs' stores out of the scan roots: a Proxmox Backup Server datastore, a virtual machine's disk, a restic, borg or kopia repository. Such a store needs every file at its own path with its own content. A delete inside it breaks it at once, and a hardlink makes two copies one file, so a program that later writes in place, as VM disks and databases do, changes both. A scan root takes in every dataset mounted below it and there is no way to exclude a path, so the roots are the only fence. The manual covers this in section 8.9.

What else the manual says points the same way:

Taken together: point dedcom at folders whose files sit still, such as ISO and template stores, media and file shares, and leave out anything guests or backup jobs keep writing.

Walkthrough: scan, review, apply

These are the keys of the default commander interface, as in the quickstart:

  1. Start. Run dedcom. On the first run, tick the notice with Space and press Enter.
  2. Open the scan wizard. F9 → "Configure and start a scan…" → Enter.
  3. Choose roots and the profile. Space on each dataset to scan, G until the intensity reads Idle, S to start.
  4. Wait. The scan walks, hashes and groups. Esc stops it cleanly, and the next start offers to resume.
  5. Open the groups. Press v twice in a panel for "groups", largest savings first. Tab to the next panel and press v until it shows "group files".
  6. Mark. In "group files", put the cursor on the file to keep and press o: its folder opens in a third panel, cursor on the file (three panels need a window at least 108 columns wide). Press F7 to make it the keeper. Go back with ←, move to a copy, press o, then → and mark it: F5 hardlink, F6 reflink or F8 delete to quarantine. Repeat for each copy (step 8).
  7. Review. F11 or x opens the confirmation. Check the "By type" line and the listed paths; Tab shows the full shell script, and S saves it as a .sh file. Enter does nothing here.
  8. Apply. Y takes the snapshots, then revalidates and applies each action. The Summary lists the rollback command, the quarantine folder and the commands that free the space.

Hardlink and reflink work only within one dataset; for copies in different datasets, delete to quarantine is the option. dedcom --classic runs the same steps as a one-screen-at-a-time wizard — Enter sets the keeper, h, c and d mark, r reviews — which also suits terminals without F-keys (classic browser).

Snapshots and quarantine on a live host

Before the first action, dedcom snapshots every dataset in the batch (<dataset>@dedcom-<timestamp>); if any snapshot fails, nothing runs. Replaced and deleted files go to .dedcom-quarantine/<timestamp>/ at the root of their dataset, with owner, permissions and xattrs intact.

zfs list -t snapshot | grep dedcom-    # what the batches left behind
zfs destroy tank@dedcom-<timestamp>    # irreversible: one per dataset
dedcom --purge-quarantine              # reports what it would delete
dedcom --purge-quarantine --yes        # irreversible: empties every quarantine

Scans from cron

dedcom --scan <path> runs without the interface, so it fits cron. It only scans; applying is interactive, with Y in the confirmation. The manual has a ready cron line with flock (headless); point it at the path command -v dedcom prints, and set Idle once in the interface first.

Next steps