07. Scanning

A scan runs as three sequential phases. When it finishes, the result is stored in dedcom.db and becomes available for review in the Browser / Commando.

walk (1/3) ──► hash (2/3) ──► group (3/3)
tree walk      BLAKE3          group assembly
               + cache         + directory
                               signatures
PhaseWhat it doesTime on 2 M filesMemory
1/3 WalkingWalk the tree; lstat each entry; write to the DBminutestens of MB
2/3 HashingBLAKE3 of each candidate; read by (dev,inode)hours (on HDD)tens of MB
3/3 GroupingSQL aggregation into groups + directory signaturesseconds–minutes~2.5 KiB/file = GiB

Phases 1 and 2 stream and checkpoint progress in chunks — you can interrupt and resume. During hashing a stop takes effect within about a second, even in the middle of a large file — Esc on the scan screen, Ctrl+C in --scan, SIGTERM, or SIGHUP from a dropped SSH session (so run a long scan over SSH in tmux or screen: nohup does not keep it going); the files being read at that moment get no hash and are read again from the start on resume. Phase 3 is a single transient memory peak; if RAM is tight, see --merkle-dirs below.

Intensity profiles (Resource Governor)

In the scan configuration wizard the G key cycles the profile:

Turbo → Balanced → Idle → Turbo
ProfileHintRead threadsPriority
Turboall cores, disk at fullnprocdefault
Balanced ◀ default2 threads, no seek-thrashmin(2, nproc)default
Idle1 thread, nice+ionice idle1nice 19 + ionice idle

⚠️ Idle is mandatory on live data. A concurrent VM or backup will only see the dedup run when the disk is idle (the ionice idle class on Linux). Turbo on /tank with running VMs means they slow down for the whole duration of the scan.

The profile is persisted in the scan's checkpoint DB — resume uses the same profile. There is no command-line flag for it: a headless --scan starts every new scan on Balanced (§11).

Filters

By size

ScanConfig.min_size (default 4096 bytes) — smaller files are skipped. There is little point in deduplicating files smaller than one filesystem block (the gain is smaller than the hardlink/reflink overhead).

ScanConfig.max_size (no default) — if set, larger files are skipped. Not yet configurable in the UI; change it in the code/checkpoint.

By extension

The CLI flag --include-ext (repeatable, comma-separated):

--include-ext <LIST> — Scan only files with these extensions (comma-separated: jpg,png,gif; flag repeatable)

Case is ignored for the letters A–Z, and a leading * and dot are stripped: '*.JPG' means jpg (quote it, or the shell expands the *). An extension may have a dot inside: tar.gz matches photos.tar.gz, and so does gz. A value with a / is refused — no file name holds one.

dedcom --scan /tank/media --include-ext jpg,heic,raw
dedcom --scan /tank/media --include-ext jpg --include-ext heic
dedcom --scan /tank/backup --include-ext tar.gz,zip

In the TUI a preset is chosen with P in the scan configuration wizard. Presets are defined in the code — Images and Office documents — and in presets.json in the state directory, whose extensions lose a leading * and dot the same way.

An empty filter (the default) means all files are scanned.

Permanent exclusions

Always skipped:

  • **/.zfs/** — ZFS snapshots (reading them is suboptimal, and they are read-only).
  • **/.dedcom-quarantine/** — our own quarantine (we don't deduplicate ourselves).

There is no way to add exclusions of your own: a root takes in everything below it, including every dataset mounted there, so choose roots that do not reach what must stay untouched (§8.9).

Hash cache (hash_cache)

The DB holds a hash_cache table keyed by (device, inode, size, mtime) — on a repeat scan of the same file, if those four attributes are unchanged, BLAKE3 is not recomputed but taken from the cache.

EnabledRepeat-scan speedWhen to disable
ON (default)Tens of seconds per million files (all from cache)Never in normal use
OFFFull re-hashing of all filesSuspicion that external software changes content without updating mtime
How to disableWhere
Persistently (for one scan)TUI: C in the scan configuration wizard
For headlessdedcom --scan /tank --no-hash-reuse

The flag's help string:

--no-hash-reuse — Disable the hash cache — re-hash all files

The cache survives a reboot (the stat fields are stable). It does not survive: a rename by external software (mv keeps the inode, but ctime changes — irrelevant to us, we look only at mtime), a copy (a new inode = a cache miss), a filesystem change.

Directory-signature algorithm

After files are hashed, phase 3/3 builds directory signatures — a hash of a directory's entire contents. Two directories with the same signature are "twin folders" (see §05 Commando).

There are two algorithms:

Default — build_dir_groups

Top-down recursion: each directory keeps its signature together with the list of its child nodes in memory.

PropertyValue
Memory~2.5 KiB per hashed file (on /tank, gigabytes)
SpeedFaster on typical trees
Group hex valueStable, readable

On large pools (millions of files) this is a transient memory peak for the whole program. On a 2.12 M-file pool the peak was +4.35 GiB.

Merkle — --merkle-dirs (opt-in)

A streaming Merkle hash bottom-up: each directory is hashed once and frees its child nodes' memory immediately.

PropertyValue
MemoryO(tree depth) — tens to hundreds of MB
SpeedComparable to the default (the streaming overhead is minimal)
Group hex valueDifferent (a Merkle hash), but the group membership is identical to the default

The flag's help string:

--merkle-dirs — (opt-in) streaming-Merkle directory signature: O(depth) memory instead of ~2.5 KiB/file. Group membership is identical to the default; per-row hex differs. Persisted in the checkpoint — resume uses the same algorithm.

dedcom --merkle-dirs                    # TUI with Merkle for the next scan
dedcom --scan /tank --merkle-dirs       # headless

When you need --merkle-dirs: on a host where ~2.5 KiB × file_count is close to free RAM or exceeds it. Before phase 3/3 dedcom prints a forecast and compares it with free RAM — if you get a red warning, or a previous scan was killed by OOM on 3/3, turn it on. A scan killed that way is still unfinished, and a resume keeps the algorithm it was started with, so start a new one: dedcom --scan /tank --merkle-dirs --no-resume (without --no-resume the flag is refused, see §11).

The algorithm is persisted in the checkpoint — resume uses the same one.

Estimating the 3/3 memory peak (for the default algorithm)

phase_3/3_peak ≈ files × 2.5 KiB

Files hashedPhase-3/3 peak (default)Decision
250,000~0.6 GiBfine
1,000,000~2.4 GiBcheck free RAM
2,000,000~4.8 GiBcompare with RAM; --merkle-dirs likely needed
5,000,000~12 GiB--merkle-dirs required

After phase 3/3 the memory is returned to the OS — it is a transient peak. After showing the result the UI holds hundreds of MB.

Byte-by-byte comparison (--verify)

After hashing, each duplicate group can additionally be compared byte by byte (a guard against a theoretical BLAKE3 collision):

The flag's help string:

--verify — Byte-by-byte comparison after hashing

dedcom --scan /tank --verify

It doubles the scan time (each file is read twice: hash + comparison). In practice a BLAKE3 collision has never been observed on real data and this check is redundant — but if you are paranoid, the flag enables it. A stop that comes during the comparison is not seen: the comparison runs to the end and the scan completes.

Re-validation before actions (--strict-verify)

This is a different check — not during the scan, but before every apply action. By default the mode is Hybrid: the keeper is hashed once per batch, the remaining checks go through a re-stat on FileIdentity. With the flag the mode is Strict: full re-hashing of target and keeper before every destructive operation.

The flag's help string:

--strict-verify — Re-validate before an action: re-hash target and keeper every time (default Hybrid — keeper once per batch)

dedcom --strict-verify      # for all applies in this session

Details — §08 Actions.

Resume — continue an unfinished scan

A scan is saved into dedcom.db in chunks:

  • Walking — after every batch of WALK_BATCH files.
  • Hashing — after every chunk of 64 files (HASH_CHUNK).
  • Grouping — NOT saved (the phase is atomic; a crash = re-running phase 3 from scratch, but the hashes are intact).

ScanStatus in the DB:

StatusMeaningResume?
WalkingInterrupted in phase 1✅
HashingInterrupted in phase 2✅
CompleteScan finished (including phase 3)—
AbortedInterrupted after phase 3 or an invariant was broken❌ — new scan

On the next start of the wizard for the same roots, a Resume overlay appears (see §05 Commando) offering "resume / open the last completed / start a new one".

What does NOT survive a resume:

  • A change of the root set (resume_probe_for_roots compares exactly).
  • A change of min_size or the extension filter (new files could enter the scan).
  • Deleting dedcom.db or the state directory.
  • All settings that affect the config (--merkle-dirs, --no-hash-reuse, --include-ext, the profile): they apply ONLY at the start of a new scan; on resume the values are read from the checkpoint. The interface ignores such flags on a resume; a headless --scan that asks for other values is refused (add --no-resume to start a new scan with them).

Storage-type override (--storage-type)

DedupCommando auto-detects a dataset's storage type (HDD / SSD / NVMe) for:

  • the read order in phase 2 (for HDD: sorting candidates by (device, inode) — −21% off the cold-scan time on 2×HDD, by reducing seeks; not needed for SSD/NVMe);
  • statistics (--stats shows the type on the scan line).

If auto-detection is wrong (for example, the host is in a VM and the disks are reported as SSD but are in reality HDD-backed) — override it:

The flag's help string:

--storage-type <TYPE> — Storage type for statistics: hdd | ssd | nvme (overrides auto-detection)

dedcom --scan /tank --storage-type hdd       # force the HDD strategy
dedcom --scan /tank --storage-type ssd
dedcom --scan /tank --storage-type nvme

Not configurable in the TUI — CLI only.

Messages and warnings during a scan

ScanProgress::Notice(String) — the Scanning screen and dedcom.log print additional messages, the most important being:

  • "Phase 3/3 peak estimate: X GiB, free: Y GiB" — the files × 2.5 KiB calculation from the phase-2 results vs available_ram_bytes. If X > Y it is printed in yellow: a risk of OOM, and you are advised to interrupt (Esc) and start a new scan with --merkle-dirs — "new" rather than "resume" in the interface, --no-resume with --scan: a resume keeps the algorithm it was started with.

Summary — typical flag combinations

ScenarioCommand
Scan /tank, gently (live VMs)TUI: F9 → scan configuration wizard → Space tank → G to Idle → S
Scan /tank, headless from cronnice -n 19 ionice -c 3 dedcom --scan /tank (a new headless scan runs on Balanced — the full cron line is in §11)
Large pool, little RAM+ --merkle-dirs (with --scan, or at TUI launch)
Doubts about hash integrity+ --verify (with --scan, or at TUI launch; 2× slower)
Suspicion that external software changes content--scan + --no-hash-reuse (TUI: C in the wizard)
Media only--scan + --include-ext jpg,heic,mp4,mov (TUI: P in the wizard)
Paranoid applydedcom --strict-verify (at TUI launch)
VM with wrong storage auto-detection--scan + --storage-type hdd

A flag given to a run that does not read it is refused with exit code 2 — for example --include-ext without --scan (§11).

What's next

Published from docs/manual/07-scanning.md at v0.9.2 · last changed 2026-09-28