07. Scanning
A scan runs as three sequential phases. When it finishes, the result is stored
in dedcom.db and becomes available for review in the Browser / Commando.
walk (1/3) ──► hash (2/3) ──► group (3/3)
tree walk BLAKE3 group assembly
+ cache + directory
signatures
| Phase | What it does | Time on 2 M files | Memory |
|---|---|---|---|
| 1/3 Walking | Walk the tree; lstat each entry; write to the DB | minutes | tens of MB |
| 2/3 Hashing | BLAKE3 of each candidate; read by (dev,inode) | hours (on HDD) | tens of MB |
| 3/3 Grouping | SQL aggregation into groups + directory signatures | seconds–minutes | ~2.5 KiB/file = GiB |
Phases 1 and 2 stream and checkpoint progress in chunks — you can interrupt and
resume. During hashing a stop takes effect within about a second, even in the
middle of a large file — Esc on the scan screen, Ctrl+C in --scan, SIGTERM,
or SIGHUP from a dropped SSH session (so run a long scan over SSH in tmux or
screen: nohup does not keep it going); the files being read at that moment
get no hash and are read again from the start on resume. Phase 3 is a single
transient memory peak; if RAM is tight, see --merkle-dirs below.
Intensity profiles (Resource Governor)
In the scan configuration wizard the G key cycles the profile:
Turbo → Balanced → Idle → Turbo
| Profile | Hint | Read threads | Priority |
|---|---|---|---|
| Turbo | all cores, disk at full | nproc | default |
| Balanced ◀ default | 2 threads, no seek-thrash | min(2, nproc) | default |
| Idle | 1 thread, nice+ionice idle | 1 | nice 19 + ionice idle |
⚠️ Idle is mandatory on live data. A concurrent VM or backup will only see the dedup run when the disk is idle (the
ionice idleclass on Linux). Turbo on/tankwith running VMs means they slow down for the whole duration of the scan.
The profile is persisted in the scan's checkpoint DB — resume uses the same
profile. There is no command-line flag for it: a headless --scan starts every
new scan on Balanced (§11).
Filters
By size
ScanConfig.min_size (default 4096 bytes) — smaller files are skipped.
There is little point in deduplicating files smaller than one filesystem block
(the gain is smaller than the hardlink/reflink overhead).
ScanConfig.max_size (no default) — if set, larger files are skipped. Not yet
configurable in the UI; change it in the code/checkpoint.
By extension
The CLI flag --include-ext (repeatable, comma-separated):
--include-ext <LIST>— Scan only files with these extensions (comma-separated: jpg,png,gif; flag repeatable)
Case is ignored for the letters A–Z, and a leading * and dot are stripped:
'*.JPG' means jpg (quote it, or the shell expands the *). An extension may have
a dot inside: tar.gz matches photos.tar.gz, and so does gz. A value with a /
is refused — no file name holds one.
dedcom --scan /tank/media --include-ext jpg,heic,raw
dedcom --scan /tank/media --include-ext jpg --include-ext heic
dedcom --scan /tank/backup --include-ext tar.gz,zip
In the TUI a preset is chosen with P in the scan configuration wizard.
Presets are defined in the code — Images and Office documents — and in presets.json in
the state directory, whose extensions lose a leading * and dot the same way.
An empty filter (the default) means all files are scanned.
Permanent exclusions
Always skipped:
**/.zfs/**— ZFS snapshots (reading them is suboptimal, and they are read-only).**/.dedcom-quarantine/**— our own quarantine (we don't deduplicate ourselves).
There is no way to add exclusions of your own: a root takes in everything below it, including every dataset mounted there, so choose roots that do not reach what must stay untouched (§8.9).
Hash cache (hash_cache)
The DB holds a hash_cache table keyed by (device, inode, size, mtime) — on a
repeat scan of the same file, if those four attributes are unchanged, BLAKE3 is
not recomputed but taken from the cache.
| Enabled | Repeat-scan speed | When to disable |
|---|---|---|
| ON (default) | Tens of seconds per million files (all from cache) | Never in normal use |
| OFF | Full re-hashing of all files | Suspicion that external software changes content without updating mtime |
| How to disable | Where |
|---|---|
| Persistently (for one scan) | TUI: C in the scan configuration wizard |
| For headless | dedcom --scan /tank --no-hash-reuse |
The flag's help string:
--no-hash-reuse— Disable the hash cache — re-hash all files
The cache survives a reboot (the
statfields are stable). It does not survive: a rename by external software (mvkeeps the inode, butctimechanges — irrelevant to us, we look only atmtime), a copy (a new inode = a cache miss), a filesystem change.
Directory-signature algorithm
After files are hashed, phase 3/3 builds directory signatures — a hash of a directory's entire contents. Two directories with the same signature are "twin folders" (see §05 Commando).
There are two algorithms:
Default — build_dir_groups
Top-down recursion: each directory keeps its signature together with the list of its child nodes in memory.
| Property | Value |
|---|---|
| Memory | ~2.5 KiB per hashed file (on /tank, gigabytes) |
| Speed | Faster on typical trees |
| Group hex value | Stable, readable |
On large pools (millions of files) this is a transient memory peak for the whole program. On a 2.12 M-file pool the peak was +4.35 GiB.
Merkle — --merkle-dirs (opt-in)
A streaming Merkle hash bottom-up: each directory is hashed once and frees its child nodes' memory immediately.
| Property | Value |
|---|---|
| Memory | O(tree depth) — tens to hundreds of MB |
| Speed | Comparable to the default (the streaming overhead is minimal) |
| Group hex value | Different (a Merkle hash), but the group membership is identical to the default |
The flag's help string:
--merkle-dirs— (opt-in) streaming-Merkle directory signature: O(depth) memory instead of ~2.5 KiB/file. Group membership is identical to the default; per-row hex differs. Persisted in the checkpoint — resume uses the same algorithm.
dedcom --merkle-dirs # TUI with Merkle for the next scan
dedcom --scan /tank --merkle-dirs # headless
When you need
--merkle-dirs: on a host where~2.5 KiB × file_countis close to free RAM or exceeds it. Before phase 3/3 dedcom prints a forecast and compares it with free RAM — if you get a red warning, or a previous scan was killed by OOM on 3/3, turn it on. A scan killed that way is still unfinished, and a resume keeps the algorithm it was started with, so start a new one:dedcom --scan /tank --merkle-dirs --no-resume(without--no-resumethe flag is refused, see §11).
The algorithm is persisted in the checkpoint — resume uses the same one.
Estimating the 3/3 memory peak (for the default algorithm)
phase_3/3_peak ≈ files × 2.5 KiB
| Files hashed | Phase-3/3 peak (default) | Decision |
|---|---|---|
| 250,000 | ~0.6 GiB | fine |
| 1,000,000 | ~2.4 GiB | check free RAM |
| 2,000,000 | ~4.8 GiB | compare with RAM; --merkle-dirs likely needed |
| 5,000,000 | ~12 GiB | --merkle-dirs required |
After phase 3/3 the memory is returned to the OS — it is a transient peak. After showing the result the UI holds hundreds of MB.
Byte-by-byte comparison (--verify)
After hashing, each duplicate group can additionally be compared byte by byte (a guard against a theoretical BLAKE3 collision):
The flag's help string:
--verify— Byte-by-byte comparison after hashing
dedcom --scan /tank --verify
It doubles the scan time (each file is read twice: hash + comparison). In practice a BLAKE3 collision has never been observed on real data and this check is redundant — but if you are paranoid, the flag enables it. A stop that comes during the comparison is not seen: the comparison runs to the end and the scan completes.
Re-validation before actions (--strict-verify)
This is a different check — not during the scan, but before every apply
action. By default the mode is Hybrid: the keeper is hashed once per
batch, the remaining checks go through a re-stat on FileIdentity. With the flag
the mode is Strict: full re-hashing of target and keeper before every
destructive operation.
The flag's help string:
--strict-verify— Re-validate before an action: re-hash target and keeper every time (default Hybrid — keeper once per batch)
dedcom --strict-verify # for all applies in this session
Details — §08 Actions.
Resume — continue an unfinished scan
A scan is saved into dedcom.db in chunks:
- Walking — after every batch of
WALK_BATCHfiles. - Hashing — after every chunk of 64 files (
HASH_CHUNK). - Grouping — NOT saved (the phase is atomic; a crash = re-running phase 3 from scratch, but the hashes are intact).
ScanStatus in the DB:
| Status | Meaning | Resume? |
|---|---|---|
| Walking | Interrupted in phase 1 | ✅ |
| Hashing | Interrupted in phase 2 | ✅ |
| Complete | Scan finished (including phase 3) | — |
| Aborted | Interrupted after phase 3 or an invariant was broken | ❌ — new scan |
On the next start of the wizard for the same roots, a Resume overlay appears (see §05 Commando) offering "resume / open the last completed / start a new one".
What does NOT survive a resume:
- A change of the root set (
resume_probe_for_rootscompares exactly).- A change of
min_sizeor the extension filter (new files could enter the scan).- Deleting
dedcom.dbor the state directory.- All settings that affect the config (
--merkle-dirs,--no-hash-reuse,--include-ext, the profile): they apply ONLY at the start of a new scan; on resume the values are read from the checkpoint. The interface ignores such flags on a resume; a headless--scanthat asks for other values is refused (add--no-resumeto start a new scan with them).
Storage-type override (--storage-type)
DedupCommando auto-detects a dataset's storage type (HDD / SSD / NVMe) for:
- the read order in phase 2 (for HDD: sorting candidates by
(device, inode)— −21% off the cold-scan time on 2×HDD, by reducing seeks; not needed for SSD/NVMe); - statistics (
--statsshows the type on the scan line).
If auto-detection is wrong (for example, the host is in a VM and the disks are reported as SSD but are in reality HDD-backed) — override it:
The flag's help string:
--storage-type <TYPE>— Storage type for statistics: hdd | ssd | nvme (overrides auto-detection)
dedcom --scan /tank --storage-type hdd # force the HDD strategy
dedcom --scan /tank --storage-type ssd
dedcom --scan /tank --storage-type nvme
Not configurable in the TUI — CLI only.
Messages and warnings during a scan
ScanProgress::Notice(String) — the Scanning screen and dedcom.log print
additional messages, the most important being:
- "Phase 3/3 peak estimate: X GiB, free: Y GiB" — the
files × 2.5 KiBcalculation from the phase-2 results vsavailable_ram_bytes. If X > Y it is printed in yellow: a risk of OOM, and you are advised to interrupt (Esc) and start a new scan with--merkle-dirs— "new" rather than "resume" in the interface,--no-resumewith--scan: a resume keeps the algorithm it was started with.
Summary — typical flag combinations
| Scenario | Command |
|---|---|
Scan /tank, gently (live VMs) | TUI: F9 → scan configuration wizard → Space tank → G to Idle → S |
Scan /tank, headless from cron | nice -n 19 ionice -c 3 dedcom --scan /tank (a new headless scan runs on Balanced — the full cron line is in §11) |
| Large pool, little RAM | + --merkle-dirs (with --scan, or at TUI launch) |
| Doubts about hash integrity | + --verify (with --scan, or at TUI launch; 2× slower) |
| Suspicion that external software changes content | --scan + --no-hash-reuse (TUI: C in the wizard) |
| Media only | --scan + --include-ext jpg,heic,mp4,mov (TUI: P in the wizard) |
| Paranoid apply | dedcom --strict-verify (at TUI launch) |
| VM with wrong storage auto-detection | --scan + --storage-type hdd |
A flag given to a run that does not read it is refused with exit code 2 — for
example --include-ext without --scan (§11).
What's next
- §08 Actions — what happens after
apply. - §11 Headless —
--scanin cron, exit codes, output format. - §13 Troubleshooting — what to do on an OOM in phase 3/3, and so on.
Published from docs/manual/07-scanning.md at v0.9.2 · last changed 2026-09-28