Four small, independent changes surfaced while running a manual
backup/restore demo end to end:
- fdbrestore now accepts -C and --cluster-file as aliases for
--dest-cluster-file. Every other tool (fdbcli, fdbbackup, fdbserver,
backup_agent) already pairs -C with --cluster-file; fdbrestore was
the lone outlier and forced scripts to special-case it.
- fdbcli now accepts --logdir as an alias for --log-dir. Every other
tool spells the flag --logdir (no hyphen); the --log-dir form in
fdbcli was the outlier here.
- fdbbackup status no longer prints "BulkLoad Compatible: no" when
Snapshot Mode is rangefile. In that mode the answer is tautologically
derivable from the line above it (no bulkdump task ever runs), so
the line carries no information. The field still prints in modes
bulkdump and both, where it can flip during an in-progress backup.
- Correct the misleading comment in FileBackupAgent.cpp that claimed
submitBulkLoadJob validates cluster preconditions server-side. It
does not. A bulkload restore against a cluster missing
shard_encode_location_metadata / enable_read_lock_on_range, or with
a storage engine that cannot ingest SSTs, succeeds at submission and
then silently stalls in "State: running, Tasks: 0/0" forever. Replace
the comment with a faithful description of the failure mode and add
a matching TODO in submitBulkLoadJob explaining where the real
validation needs to live (somewhere with cluster-side knob
visibility) and why it cannot be done from fdbclient.
* Improve backup/restore/bulkload observability and reliability
Restore progress tracking:
- Add sub-phase counts (Submitted/Triggered/Running/TotalTasks) to
fdbrestore status output so users can see task submission progress
- Add getBulkLoadTaskProgress() to scan task states during restore
- Replace monitorBulkLoadJobCompletion with progress-tracking variant
that updates RestoreConfig counters every 5 seconds
Backup mode=BOTH fixes:
- Set bulkDumpJobId on BackupConfig so status shows BulkLoad Compatible
- Add bulkDumpSnapshotEndVersion for proper getLatestRestorableVersion
- Fix firstSnapshotEndVersion: only set from rangefile in mode=BOTH
- Skip empty BulkDump snapshots (totalSize=0) in rangefile restore
BulkLoad restore reliability:
- Abort restore immediately when no bulkdump data found (not retry forever)
- Add monitorBulkLoadModeAndSpawnActors so DD picks up bulkload jobs
submitted after DD initialization (required for restore workflow)
Audit validate_restore fixes:
- Read source data via database transaction instead of local SS to
avoid missing keys at shard boundaries after bulkload restore
- Add fast-path detection for completely empty baseline or source
- Retry on server_overloaded errors during audit
- Add AUDIT_RESTORE_BATCH_KEY_LIMIT and AUDIT_PROGRESS_PERSIST_BYTES_INTERVAL
knobs for tuning audit performance
CLI cleanup:
- Combine redundant task lines in bulkload/bulkdump status output
- Remove misleading health score and optimization recommendations
- Raise bulk task stall threshold from 60s to 600s (SST downloads
from blobstore routinely take 5-10 minutes)
* Address PR review: incrementalBackup check for mode=BOTH, snapshot type label, retry logging
- Add incrementalBackup fallback in getLatestRestorableVersion for mode==2
(was missing unlike modes 0 and 1)
- Note snapshot type heuristic as imperfect in describe output
- Log retries in getBulkLoadTaskProgress
* Parse snapshot type from filename instead of inferring from heuristic
KeyspaceSnapshotFile now has a snapshotType field parsed from the
snapshot filename (format: snapshot,begin,end,size,type). The type
suffix was already written but discarded during parsing. Now captured
and used for reliable bulkdump vs rangefile identification in
fdbbackup describe output.
* Simplify getBulkLoadTaskProgress: always set LOCK_AWARE, drop lockAware parameter
* Use Transaction instead of ReadYourWritesTransaction in getBulkLoadTaskProgress
Read-only function does not need write tracking overhead.
Uses Transaction* overload of krmGetRanges.