foundationdb/fdbserver/MemoryTrackerTest.cpp

Ignoring revisions in .git-blame-ignore-revs. Click here to bypass and see the normal blame view.

770 lines
25 KiB
C++
Raw Normal View History

call-site aware memory tracking (#13344) * Initial memory tracking design doc draft, plus first round of review comments by me with // TODO annotations * design/memory-tracker: address first-round review TODOs Resolves all // TODO annotations from the initial draft: - live-block table now optional via MEMORY_TRACKING_LIVE_TRACKING knob - drop the mmap slab pool; std::malloc + in-tracker flag is sufficient - add MEMORY_TRACKING_FORCE_SAMPLE_BYTES so large allocations are always captured regardless of the count-rate sampler - add live block / byte totals to MemoryTrackerSummary - replace the manual coverage spot-check with a sentinel-function unit test that introspects the aggregation table directly - leave ALLOC_INSTRUMENTATION alone; new hooks sit next to (not replacing) existing conditional ones - drop SJLJ jargon, trim A4 alternative now that force-sample-large collapses the byte-rate-vs-count-rate question * flow: add sampled per-call-site memory tracker Adds a sampled memory attribution layer (flow/MemoryTracker.{h,cpp}) hooked into the three primary allocation paths — global operator new/delete, FastAllocator, and ArenaBlock::create — plus a periodic TraceEvent dump driven from SystemMonitor. Knobs gate sample rate, force-sample threshold, report cadence, top-N, and capture depth; prod default is off. See design/memory-tracker.md. Test fixes uncovered while bringing the unit tests up: - memTrackerResetForTest now resets the per-thread sample counter and force-sample threshold. Without this, a test that exercised the off-switch path left gMemTrackerCounter at INT_MAX, which then silently suppressed sampling for the remainder of the run. - Slow-path reseed special-cases inverse==1 to keep the counter at 1. The general formula 1 + r % (2*inverse) yields counter values 1 or 2 at inverse==1, sampling only ~67% of allocations rather than every one, which broke a test that asserts exact alloc counts. - Sentinel functions in MemoryTrackerTest.cpp now route the allocated pointer through an asm-volatile escape() helper. Clang -O3 was eliding the new/delete pair (P0593 heap fusion), so the test's allocations never reached the operator-new override. * design/memory-tracker: address second-round review - threshold-based reporting (80 MB default, ~1% of 8 GB target RSS) replaces fixed top-N - prod report interval 60 s -> 10 min; sim stays 30 s - single combined MemoryTrackerAddrCmd event with one addr2line invocation per dump (positional mapping back to sites), keeping frame 0 — old design's per-site format_backtrace dropped the leaf alloc frame - new R12 "Side-thread coverage" + "Side-thread safety" subsection documenting the FP-elision crash mode found via joshua repros (RandomUnitTests / IThreadPool seeds segfaulted in captureStackFP when walking from FastAllocator<N>::~ThreadData into glibc's FP-elided pthread shutdown machinery) and the stack-bounds mitigation via pthread_getattr_np - R7 wording: live bytes (not cumulative) - CallSite struct in design overview aligned with implementation; ForceSampledCount promoted from prose-only to struct + emitted detail; exemplarFrames sized to MEMORY_TRACKER_MAX_FRAMES (=10) matching the FRAMES knob's stated 1-10 range - LIVE_TRACKING=false degraded-mode interaction documented - stale "slab pool" refs removed (std::malloc was already in code) and fdbserver.cpp self-contradiction resolved - R3 / Rollout reconciled (table shows steady-state, step 1 lands at 0) - 4-6 frame count flagged as initial estimate, subject to refinement - file:line citations stripped from path references (line numbers drift; symbol names are stable) * flow: threshold-based memory tracker reporting + side-thread safety Implements the second-round design changes in flow/MemoryTracker.{cpp,h}, flow/Knobs.{cpp,h}, flow/SystemMonitor.cpp. - MEMORY_TRACKING_TOP_N -> MEMORY_TRACKING_REPORT_BYTES_THRESHOLD (int64_t, default 80,000,000). MEMORY_TRACKING_REPORT_INTERVAL prod default 60.0 -> 600.0; sim still 30.0. - memTrackerDump(int topN) -> memTrackerDump(int64_t bytesThreshold). Filters by liveBytes (or cumulativeBytes when LIVE_TRACKING=false) >= threshold; emits MemoryTrackerSite per qualifying site plus one MemoryTrackerAddrCmd event with a single addr2line invocation covering every qualifying site's frames in dump order. The Summary event picks up SitesReported and ReportBytesThreshold details. - AddrCmd is built directly here (not via platform::format_backtrace, which deliberately drops index 0 for its single-site use case); the leaf alloc frame is preserved. - captureFramesFP gains a per-thread stack-bounds check via pthread_getattr_np + pthread_attr_getstack, cached in TLS. Without it, walking the FP chain from FastAllocator<N>::~ThreadData into glibc's FP-elided pthread shutdown machinery follows an uninitialized saved-FP slot and dereferences garbage. Fixes the joshua-found segfaults on RandomUnitTests seeds 3288611985, 3731245491, and 2219741568 (all in the IThreadPool worker-exit path). See design/memory-tracker.md "Side-thread safety". * flow/MemoryTracker: stub the FP walker on non-Linux pthread_getattr_np is glibc-specific and the macOS build broke on it. Frame-pointer walking on macOS is also unreliable on its own (system runtime has -fomit-frame-pointer in places we can't control), so a "loose bounds" workaround would still risk crashes. FDB is required to compile on macOS but is not run in production there. Gate initStackBoundsForThread + the real captureFramesFP on __linux__; provide a return-0 stub on non-Linux. The rest of the tracker (sample counters, aggregation, dump) still compiles and runs; per-call-site reports on macOS will just lack stack attribution. * flow: clang-format fixup for memory-tracker files Whitespace-only. Catches up flow/Arena.cpp and flow/MemoryTrackerTest.cpp with the project's clang-format style; the original implementation commit (d587f82b) slipped these past the format pre-flight. * edit for clarity, brevity, and uniform voice * flow/Arena: fix double-tracking on the >256/huge ArenaBlock paths ArenaBlock::create's >256 and huge paths go through allocateAndMaybeKeepalive (`new uint8_t[]`), which fires the global operator new[] hook in addition to the explicit memTrackerOnAlloc that fires immediately after. Two sites tracked the same pointer; on free only the explicit-Arena fingerprint was debited, so the operator new[] fingerprint accumulated liveBytes monotonically and LiveBytesTotal/LiveBlocksTotal skewed by +n/+1 per arena alloc/free pair. Reported as B1 in the PR review. Fix at the Arena layer (so non-arena allocateAndMaybeKeepalive callers in serialize.h's PacketBuffer code remain attributed at the operator-new layer): a MemTrackerSuppress RAII helper held across the underlying new[]/delete[] in ArenaBlock::create and ArenaBlock::destroyLeaf. Adds accounting tests for FastAllocator<32>, Arena small, Arena medium (the B1 path), and Arena huge. The load-bearing assertion is "exactly one site has the sentinel's frames AND nonzero bytes" -- fails pre-fix for medium/huge with sites=2. Tests gate on __linux__ since captureFramesFP is a no-op on macOS. * flow/memory-tracker: address PR review follow-ups - Knobs.cpp: sim default for MEMORY_TRACKING_REPORT_BYTES_THRESHOLD drops 80 MB -> 1 MB so sim dumps surface more sites for manual sanity-checking. Prod unchanged. - B2: memTrackerForEachSite holds MemTrackerSuppress across the callback loop so callbacks that allocate (e.g. fprintf failure dumps) don't re-enter tracking under SAMPLE_INVERSE=1. - B8: drop the always-zero g_reentrantBailouts and its SamplesDroppedReentry summary detail. Wiring it would require an atomic in the inline hot path, which R1 forbids. Design doc updated. - B3/B6/B7/B9: explanatory comments only -- frame-strip-count inlining assumption (B3), unstable sort acceptable under R6 (B6), shared xorshift seed acceptable in practice (B7), MEMORY_TRACKING_LIVE_TRACKING is startup-only (B9). * flow/MemoryTracker: one addr2line per site, drop chunking Move the addr2line command back onto MemoryTrackerSite as a per-site AddrCmd detail and remove the MemoryTrackerAddrCmd event entirely. Each AddrCmd carries exactly that site's stack -- short, well under the trace-detail truncation cap, ready to paste. Replaces the consolidated-then-chunked-into-byte-buckets approach, which split stacks across chunk boundaries and made raw events hard to read. * add standard Apache 2.0 license headers to new memory-tracker files The three new files (flow/MemoryTracker.{cpp,h} and flow/MemoryTrackerTest.cpp) shipped without the project-standard copyright/license block. Adds the standard 19-line header to each; file-purpose comments stay below it, switched to // line comments to keep the license block visually distinct. AGENTS.md gains a short "Source File Headers" section so the next contributor doesn't repeat the omission. * address review comments; reduce cost of enable check; maintain net estimates so users dont have to do it manually * formatting * unit test bug fix * gglass review comments on memory-tracker.md design doc * design doc updates, and fill in a plan for remaining tests/benchmarks * delete useless simulation section * flow/bench: add memory-tracker microbenchmarks; record measured overhead Add flow/bench/BenchMemoryTracker.cpp (Google Benchmark) measuring the tracker's per-op cost via three benchmarks -- raw malloc/free baseline, end-to-end operator new/delete, and isolated memTrackerOnAlloc/OnFree -- each at sample inverse 0 / 100 / 1. Auto-picked up by the flow_bench CONFIGURE_DEPENDS glob; run with: bin/flow_bench --benchmark_filter=memtracker Replace the placeholder "< 5% delta" targets in the design doc's Microbenchmarks section with the measured numbers and a projected per-second overhead at an assumed 100K alloc/sec. Headline: ~1.9 ns/pair disabled, ~5.6 ns/pair at the production 1% rate (~0.056% of a core at 100K/s, ~1/17th of R0's 1% ceiling); the every-allocation rows are labeled a buggified worst case, not a default. Testing: built flow_bench on the dev pod (clean, -Werror) and ran the memtracker filter six times; results stable to within a few percent across runs. * design/memory-tracker: describe the coverage test as implemented, not proposed The "Coverage spot-check via sentinel functions" section described the already-implemented `coverage` test in proposal tense ("Add an introspection API:", "The unit test:"), which read as future work. Reword to present tense referencing the actual `coverage`/`*Accounting` tests and the existing `memTrackerForEachSite` API. No remaining references to unwritten test cases. * remove extraneous detail from requirements section * substantially revise microbenchmark results based on seeing 2M allocations/frees on a CPU-maxed storage server * add script to drive A/B experiment for sampled memory allocation tracking * flow/MemoryTracker: move global operator new/delete into a server-only TU The global operator new/delete replacements that route allocations through the memory tracker lived in flow/MemoryTracker.cpp, i.e. in the flow static library. flow is linked into libfdb_c and every client binding, so the interposition shipped into client artifacts and could interpose the whole host process's allocator even with sampling off. Move them into a new fdbserver/GlobalNewDelete.cpp, compiled directly into the fdbserver executable (which clients never link), mirroring where the legacy ALLOC_INSTRUMENTATION overrides already lived. This also makes the replaceable-symbol interposition reliable (guaranteed in the final link) rather than dependent on static-archive pull-in. The legacy ALLOC_INSTRUMENTATION overrides move into the same file, selected by #if defined(ALLOC_INSTRUMENTATION) / #else, so exactly one set of global operators is ever defined. No CMake change is needed (fdbserver/*.cpp is globbed). Also reconcile the design doc with the code: override placement, the Files section (drop the unnecessary flow/CMakeLists.txt edit), the sim knob-table values (REPORT_BYTES_THRESHOLD, SAMPLE_INVERSE prod default), the FORCE_SAMPLE_BYTES "-1 disables" sentinel wording, and drop the stale "buggify inverse to 1" note -- the every-allocation and sampled/weighted paths are already pinned deterministically by MemoryTrackerTest.cpp. Testing: build green (run-ccmk5); fdbserver -r unittests -f /flow/MemoryTracker/ runs all 10 memory-tracker unit tests, 0 failed. * contrib/mako_ab_memtracker: disable RocksDB direct I/O for tmpfs runs RocksDB opens its DB with O_DIRECT by default; /mnt/ram (tmpfs), where the harness puts its data dir, does not support direct I/O, so the storage engine fails to Open and the cluster never configures (fdbcli "configure new" hangs, mako never starts). Pass the existing ROCKSDB_USE_DIRECT_READS / ROCKSDB_USE_DIRECT_IO_FLUSH_COMPACTION knobs (=0) for the rocksdb arm only -- no new knobs added. redwood is unaffected. * fdbserver/bench: move memory-tracker microbench to fdbserver_bench The global operator new/delete override lives in fdbserver/GlobalNewDelete.cpp (server-only, for client isolation), so flow_bench -- which links only flow -- could not exercise it: bench_memtracker_operator_new hit libc++'s operator new and its Arg(100)/Arg(1) rows were identical to Arg(0). Move BenchMemoryTracker.cpp to a new fdbserver/bench/ whose CMake compiles GlobalNewDelete.cpp into the fdbserver_bench executable (via ADDL_SRCS), so the real override is a strong definition in the bench link and the operator-new benchmark measures the actual hooked path. It links only flow, not the fdbserver dependency graph. Adds a BenchMain.cpp (BENCHMARK_MAIN equivalent) and wires the subdirectory into fdbserver/CMakeLists.txt. Testing: fdbserver_bench builds; bench_memtracker_operator_new now shows distinct off / 1% / every-alloc costs (11.4 / 14.4 / 80.9 ns), confirming the override fires. * design/memory-tracker: reconcile with code and slim down Reconcile the doc with the implementation and trim implementation detail that duplicated the code and had begun to drift: - Fix drift found in review: degraded-mode live/peak fields stay 0 (not "tracking the cumulatives"); the dump thresholds on the estimated fields (fix the memTrackerDump header comment too); the live-block table is allocated but empty when live-tracking is off (not "never allocated"); list the test file and its forceLinkMemoryTrackerTests() wiring. - Slim ~220 lines: replace the CallSite struct, the full captureStackFP source, both TraceEvent .detail() schemas, and the file-by-file inventory with prose that defers exact structs/keys/constants to the code; soften the knob table to intent (authoritative defaults live in Knobs.cpp). - Refresh the microbench numbers from fdbserver_bench and annotate the host (AMD EPYC 9R14, clang -O3); replace the A/B placeholder with the measured -16.5% (redwood) / -9.9% (rocksdb); note it is a point-in-time snapshot. - Frame R0 as the target the v1 single-lock design does not yet meet (ships off pending lock sharding); add a one-line note on table teardown/fork behavior. * fix clang-tidy error * mako wrapper scripts: take care to clobber the ramdisk before runs to avoid low-space throttling * design/memory-tracker: note the two mako A/B harnesses and the off-state result Point at contrib/mako_ab_memtracker.py (sampling off vs on) and contrib/mako_ab_binaries.py (vanilla main vs this PR built with tracking off), and record the latter's measured off-state overhead of -0.53% (redwood) / -0.10% (rocksdb) on a shared base commit — well within R0's 1% ceiling. * Add the default code review guidance to AGENTS.md * address big brother review comments; add some braces and trim some generated comments (hard to believe, but true) * clang-tidy again * Final read-through of this PR. -- Add or enhance a few comments on important items (performance & reliability) -- Delete some misc agent-written comments, typically exhibiting recency bias e.g. naming bugs identified in code review passes or describing mundane earlier bugs -- Add braces (InsertBraces style) * memory-tracker: address round-4 review Correctness/robustness: - operator new now runs the installed std::new_handler retry loop, so an allocation failure (including the tracker's own map growth) reaches FDB's platform::outOfMemory / FDB_EXIT_NO_MEM path instead of throwing past it. - Reentrancy guard restored via MemTrackerSuppress RAII on every path (hot path, dump, reset) so an exception can't permanently disable tracking on a thread. - Sampling reseed draws from [1, 2N-1] (mean exactly N); the old [1, 2N] biased the Est* estimate low by ~0.5/N. Design: - Drop runtime enable/disable (now an explicit Non-requirement): the sample knob is read at startup only; park the counter when off. Removes the DISABLED_RESEED re-park and the on/off reconciliation logic. - Simulation samples 1-in-10 (was 1-in-2). Portability (Windows is not build-tested here; lean on existing abstractions): - Aligned operator new uses platform::aligned_alloc/aligned_free (overflow-guarded) instead of posix_memalign; guard <pthread.h> under __linux__; add a force_noinline macro (GNU-only, empty elsewhere); portable volatile-sink escape() in the test. Tooling/tests/docs: - mako A/B scripts: validate the scratch mount (realpath + tmpfs + denylist) before rm -rf, locate mako_storage_bench.sh via __file__, and black-format. - Add operatorNewHonorsNewHandler and samplingRate tests; drop enableAfterOff. - Reconcile the design doc with the code; prune low-value comments. Testing: - /flow/MemoryTracker/* unit tests: 11 pass, 0 fail (bin/fdbserver -r unittests). - Joshua 100k (correctness-8.0.0): 99,995 pass / 1 fail. The single failure is NativeCdcAssignmentPublication -- a NativeCdc test from main, not this PR (CommitProxy failed_to_progress -> QuietDatabase DataDistributionActive -> NativeCdcEndToEnd workload start timed_out). It reproduces bit-identically (seed 2939264355, unseed 8776) with the tracker built off (sim sample_inverse=0, confirmed via the MemoryTrackerSummary SampleInverse trace field), so it is a pre-existing rare CDC/QuietDatabase flake unrelated to this change. * memory-tracker: fix GCC IPA-clone breaking frame-attribution tests Under GCC -O3, IPA constant-propagation cloning (-fipa-cp-clone) specializes the MemoryTrackerTest sentinels (each called with a constant N) into `.constprop` clones emitted at a different address than the function symbol. The executed code -- and thus the captured return addresses -- live in the clone, so the tests' frameInside(frame, &sentinel) window missed them: fastAlloc32Accounting aborted (sitesWithSentinelFrames == 0) in GCC CI while clang passed. Add `noclone` to force_noinline on GCC so each sentinel stays a single body at the address &fn yields. clang doesn't support noclone (and doesn't clone this way), so it keeps noinline only; other compilers stay empty. Testing: /flow/MemoryTracker/* passes 11/11 under both a gcc-toolset-13 build and a clang build. * memory-tracker: cheaper disabled alloc hot path (per-thread off flag) When sampling is off, memTrackerOnAlloc previously still did three TLS accesses every allocation -- a gInMemTracker load, a gMemTrackerCounter load-decrement- STORE, and a gForceSampleBytes load. Add a per-thread gMemTrackerOff flag, checked first, that a thread sets once its slow path observes sampling is off; the disabled alloc path then short-circuits on a single TLS load + branch (no counter store, no gForceSampleBytes load). The flag is per-thread, not global, on purpose: the counter still bootstraps sampling per thread (first alloc reaches the slow path and reads the knob), so a single global gate set by an early main-thread allocation before FLOW_KNOBS is ready would wrongly disable worker threads that bootstrap later. Free keeps the global g_memTrackerEnabled gate (a free-only thread must see global state to debit). memTrackerResetForTest clears the flag. Removes the now-dead INT_MAX counter parking. The saving is real but small (~2 loads + 1 store per alloc, well under the mako A/B's run-to-run noise), so the redwood off-state A/B shows no resolvable change; the win is on principle / at high allocation rates. Testing: /flow/MemoryTracker/* passes 11/11 (bin/fdbserver -r unittests). * memory-tracker: add FDB_MEMORY_TRACKER compile-time gate + per-path microbench Add a compile-time switch, FDB_MEMORY_TRACKER (CMake option, default ON), that removes the feature entirely when set to 0: the header hooks become no-op inlines, flow/MemoryTracker.cpp and the MemoryTrackerTest TEST_CASEs are #if'd out (the forceLink stub stays), and fdbserver/GlobalNewDelete.cpp defines no global operator new/delete override (libc++'s allocator is used, which still honors the installed new_handler). This gives operators an escape hatch to zero always-compiled footprint, and lets the microbench measure present-but-disabled vs absent cost per allocation path. Extend fdbserver/bench/BenchMemoryTracker.cpp with plain per-path alloc/free loops (operator new[] at several sizes, FastAllocator<64/96/256>, Arena medium/huge) so the same binary built =1 (tracker present, sampling off by default) vs =0 (absent) isolates each path's unweighted off-state overhead. Measured (EPYC 9R14 @ 3.7 GHz, ns/op, =1 minus =0): operator new[] small ~+1.1 ns (~11%); huge ~+2.5 ns FastAllocator<N> ~0 (within the ~0.6 ns cross-build noise floor) Arena block ~+3.3 ns (medium) / +4.8 ns (huge) The per-op-costly paths (operator new, Arena) are the low-volume ones; the high-volume path (FastAllocator, ~82% of redwood allocs) is ~free, so the weighted off-state overhead is ~0.1-0.2% of a core -- under the R0 target, though R0 is now framed as an unproven target with this gate as the escape hatch. Testing: /flow/MemoryTracker/* passes 11/11 (default build); FDB_MEMORY_TRACKER=OFF builds clean under -Werror. * clang-tidy fix * comment about why fdbserver/bench exists * address Codex round 5 review comments (MSVC build, startup sequencing, memory allocation failures on tracker-internal bookkeeping, cross-compile, contrib script cleanup) * add another disclaimer about MSVC/Windows being best effort * memory-tracker: fail-open ordering + deflake huge-arena unit test memTrackerSampleAlloc now performs both table insertions (aggregation-map and live-map nodes) before mutating any per-site or global counter, so if the tracker's own map growth throws std::bad_alloc the exception unwinds with all totals in lockstep and the fail-open catch simply drops the sample. This addresses an adversarial-review finding. Note such a failure does not occur in practice: FDB "OOM" is an RSS threshold enforced by fdbmonitor (typically ~12-16 GB against an ~8 GB target), not a malloc/operator-new failure. This feature is for finding leaks and untuned allocations that drive RSS growth, well short of any allocation failure -- documented in flow/MemoryTracker.cpp and design/memory-tracker.md so reviewers don't over-index on the bad_alloc path. arenaHugeAccounting now identifies the huge blocks by their ~100 KB size signature instead of the frame-pointer sentinel. The huge-Arena path is the deepest tracked call chain; under a real (non-simulation) network the best-effort frame-pointer walker is per-allocation nondeterministic there and cannot reliably attribute the blocks to the test's frame, so the sentinel-frame assertion flaked (~1-4% of runs). The recorded block size is correct regardless of which frames were captured, no incidental or foreign-thread allocation comes near 100 KB, and double-tracking still shows up as 2N. The shallower fastAlloc32/arenaSmall/arenaMedium/operatorNew accounting tests keep exact sentinel-frame attribution (reliable on their shorter paths) as the regression guard. Testing: - /flow/MemoryTracker/* via `fdbserver -r unittests`: 500 runs across two seed spaces (including the sequence that previously flaked), 14/14 test cases pass every run, 0 failures. - Builds clean with FDB_MEMORY_TRACKER on and off under -Werror; clang-format and clang-tidy clean on the changed files. - Joshua 100k on the parent commit: ensemble 20260728-160205-gglass-0b30fa74f2605807, ended=100000 pass=100000 fail=0. These changes are a pure counter-update reorder (no simulation-reachable behavior change -- bad_alloc is not injected in sim), a unit-test-only change, and documentation, so that result remains representative. * memory-tracker: note why the fdbserver-local bench binary exists Add a short comment to BenchMain.cpp explaining that most microbenchmarks belong in flow/bench and this fdbserver-local benchmark binary exists only for benchmarks that must link fdbserver-only code, pointing at BenchMemoryTracker.cpp for the detailed rationale rather than duplicating it.
2026-08-01 01:12:50 +08:00
/*
* MemoryTrackerTest.cpp
*
* This source file is part of the FoundationDB open source project
*
* Copyright 2013-2026 Apple Inc. and the FoundationDB project authors
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
// Unit tests for the per-call-site memory tracker.
//
// The "coverage" test uses sentinel functions: each sentinel triggers exactly
// one allocation path (operator new, FastAllocator, Arena), and the test
// confirms that some call site in the aggregation table contains a frame
// inside that sentinel's body. We compare raw return-address values against
// function-pointer values at runtime, so this works on stripped builds with
// no symbolization.
#include "flow/Arena.h"
#include "flow/FastAlloc.h"
#include "flow/Knobs.h"
#include "flow/MemoryTracker.h"
#include "flow/Platform.h"
#include "flow/UnitTest.h"
#include <climits>
#include <cstdint>
#include <cstdlib>
#include <limits>
#include <new>
#include <vector>
// Force this TU to link. The TEST_CASE macro registers via a static
// initializer; in a static library, a TU containing only static initializers
// gets dropped by the linker because nothing references its symbols.
// fdbserver/workloads/UnitTests.cpp calls this function to keep the TU.
void forceLinkMemoryTrackerTests() {}
#if FDB_MEMORY_TRACKER
namespace {
// A sentinel is an out-of-line function that performs exactly one kind of
// allocation, then returns its own address. We use the returned address to
// recognize captured stack frames that fell inside the sentinel's body.
constexpr uintptr_t SENTINEL_FUNC_SIZE = 4096;
// Defeat clang -O3 heap elision (P0593): if the allocated pointer doesn't
// escape, the compiler is free to drop the new/delete pair entirely, which
// then never reaches our operator-new override and the test sees zero samples.
void* volatile gEscapeSink;
inline void escape(void* p) {
gEscapeSink = p;
}
bool frameInside(void* frame, void* sentinel) {
uintptr_t f = reinterpret_cast<uintptr_t>(frame);
uintptr_t s = reinterpret_cast<uintptr_t>(sentinel);
return f >= s && f < s + SENTINEL_FUNC_SIZE;
}
force_noinline void* triggerOperatorNewSentinel(int n, int k) {
for (int i = 0; i < n; i++) {
auto* p = new int[k];
p[0] = i;
escape(p);
delete[] p;
}
return reinterpret_cast<void*>(&triggerOperatorNewSentinel);
}
force_noinline void* triggerFastAllocSentinel(int n) {
for (int i = 0; i < n; i++) {
void* p = FastAllocator<32>::allocate();
escape(p);
FastAllocator<32>::release(p);
}
return reinterpret_cast<void*>(&triggerFastAllocSentinel);
}
force_noinline void* triggerArenaSentinel(int n) {
// Force ArenaBlock::create by allocating large enough chunks to exceed
// the small-block threshold.
for (int i = 0; i < n; i++) {
Arena a;
// One ~512-byte allocation per arena -> goes through allocateAndMaybeKeepalive
// path which has the explicit Arena hook.
auto* p = new (a) uint8_t[600];
escape(p);
}
return reinterpret_cast<void*>(&triggerArenaSentinel);
}
// ---------------------------------------------------------------------------
// Accounting tests: verify byte/block counts come out right per allocation path
// and there's no double-tracking.
force_noinline void* allocateArenaMediumSentinel(int n, std::vector<Arena>& arenas) {
for (int i = 0; i < n; i++) {
arenas.emplace_back();
auto* p = new (arenas.back()) uint8_t[600];
escape(p);
}
return reinterpret_cast<void*>(&allocateArenaMediumSentinel);
}
force_noinline void* allocateArenaHugeSentinel(int n, std::vector<Arena>& arenas) {
for (int i = 0; i < n; i++) {
arenas.emplace_back();
auto* p = new (arenas.back()) uint8_t[100000];
escape(p);
}
return reinterpret_cast<void*>(&allocateArenaHugeSentinel);
}
force_noinline void* allocateArenaSmallSentinel(int n, std::vector<Arena>& arenas) {
for (int i = 0; i < n; i++) {
arenas.emplace_back();
auto* p = new (arenas.back()) uint8_t[64];
escape(p);
}
return reinterpret_cast<void*>(&allocateArenaSmallSentinel);
}
force_noinline void* allocateOperatorNewSentinel(int n, int k, std::vector<int*>& ptrs) {
for (int i = 0; i < n; i++) {
auto* p = new int[k];
escape(p);
ptrs.push_back(p);
}
return reinterpret_cast<void*>(&allocateOperatorNewSentinel);
}
force_noinline void* allocateFastAlloc32Sentinel(int n, std::vector<void*>& ptrs) {
for (int i = 0; i < n; i++) {
void* p = FastAllocator<32>::allocate();
escape(p);
ptrs.push_back(p);
}
return reinterpret_cast<void*>(&allocateFastAlloc32Sentinel);
}
force_noinline void releaseFastAlloc32(std::vector<void*>& ptrs) {
for (void* p : ptrs) {
FastAllocator<32>::release(p);
}
ptrs.clear();
}
struct AccountingSummary {
int sitesWithSentinelFrames = 0;
int64_t cumBytesSentinel = 0;
int64_t cumAllocsSentinel = 0;
int64_t liveBytesSentinel = 0;
int64_t liveCountSentinel = 0;
int totalSites = 0;
};
AccountingSummary collectAccounting(void* sentinel) {
AccountingSummary acc;
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
acc.totalSites++;
bool touches = false;
for (int i = 0; i < s.exemplarFrameCount; i++) {
if (frameInside(s.exemplarFrames[i], sentinel)) {
touches = true;
break;
}
}
if (touches && s.cumulativeBytes > 0) {
acc.sitesWithSentinelFrames++;
acc.cumBytesSentinel += s.cumulativeBytes;
acc.cumAllocsSentinel += s.cumulativeAllocs;
acc.liveBytesSentinel += s.liveBytes;
acc.liveCountSentinel += s.liveCount;
}
});
return acc;
}
void dumpSitesForFailure(const char* tag) {
fprintf(stderr, "[%s] dumping all tracker sites:\n", tag);
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
fprintf(stderr,
" fp=%016llx liveBytes=%lld liveCount=%lld cumBytes=%lld cumAllocs=%lld frames=",
(unsigned long long)s.fingerprint,
(long long)s.liveBytes,
(long long)s.liveCount,
(long long)s.cumulativeBytes,
(long long)s.cumulativeAllocs);
for (int i = 0; i < s.exemplarFrameCount; i++) {
fprintf(stderr, "%p ", s.exemplarFrames[i]);
}
fprintf(stderr, "\n");
});
}
class KnobOverride {
public:
explicit KnobOverride(int inverse = 1) : prevInverse(FLOW_KNOBS->MEMORY_TRACKING_SAMPLE_INVERSE) {
auto* k = const_cast<FlowKnobs*>(FLOW_KNOBS);
k->MEMORY_TRACKING_SAMPLE_INVERSE = inverse;
}
~KnobOverride() {
auto* k = const_cast<FlowKnobs*>(FLOW_KNOBS);
k->MEMORY_TRACKING_SAMPLE_INVERSE = prevInverse;
}
private:
int prevInverse;
};
} // namespace
TEST_CASE("/flow/MemoryTracker/coverage") {
#ifndef __linux__
// captureFramesFP is a no-op stub on non-Linux (FP walking through libc
// can't be made reliable on macOS); tests that inspect captured frames
// have nothing to inspect. Skip cleanly. The tracker still compiles
// and the non-frame tests (offSwitch, freeOfUntrackedPtrIsNoop) still
// run.
return Void();
#endif
// Sample everything, reset, run sentinels, check.
KnobOverride ko;
memTrackerResetForTest();
void* opNew = triggerOperatorNewSentinel(50, 4);
void* fastAlloc = triggerFastAllocSentinel(50);
void* arena = triggerArenaSentinel(50);
bool foundOpNew = false;
bool foundFastAlloc = false;
bool foundArena = false;
int siteCount = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
siteCount++;
for (int i = 0; i < s.exemplarFrameCount; i++) {
if (frameInside(s.exemplarFrames[i], opNew)) {
foundOpNew = true;
}
if (frameInside(s.exemplarFrames[i], fastAlloc)) {
foundFastAlloc = true;
}
if (frameInside(s.exemplarFrames[i], arena)) {
foundArena = true;
}
}
});
if (!foundOpNew || !foundFastAlloc || !foundArena) {
fprintf(stderr,
"MemoryTracker/coverage: sites=%d opNewSentinel=%p fastAllocSentinel=%p arenaSentinel=%p\n",
siteCount,
opNew,
fastAlloc,
arena);
fprintf(stderr,
"MemoryTracker/coverage: foundOpNew=%d foundFastAlloc=%d foundArena=%d\n",
foundOpNew,
foundFastAlloc,
foundArena);
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
fprintf(stderr,
" site fp=%016llx liveBytes=%lld cumAllocs=%lld frames=",
(unsigned long long)s.fingerprint,
(long long)s.liveBytes,
(long long)s.cumulativeAllocs);
for (int i = 0; i < s.exemplarFrameCount; i++) {
fprintf(stderr, "%p ", s.exemplarFrames[i]);
}
fprintf(stderr, "\n");
});
}
ASSERT(foundOpNew);
ASSERT(foundFastAlloc);
ASSERT(foundArena);
return Void();
}
TEST_CASE("/flow/MemoryTracker/offSwitch") {
// With sample inverse 0, no allocations are attributed. memTrackerResetForTest
// arms this thread's off-latch straight from the knob (as memTrackerInit does
// at startup), so even the first allocation short-circuits.
auto* k = const_cast<FlowKnobs*>(FLOW_KNOBS);
int prev = k->MEMORY_TRACKING_SAMPLE_INVERSE;
k->MEMORY_TRACKING_SAMPLE_INVERSE = 0;
memTrackerResetForTest();
for (int i = 0; i < 100; i++) {
auto* p = new int[4];
p[0] = i;
delete[] p;
}
int siteCount = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite&) { siteCount++; });
ASSERT_EQ(siteCount, 0);
// The enabled flag gates the free hot path: with sampling off it must be
// false, so memTrackerOnFree short-circuits before taking g_mtLock.
ASSERT(!g_memTrackerEnabled.value.load(std::memory_order_relaxed));
k->MEMORY_TRACKING_SAMPLE_INVERSE = prev;
return Void();
}
TEST_CASE("/flow/MemoryTracker/operatorNewHonorsNewHandler") {
// The global operator new override (fdbserver/GlobalNewDelete.cpp) must run
// the std::new_handler retry loop so an allocation failure reaches FDB's OOM
// path instead of throwing straight past it. (Where the override isn't linked,
// the standard library's operator new provides the same contract, so this
// still passes.)
static bool handlerRan;
handlerRan = false;
std::new_handler prev = std::set_new_handler([]() {
handlerRan = true;
throw std::bad_alloc(); // break the retry loop
});
bool caught = false;
try {
// volatile so the compiler can't fold the size and warn (-Walloc-size); malloc
// reliably fails for SIZE_MAX, driving the handler loop.
volatile std::size_t huge = std::numeric_limits<std::size_t>::max();
void* p = ::operator new(huge);
escape(p);
} catch (const std::bad_alloc&) {
caught = true;
}
std::set_new_handler(prev);
ASSERT(handlerRan);
ASSERT(caught);
return Void();
}
TEST_CASE("/flow/MemoryTracker/samplingRate") {
// The reseed gap is uniform on [1, 2N-1] (mean N), so at inverse N the sampled
// fraction should be ~1/N. Wide bounds keep it non-flaky across RNG state.
constexpr int N = 10;
constexpr int ALLOCS = 200000;
KnobOverride ko(N);
memTrackerResetForTest();
std::vector<int*> ptrs;
ptrs.reserve(ALLOCS);
for (int i = 0; i < ALLOCS; i++) {
auto* p = new int[4];
escape(p);
ptrs.push_back(p);
}
int64_t sampled = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
if (s.forceSampledCount == 0) {
sampled += s.cumulativeAllocs;
}
});
for (auto* p : ptrs) {
delete[] p;
}
ptrs.clear();
double frac = double(sampled) / ALLOCS;
ASSERT(frac > 0.06 && frac < 0.15); // expect ~0.1
memTrackerResetForTest();
return Void();
}
TEST_CASE("/flow/MemoryTracker/freeOfUntrackedPtrIsNoop") {
// memTrackerOnFree on a pointer the tracker never recorded must be a no-op.
KnobOverride ko;
memTrackerResetForTest();
int x = 0;
memTrackerOnFree(&x); // not in any table
memTrackerOnFree(nullptr);
int siteCount = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite&) { siteCount++; });
ASSERT_EQ(siteCount, 0);
return Void();
}
TEST_CASE("/flow/MemoryTracker/cumulativeIsMonotonic") {
#ifndef __linux__
return Void(); // see /coverage for rationale
#endif
// liveCount must return to ~0 after we free everything we allocated;
// cumulativeAllocs must NOT decrement.
KnobOverride ko;
memTrackerResetForTest();
void* sentinel = triggerOperatorNewSentinel(100, 8);
int64_t maxCumulative = 0;
int64_t finalLive = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
for (int i = 0; i < s.exemplarFrameCount; i++) {
if (frameInside(s.exemplarFrames[i], sentinel)) {
if (s.cumulativeAllocs > maxCumulative) {
maxCumulative = s.cumulativeAllocs;
}
finalLive += s.liveCount;
}
}
});
ASSERT(maxCumulative >= 100);
ASSERT_EQ(finalLive, 0); // every alloc was paired with delete
return Void();
}
TEST_CASE("/flow/MemoryTracker/estimateScaling") {
// End-to-end estimate check. With a fixed inverse N > 1 and no force-sampled
// blocks, every sample at a site carries weight N, so the site's estimated
// usage must be *exactly* N times its raw sampled counters. This verifies
// the reported Est* numbers without depending on which specific allocations
// happened to be sampled. Runs on all platforms (no frame inspection).
constexpr int N = 8;
KnobOverride ko(N);
memTrackerResetForTest();
// Small allocations far below the force-sample threshold, so none
// are force-sampled and every sampled block gets weight N.
std::vector<int*> ptrs;
ptrs.reserve(5000);
for (int i = 0; i < 5000; i++) {
auto* p = new int[4];
escape(p);
ptrs.push_back(p);
}
int checked = 0;
int64_t rawLiveBefore = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
if (s.forceSampledCount != 0) {
return; // ignore any incidental force-sampled site (weight 1, not N)
}
ASSERT_EQ(s.estCumulativeBytes, s.cumulativeBytes * N);
ASSERT_EQ(s.estCumulativeAllocs, s.cumulativeAllocs * N);
ASSERT_EQ(s.estLiveBytes, s.liveBytes * N);
ASSERT_EQ(s.estLiveCount, s.liveCount * N);
ASSERT_EQ(s.estPeakBytes, s.peakBytes * N);
rawLiveBefore += s.liveBytes;
checked++;
});
ASSERT(checked > 0);
ASSERT(rawLiveBefore > 0);
for (auto* p : ptrs) {
delete[] p;
}
ptrs.clear();
// Symmetric debit: the per-site scaling invariant must still hold after the
// frees (each free debits the estimate by exactly weight×size), and the live
// total must have dropped. We check the invariant rather than "live == 0"
// because incidental still-live allocations (e.g. the ptrs vector's own
// backing buffer) legitimately remain tracked.
int64_t rawLiveAfter = 0;
int64_t estLiveAfter = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
if (s.forceSampledCount != 0) {
return;
}
ASSERT_EQ(s.estLiveBytes, s.liveBytes * N);
rawLiveAfter += s.liveBytes;
estLiveAfter += s.estLiveBytes;
});
ASSERT_EQ(estLiveAfter, rawLiveAfter * N);
ASSERT(rawLiveAfter < rawLiveBefore); // the freed blocks were debited
memTrackerResetForTest();
return Void();
}
// ---------------------------------------------------------------------------
// Accounting tests. The "sites with the sentinel's frames" assertion is the
// main check here.
TEST_CASE("/flow/MemoryTracker/fastAlloc32Accounting") {
#ifndef __linux__
return Void(); // see /coverage for rationale
#endif
KnobOverride ko;
constexpr int N = 30;
std::vector<void*> ptrs;
ptrs.reserve(N);
memTrackerResetForTest();
void* sentinel = allocateFastAlloc32Sentinel(N, ptrs);
auto pre = collectAccounting(sentinel);
if (pre.sitesWithSentinelFrames != 1) {
dumpSitesForFailure("fastAlloc32Accounting/post-alloc");
}
ASSERT_EQ(pre.sitesWithSentinelFrames, 1);
ASSERT_EQ(pre.cumAllocsSentinel, N);
ASSERT_EQ(pre.liveCountSentinel, N);
ASSERT_EQ(pre.cumBytesSentinel, int64_t(N) * 32);
ASSERT_EQ(pre.liveBytesSentinel, pre.cumBytesSentinel);
// Global totals are intentionally not asserted: at inverse=1 a foreign-thread
// allocation in the window would break a strict global equality (flaky at
// Joshua scale); the sentinel-scoped checks above pin the regression.
releaseFastAlloc32(ptrs);
auto post = collectAccounting(sentinel);
ASSERT_EQ(post.liveBytesSentinel, 0);
ASSERT_EQ(post.liveCountSentinel, 0);
// Global live totals intentionally not asserted (flaky at inverse=1; see above).
ASSERT_EQ(post.cumAllocsSentinel, N); // cumulative never decrements
return Void();
}
TEST_CASE("/flow/MemoryTracker/arenaSmallAccounting") {
#ifndef __linux__
return Void(); // see /coverage for rationale
#endif
KnobOverride ko;
constexpr int N = 30;
std::vector<Arena> arenas;
arenas.reserve(N);
memTrackerResetForTest();
void* sentinel = allocateArenaSmallSentinel(N, arenas);
auto pre = collectAccounting(sentinel);
if (pre.sitesWithSentinelFrames != 1) {
dumpSitesForFailure("arenaSmallAccounting/post-alloc");
}
ASSERT_EQ(pre.sitesWithSentinelFrames, 1);
ASSERT_EQ(pre.cumAllocsSentinel, N);
ASSERT_EQ(pre.liveCountSentinel, N);
ASSERT_EQ(pre.liveBytesSentinel, pre.cumBytesSentinel);
// Global totals are intentionally not asserted: at inverse=1 a foreign-thread
// allocation in the window would break a strict global equality (flaky at
// Joshua scale); the sentinel-scoped checks above pin the regression.
arenas.clear();
auto post = collectAccounting(sentinel);
ASSERT_EQ(post.liveBytesSentinel, 0);
ASSERT_EQ(post.liveCountSentinel, 0);
// Global live totals intentionally not asserted (flaky at inverse=1; see above).
ASSERT_EQ(post.cumAllocsSentinel, N);
return Void();
}
TEST_CASE("/flow/MemoryTracker/arenaMediumAccounting") {
#ifndef __linux__
return Void(); // see /coverage for rationale
#endif
// Make sure arenas aren't counted twice, once due to their direct
// instrumentation and a second time due to their use of operator new.
KnobOverride ko;
constexpr int N = 30;
std::vector<Arena> arenas;
arenas.reserve(N);
memTrackerResetForTest();
void* sentinel = allocateArenaMediumSentinel(N, arenas);
auto pre = collectAccounting(sentinel);
if (pre.sitesWithSentinelFrames != 1) {
dumpSitesForFailure("arenaMediumAccounting/post-alloc");
}
ASSERT_EQ(pre.sitesWithSentinelFrames, 1);
ASSERT_EQ(pre.cumAllocsSentinel, N);
ASSERT_EQ(pre.liveCountSentinel, N);
ASSERT_EQ(pre.liveBytesSentinel, pre.cumBytesSentinel);
// Global totals intentionally not asserted (flaky at inverse=1; see above).
arenas.clear();
auto post = collectAccounting(sentinel);
ASSERT_EQ(post.liveBytesSentinel, 0);
ASSERT_EQ(post.liveCountSentinel, 0);
// Global live totals intentionally not asserted (flaky at inverse=1; see above).
ASSERT_EQ(post.cumAllocsSentinel, N);
return Void();
}
TEST_CASE("/flow/MemoryTracker/arenaHugeAccounting") {
#ifndef __linux__
return Void(); // see /coverage for rationale
#endif
// Verify the huge Arena-block path (reqSize >= LARGE), including that a block is
// tracked once (not double-counted by both the explicit Arena hook and the inner
// operator new[]). This is the deepest tracked call chain, and under a real
// (non-simulation) network its frame-pointer backtrace is unreliable — the
// best-effort walker may, run to run and even alloc to alloc, fail to climb to
// the test's frame or attribute the blocks to varying fingerprints. So instead of
// the sentinel-frame approach the other *Accounting tests use, we identify the
// huge blocks by their unmistakable ~100 KB size signature: the recorded size is
// correct regardless of which frames were captured, and no incidental or
// foreign-thread allocation comes anywhere near this large.
KnobOverride ko;
constexpr int N = 10;
constexpr int64_t BLOCK = 100000;
constexpr int64_t HUGE_MIN = 90000; // mean bytes/alloc of a huge site; nothing else is this big
std::vector<Arena> arenas;
arenas.reserve(N);
memTrackerResetForTest();
allocateArenaHugeSentinel(N, arenas);
int64_t cumAllocs = 0, cumBytes = 0, liveBytes = 0, liveCount = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
if (s.cumulativeAllocs > 0 && s.cumulativeBytes / s.cumulativeAllocs >= HUGE_MIN) {
cumAllocs += s.cumulativeAllocs;
cumBytes += s.cumulativeBytes;
liveBytes += s.liveBytes;
liveCount += s.liveCount;
}
});
// Exactly N huge blocks, each tracked once (double-tracking would show as 2N),
// all currently live, with live bytes == cumulative bytes (nothing freed yet).
ASSERT_EQ(cumAllocs, N);
ASSERT_EQ(liveCount, N);
ASSERT_EQ(liveBytes, cumBytes);
ASSERT(cumBytes >= int64_t(N) * BLOCK);
arenas.clear();
int64_t cumAllocsPost = 0, liveBytesPost = 0, liveCountPost = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite& s) {
if (s.cumulativeAllocs > 0 && s.cumulativeBytes / s.cumulativeAllocs >= HUGE_MIN) {
cumAllocsPost += s.cumulativeAllocs;
liveBytesPost += s.liveBytes;
liveCountPost += s.liveCount;
}
});
// After freeing, the huge blocks are debited: live returns to 0, cumulative persists.
ASSERT_EQ(liveBytesPost, 0);
ASSERT_EQ(liveCountPost, 0);
ASSERT_EQ(cumAllocsPost, N);
return Void();
}
TEST_CASE("/flow/MemoryTracker/operatorNewAccounting") {
#ifndef __linux__
return Void(); // see /coverage for rationale
#endif
KnobOverride ko;
constexpr int N = 30;
constexpr int K = 8; // new int[8] -> 32 bytes; int is trivial so no array cookie
std::vector<int*> ptrs;
ptrs.reserve(N);
memTrackerResetForTest();
void* sentinel = allocateOperatorNewSentinel(N, K, ptrs);
auto pre = collectAccounting(sentinel);
if (pre.sitesWithSentinelFrames != 1) {
dumpSitesForFailure("operatorNewAccounting/post-alloc");
}
ASSERT_EQ(pre.sitesWithSentinelFrames, 1);
ASSERT_EQ(pre.cumAllocsSentinel, N);
ASSERT_EQ(pre.liveCountSentinel, N);
ASSERT_EQ(pre.cumBytesSentinel, static_cast<int64_t>(N) * K * static_cast<int64_t>(sizeof(int)));
ASSERT_EQ(pre.liveBytesSentinel, pre.cumBytesSentinel);
// Global totals intentionally not asserted (flaky at inverse=1; see above).
for (auto* p : ptrs) {
delete[] p;
}
ptrs.clear();
auto post = collectAccounting(sentinel);
ASSERT_EQ(post.liveBytesSentinel, 0);
ASSERT_EQ(post.liveCountSentinel, 0);
// Global live totals intentionally not asserted (flaky at inverse=1; see above).
ASSERT_EQ(post.cumAllocsSentinel, N);
return Void();
}
TEST_CASE("/flow/MemoryTracker/failOpenOnMetadataAllocFailure") {
// The tracker must fail open: if its own metadata allocation throws, the
// underlying user allocation still succeeds and tracking recovers on this
// thread afterward (the reentrancy guard is restored, not leaked). Runs on all
// platforms — no frame inspection.
KnobOverride ko; // inverse = 1: sample every allocation
memTrackerResetForTest();
// Arm the one-shot; it is consumed by the next sampled allocation, which throws
// inside the tracker. new/delete must not observe that exception.
memTrackerFailNextSampleForTest();
int* p = new int[4];
ASSERT(p != nullptr);
p[0] = 42;
int observed = p[0];
delete[] p;
ASSERT_EQ(observed, 42);
// Tracking must still work after the injected failure.
memTrackerResetForTest();
std::vector<int*> ptrs;
ptrs.reserve(8);
for (int i = 0; i < 8; i++) {
auto* q = new int[4];
escape(q);
ptrs.push_back(q);
}
int siteCount = 0;
memTrackerForEachSite([&](const MemoryTrackerCallSite&) { siteCount++; });
for (auto* q : ptrs) {
delete[] q;
}
ASSERT(siteCount > 0);
return Void();
}
TEST_CASE("/flow/MemoryTracker/initEnablesFromKnob") {
// Exercises the *production* enablement path (memTrackerInit), not the test-only
// reset — the reviewer noted that memTrackerResetForTest masks the startup race.
// memTrackerInit must set the global enabled flag and this thread's fast-path
// off-latch straight from MEMORY_TRACKING_SAMPLE_INVERSE, and (the regression)
// must re-arm a thread that a prior off-configuration had already latched off.
auto* k = const_cast<FlowKnobs*>(FLOW_KNOBS);
int prev = k->MEMORY_TRACKING_SAMPLE_INVERSE;
// Sampling off: init publishes disabled and latches this thread off. This is the
// state an early startup allocation used to get stuck in before init existed.
k->MEMORY_TRACKING_SAMPLE_INVERSE = 0;
memTrackerInit();
ASSERT(!g_memTrackerEnabled.value.load(std::memory_order_relaxed));
ASSERT(gMemTrackerOff); // alloc hot path short-circuits
// Knobs now configured with sampling on: init must re-arm THIS thread. The bug
// was that nothing re-armed the network thread once it latched off before the
// knobs were ready, so sampling stayed dead for the life of the process.
k->MEMORY_TRACKING_SAMPLE_INVERSE = 8;
memTrackerInit();
ASSERT(g_memTrackerEnabled.value.load(std::memory_order_relaxed));
ASSERT(!gMemTrackerOff); // re-armed: alloc hot path now reaches the sampler
// And the reverse transition: a subsequent off-configuration re-latches it.
k->MEMORY_TRACKING_SAMPLE_INVERSE = 0;
memTrackerInit();
ASSERT(!g_memTrackerEnabled.value.load(std::memory_order_relaxed));
ASSERT(gMemTrackerOff);
k->MEMORY_TRACKING_SAMPLE_INVERSE = prev;
memTrackerInit(); // restore tracker state from the baseline knob for later tests
return Void();
}
#endif // FDB_MEMORY_TRACKER