max_memory_mb cap and is OOM-killed every ~3 hours on Zoesmartbotic-database 1.7.5-1 (deb), installed on Debian 13 (trixie).zoe.shadowman.fsociety.hu).The DB process grows past its configured max_memory_mb budget and keeps climbing until the kernel OOM-killer reaps it. The internal eviction logic is firing — estimatedMemoryBytes_ gets reduced — but actual RSS does not. This means allocations are leaking out of an accounting path that the budget tracks, OR are happening in code paths the budget never tracked.
shadowman-server, shadowman-gateway, shadowman-llm, shadowman-tools, shadowman-stt, shadowman-tts, shadowman-cron, smartbotic-database, redis-server). Database is by far the largest./etc/smartbotic-database/conf.d/99-local.json (operator override):
{ "storage": { "memory": {
"max_memory_mb": 6144,
"eviction_threshold_percent": 90,
"eviction_target_percent": 75
}}}
Defaults from base + 50-shadowman.json for everything else, notably:
max_file_size_mb: 500
grpc.max_receive_message_size_mb: 100
grpc.max_send_message_size_mb: 100
grpc.resource_quota_memory_mb: 256
snapshot_interval_sec: 3600
encryption.enabled: true
Recent OOM samples from dmesg -T --level=err,warn (every line is smartbotic-data as the victim — the named "invoker" is just whoever was unlucky enough to allocate next):
Mon May 11 23:41:29 — killed pid 441106 anon-rss:11757612 kB total-vm:21405132 kB
Tue May 12 01:29:19 — killed pid 452374 anon-rss:11789832 kB total-vm:20891196 kB
Tue May 12 04:06:31 — killed pid 455113 anon-rss:11753480 kB total-vm:20529072 kB
Tue May 12 07:05:14 — killed pid 458822 anon-rss:11783236 kB total-vm:20211244 kB
Tue May 12 09:02:12 — killed pid 468259 anon-rss:11739336 kB total-vm:21097712 kB
Tue May 12 12:01:57 — killed pid 472711 anon-rss:11599740 kB total-vm:19878288 kB
Tue May 12 15:02:59 — killed pid 487257 anon-rss:11631612 kB total-vm:19682928 kB
Wed May 13 00:07:19 — killed pid 531382 anon-rss:11750160 kB total-vm:19053096 kB
Pattern: anon-rss clamps to ~11.7 GB right before kill (essentially "all of RAM"), total-vm ~20 GB. Memory grows linearly until the kernel intervenes — there is no plateau anywhere near 6 GB (the configured cap) or 7.6 GB (85% emergency threshold).
estimatedMemoryBytes_ accounting code is correct on the document path (every insert/replace/delete updates the atomic counter via estimateDocumentSize).This is the smoking gun: the leak is in untracked allocations, or in heap fragmentation that prevents pages from being returned to the OS after eviction frees them.
50-shadowman.json mentions a 2026-04-22 incident).nlohmann::json nodes (numbers, strings, object/array nodes). glibc's malloc uses per-thread arenas, and small-allocation frees often don't return pages to the OS — they go on the thread-local free list. RSS stays high; the next alloc reuses those pages, so it doesn't grow infinitely, but it never shrinks either.malloc_trim(0) call after each eviction burst (or every N minutes) — see if RSS drops.MALLOC_ARENA_MAX=2 in the systemd unit's Environment= — limits arena count, reduces fragmentation.LD_PRELOAD= in the unit) — both are dramatically better at returning pages after fragmentation. This is a one-line change to verify.malloc_trim from the eviction path.snapshot_interval_sec: 3600. Every hour the entire state gets serialized.service/src/persistence/ — look for whether snapshots stream to disk incrementally or buffer the whole thing in memory.:00 (hourly), this is it.:02:12 and :01:57 past the hour — suggestive but not conclusive. Add a snapshot-start / snapshot-end log line and correlate.max_file_size_mb: 500. Each in-flight PutFile could allocate up to 500 MB if implemented as buffer-in-memory.grpc.max_concurrent_file_streams: 10 → worst case 5 GB simultaneous.service/src/files/ (or wherever the file-stream RPCs live) handles streamed chunks. Are they accumulated into a contiguous buffer? Or streamed to disk as they arrive?replication.enabled: false), so the outbox code shouldn't be active. Mark this low-priority but worth a quick grep.14eccea — "client retry writes on DEADLINE_EXCEEDED/RESOURCE_EXHAUSTED/UNAVAILABLE") added retry logic. If retried-then-cancelled streams don't release their gRPC buffers, that could trickle.grpcpp_sync_ser appears as an OOM-killer invoker in the dmesg log (the gRPC server's worker pool name), which means gRPC was actively allocating at kill time.encryption.enabled: true. The AES-256-GCM path likely staging-buffers each write before persisting. If those buffers don't shrink between writes (e.g. cached per-thread), heavy write workloads could pile pages.max_memory_mb: 6144).RES=$(pgrep smartbotic-data); while sleep 30; do ps -o rss= -p $RES; done > /tmp/rss.log to log RSS every 30 s.tests/load/ harness, scaled to the same document count Zoe runs.max_memory_mb.If reproducing on a smaller config is fine, drop max_memory_mb to 512 locally and watch whether RSS exceeds 1 GB. The leak should be proportional.
service/src/memory_store.{hpp,cpp} — the budget + eviction logic. estimateDocumentSize (line ~2390) is the only path that updates estimatedMemoryBytes_. Anything allocating outside MemoryStore::insertImpl/replaceImpl/deleteImpl paths is untracked.service/src/persistence/* — snapshot + WAL handlers. Look for buffered serialization.service/src/files/* — file-streaming RPCs.service/src/database_service.cpp:363 (config.maxMemoryMb = memory.value(...)) — confirms the config-binding path.The fix is acceptable when all of these hold:
1.25 × max_memory_mb for at least 8 hours of continuous load (we allow 25% headroom for allocator overhead, but not 100%).estimatedMemoryBytes_ tracks them.tests/load/ reproduces the original failure on a known-bad commit and passes on the fix commit.CHANGELOG.md (or commit message) documents what the leak was and why the fix works.v1.7.6 (or v1.8.0).packaging/. The repo machinery follows the same add-packages.sh → create-repo.sh --suite trixie → sync-repo.sh flow used by shadowman-cpp.external/smartbotic-database/VERSION doesn't pin a hard version, so a new deb on the FTP repo is enough; on Zoe apt install --only-upgrade smartbotic-database should take it.Environment=LD_PRELOAD=/usr/lib/x86_64-linux-gnu/libjemalloc.so.2 in the unit, restart, watch RSS. If RSS now stays under 6 GB, ship that as the immediate mitigation while you investigate (1).Stopped all shadowman-* services on Zoe, kept smartbotic-database running, sampled RSS every 60 s for 11 minutes (zero clients connected, no gRPC traffic):
RSS (KB) CPU%
6399968 59.9 t=0
6401836 58.7
6406924 57.7
6409024 56.7
6412340 55.7
6413276 54.7
6420368 53.6
6420920 52.6
6421788 51.6
6425260 50.7
6426356 49.8
6427176 48.9 t=11
Idle leak rate: ~2.4 MB/min (27 MB over 11 minutes). The CPU decline reflects the eviction loop catching up post-emergency-pressure; the memory growth is monotonic with zero client load.
Strong implication: this isn't an access-pattern problem — it's a real allocator-level leak in a background thread (likely the eviction loop itself, the WAL fsync thread, or the snapshot thread). Anything that allocates without going through estimateDocumentSize (which is what updates estimatedMemoryBytes_) will leak invisibly to the budget.
We also found that the docs-on-disk picture is much smaller than the tracked working set:
The gap is dominated by:
estimateDocumentSize's 2.5x heap multiplier (memory_store.cpp:2402) — applied to live JSON nodes.conversation_context_usage updated every LLM turn) accumulate hundreds of DocumentVersion entries in versionHistory[docId] — each one a full copy of the doc data, also counted with the 2.5x multiplier. A single conversation with 200 turns × ~5 KB context_usage doc × 2.5 ≈ 2.5 MB per doc; multiply by 10 conversations + similar churn on conversation_tool_state, conversation_active_skills, metrics_rollup_*, etc. and you get GBs of version history in RAM.Shadowman-cpp is shipping migration 062 (alter_collection setting max_versions: 1-3 on hot-update collections) which will collapse this. But it's a workaround — the real question for smartbotic-database is whether unbounded max_versions=0 should still be the default. A sensible default of 5 or 10 (with explicit unlimited for collections that opt in) would prevent this category of footgun.
1a. Before any heap-profiling work, measure idle leak per-thread: gdb attach the live process, capture info threads + per-thread stacks, see which thread is allocating. The eviction loop (MemoryStore::expirationLoop?) and the WAL/snapshot threads are the suspects.
1b. Add a periodic malloc_info(0, stderr) dump (or use mallctl stats if jemalloc) to compare arena usage between a freshly-started process and one that's been idle for an hour. Confirms whether the leak is in malloc arena fragmentation or in genuinely-unreleased allocations.
After deploying v1.8.0 on Zoe the workload pattern produced this:
Before MALLOC_ARENA_MAX: After MALLOC_ARENA_MAX=2:
VmData: 11.55 GB VmData: 5.70 GB
VmRSS: 7.97 GB VmRSS: 5.57 GB
VmSwap: 3.27 GB VmSwap: 0 KB
VmHWM: 10.77 GB VmHWM: 8.29 GB
Tracked: 4995 MB Tracked: ~4995 MB (same workload)
Setting MALLOC_ARENA_MAX=2 via the systemd unit closed the 6.5 GB gap between tracked memory and committed heap to ~500 MB. The DB's tracker was correct; the rest was glibc per-thread arena fragmentation. With 22 threads, each thread gets its own arena (by default 8 × num_cpus), small frees go on per-thread freelists and never make it back to madvise(MADV_DONTNEED). The evictDocument history-pruning fix from v1.8.0 IS valid (and necessary), but its effect was masked on Zoe because eviction never fired — the workload kept memory under the 85% threshold, so the leak surface that triggered was fragmentation, not unreleased history.
Recommendation for v1.8.x:
MALLOC_ARENA_MAX=2 as a recommended environment variable in the deb's systemd unit (or set it directly in the service file). Operators won't discover this on their own and we lost a working day to it.jemalloc or mimalloc instead — both return memory to the OS far more aggressively than glibc and remove the env-var requirement. mimalloc in particular has measured 30-50% RSS reductions for similar high-thread-count C++ servers.mallopt(M_ARENA_MAX, 2) programmatic equivalent in main.cpp would make this self-contained without env vars.The boot-time parallel snapshot deserialize from v1.8.0 also worked beautifully — DB went from systemd start to READY in <1 second on Zoe's 5 GB snapshot (sequential pre-v1.8 would have been ~30 s).
After running on Zoe for 30 minutes under real workload (single user, normal LLM/tool traffic — nothing heavy), v1.8.0 with MALLOC_ARENA_MAX=2 produced this growth pattern:
Internal tracker (Memory check log line):
09:44 — 4988 MB estimated, 1854705 docs, 0 evicted
09:54 — 4994 MB estimated, 1858821 docs, 0 evicted
10:03 — 4999 MB estimated, 1862015 docs, 0 evicted (+11 MB tracked / 19 min)
Process VM (from /proc/$pid/status):
After restart (~09:34): VmRSS=5.57 GB VmData=5.70 GB ← matches tracker
30 minutes later: VmRSS=10.3 GB VmData=10.5 GB ← +4.8 GB over 30 min
The leak is real, untracked by estimatedMemoryBytes_, and reproduces consistently. ~160 MB/min under modest workload. Eviction never fires (DB stays at 69% reported pressure) so the v1.8.0 evictDocument fix never engages.
Doc-count growth is only ~250-600/min (mostly metrics_events + small rollups), accounting for ~10-15 MB/min of legitimate tracked growth. The remaining ~145 MB/min is allocations the tracker doesn't see. Candidates from earlier suspect list (re-ranked now that we have this data):
insert/update builds a WAL record. If the WAL writer's per-thread buffer grows-but-doesn't-shrink, this scales with write traffic. The 250-600 doc/min write rate matches the ~150 MB/min leak rate./etc/systemd/system/smartbotic-database.service.d/arena-cap.conf on Zoe:
[Service]
Environment=MALLOC_ARENA_MAX=2
MemoryMax=9G
MemoryHigh=7G
RuntimeMaxSec=5400 ; 90-min watchdog — DB restarts cleanly before VmData reaches MemoryMax
This bounds the damage but is a hack — clients see a 3-5 s outage every 90 minutes when the watchdog fires. Acceptable while upstream debugs but not the long-term answer.
heaptrack /usr/bin/smartbotic-database --config /etc/smartbotic-database/config.json, drive realistic write traffic for 30 min, then heaptrack_print to find the top allocator paths with no matching frees. Should immediately surface the second leak — far cheaper than guessing.grpc::Arena has Reset()) instead of letting it grow with the channel.The good news from this datapoint: the v1.8.0 changes ARE working as advertised — the boot-time parallel deserialize halved snapshot recovery time, and the eviction history-prune is correct (we just haven't been able to trigger eviction). The remaining leak is a different bug that pre-dates the v1.8.0 work.
v1.9.0 deployed on Zoe. The off-heap history architecture closed the tracker gap to zero:
Pre-1.9.0 baseline (1.8.2):
Tracker: 5017 MB
VmRSS: 6.26 GB
VmData: 8.31 GB
Growth/min: ~3.5 MB (under modest load)
Post-1.9.0 (boot done, 11 min of light traffic):
Tracker: 1396 MB ← was 5017 MB (-3.6 GB, version history off heap)
VmRSS: 6.26 GB ← unchanged (the gap moved from "leak" to "arena retention")
VmData: 6.41 GB
Growth/min: 0 MB on tracker, 2.7 MB/min on VmData
After SIGUSR2 (malloc_trim from v1.8.3):
VmRSS: 3.51 GB ← dropped 2.76 GB in one call
System free: 5.1 GB ← was 2.5 GB
Tracker: 1.4 GB
Gap (Vm-track): ~2.1 GB ← reasonable: binary + gRPC channels + thread stacks + page tables
So:
Call malloc_trim(0) automatically once at the end of MemoryStore::loadSnapshot (or wherever boot-time deserialize completes). The pages are freeable; they're just sitting on the per-thread free-list because glibc doesn't proactively check thresholds during quiet periods. One-line addition next to the existing mallopt(M_ARENA_MAX, 2) from v1.8.1. Saves operators ~3 GB of RSS post-boot, makes the OOM-killer score much friendlier on tight VMs.
Optional: trim periodically (every 5 minutes? every snapshot completion?) for ongoing arena hygiene. The SIGUSR2 handler from v1.8.3 already proves the operation is safe under load. Production allocators like jemalloc / mimalloc do this automatically.
/etc/systemd/system/smartbotic-database.service.d/zoe-leak-watchdog.conf removed — 1.9.0 makes RuntimeMaxSec unnecessary..hlog file growth, which is the new constraint.Setup: app named "DB Stress Probe" with a 5-field cron schedule (* * * * *) that on each tick inserts 200 rows into probe_events, upserts probe_stats (single doc with max_versions: 5), and trims probe_events to the latest 1000 rows. After bumping apps_max_collection_rows from 500 to 100000 so the trim/insert ratio is realistic, ran 14 minutes / 9 successful cron ticks.
start end delta rate
Tracker (MB): 1408 1409 +1 0.07 MB/min
VmRSS (MB): 6112 6115 +3 0.19 MB/min
VmData (MB): 6275 6277 +2 0.17 MB/min
Docs: 1,873,653 1,890,558 +16,905 ~1200/min
Swap: 0 0 — —
Pressure: 19% 19% — normal
CPU: 7.3% 5.2% decreasing —
The architecture decoupling is complete. Write churn flows through WAL + hlog files on disk; memory is bounded by max_memory_mb and doesn't grow with workload. Compared to the pre-fix baseline (~200+ MB/min under load on 1.7.5), this is a >1000× improvement at steady-state.
The auto-trim from v1.9.1 doesn't fire here because no snapshot completed during the window (hourly snapshots; window was 13:19–13:34, next snapshot at ~13:53). When the snapshot lands, RSS will likely drop a few hundred MB closer to the tracker — observed previously on this run that snapshot trims drop RSS by 500-3000 MB depending on prior activity.
$ ls /etc/systemd/system/smartbotic-database.service.d/
# (empty — drop-in dir doesn't exist)
$ apt-cache policy smartbotic-database
Installed: 1.9.1-1
Candidate: 1.9.1-1
$ cat /etc/smartbotic-database/conf.d/99-local.json
{
"storage": {
"memory": {
"max_memory_mb": 7168,
"eviction_threshold_percent": 85,
"eviction_target_percent": 70
}
}
}
The 7 GB cap is the only customization, and even that is generous — the actual working set under load is ~1.4 GB tracked / ~6.3 GB resident. Could safely drop to 2-3 GB if you want to test eviction-under-pressure behavior; currently eviction stays dormant which is fine.
v1.9.1's malloc_trim(0) at end of loadSnapshot/createSnapshot is necessary but not sufficient on the Zoe workload pattern:
DB restart at ~14:42 (triggered by shadowman-common dpkg upgrade)
~23 min later:
Tracker: 1422 MB
VmRSS: 9.31 GB ← 7.9 GB of glibc arena pages
System avail: 1.6 GB ← squeezing the rest of the box
After manual SIGUSR2 (malloc_trim):
VmRSS: 3.60 GB ← -5.7 GB freed
System avail: 7.0 GB
The auto-trim only fires on snapshot boundaries (hourly). Between a fresh start and the first snapshot, the DB accumulates GBs of arena pages from:
loadSnapshot)replayWal allocates and frees a lot of small Document JSON nodes)Suggestion for v1.9.2 (small, one place to add):
MemoryStore's minute-tick "Memory check" path that already logs Memory check: X MB estimated (Y%, pressure=Z), N docs, M evicted. Add a malloc_trim(0) every Nth tick (5? 10?). Behind the same #if defined(__GLIBC__) guard as the existing trims.replayWal (the post-snapshot WAL fast-forward in startup). Cheaper because it's still bounded to the boot path, but doesn't help with mid-life arena bloat.This is the only thing operators on tight VMs still trip on — the actual leak is fixed.
If anything's unclear, the live evidence is on shadowman-zoe: dmesg -T --level=err,warn | grep -iE 'oom|killed process' shows the kill history, and systemctl status smartbotic-database shows the current RSS via Memory:. shadowman-cpp's CLAUDE.md has the Zoe SSH context (creds in /data/shadowman-cpp/.env).