Authored after the 2026-04-22 Zoe incident where smartbotic-database's in-memory eviction blocked a shadowman-cpp LLM streaming turn long enough that several
messagesrow writes failed withDeadline Exceededand never landed. The user-visible symptom was "missing assistant text" after refresh.
23:18:32 MemoryStore loaded snapshot — 148,233 docs / 404 MB / cap 512 MB / threshold 80%=409 MB
23:18:32–23:23:32 Steady warm-up; +600 docs/min from normal LLM activity
23:23:52 estimatedMemoryBytes crossed 409 MB threshold
→ "Memory at 409 MB (80%), starting eviction to free 102 MB"
23:23:55 "Eviction complete: 137,310 documents evicted" ← 92% of dataset in 3 s
23:23:55+ shadowman-server log: "Client::insert failed: Deadline Exceeded"
repeating for ~60 s as the WAL-fallback path serves cold reads
23:35:18 Same insert deadline still tripping intermittently
~ user "wth, now I can access (before ERR_CONNECTION_CLOSED), but
now I see the onboarding modal"
~ user "missing LLM messages" after refresh
These are two different things and the distinction is the whole point of this plan.
evictDocuments() removes the doc from
the in-memory index. The doc is still on disk in the WAL + most
recent snapshot. A subsequent get / find falls through to the
WAL-fallback path (memory_store.cpp:808 // Phase 2: Search evicted
documents from WAL) and re-loads.messages->appendContent(...) raised, the surrounding try { } catch
(...) { SPDLOG_ERROR("DB write error: ..."); } swallowed it, the
WS chunk had already streamed to the WebUI, and on refresh the row
was missing.The eviction in itself was working as designed. The data loss was a second-order effect: writes were starved during the eviction window because (a) the WAL-fallback path serializes find traffic behind disk I/O, and (b) shadowman-cpp's client deadline didn't allow for a single slow read.
Cause A — bulk eviction step is too aggressive.
memory_store.cpp:2369 evictDocuments() selects every document needed
to drop from threshold (80%) all the way to target (60%) in one
pass. With a 512 MB cap, 80→60 means freeing 102 MB at once, which
maps to 137k docs at the observed average size. There is no cooldown
or chunking between batches.
Cause B — selection holds shared locks across the whole dataset.
selectDocumentsForEviction() (line 2404) takes globalMutex_
shared and iterates every collection's documents map sorting by
lastAccessedAt. For 148k docs this is a multi-second sort on the
critical path of the evictionLoop. Concurrent reads work, but any
write that needs globalMutex_ exclusively (e.g. createCollection,
collection metadata changes, view writes) blocks for the duration.
Cause C — WAL-fallback path is slow and serializes.
After the eviction, find and get for any of the 137k evicted IDs
trips getEvictedDocumentIds() (line 2603) which scans the
collection's eviction stubs. Each fallback then reads from WAL/disk.
The per-collection coll->mutex is held during the WAL scan. Multiple
concurrent finds against the same collection serialize. With
shadowman-server polling messages + the agent loop reading history
during a turn, this saturates instantly.
Cause D — eviction policy is too coarse. The policy is "sort by lastAccessedAt across all eligible docs, take N from the cold end." It's a true LRU, but:
messages, agent_task_messages)
the same as cold archive collections (metrics_events after
rollup).lastAccessedAt
hasn't been touched since (e.g. fire-and-forget messages — no read
between insert and eviction).document_chunks, kb_chunks) cost ~16 KB
per row vs ~1 KB for a plain doc. They're proportionally over-
represented in the "evict to free X bytes" target, taking a
disproportionate share of the bulk drop.Cause E — shadowman-cpp's 5 s gRPC deadline doesn't allow for a cold read. That's a consumer fix (separate plan), but worth noting: even a perfect eviction policy will occasionally need to page in a cold doc, and a 5 s ceiling is too tight when the data path is on spinning rust or under contention. The DB shouldn't assume callers have generous deadlines, but it should also not require them.
/etc/smartbotic-database/conf.d/*.jsonToday the database reads exactly one file: /etc/smartbotic-database/
config.json. Consumer projects (shadowman-cpp, callerai-storage's
historical successor, anything else multi-tenant) can't tune the DB
without either editing the upstream config (which dpkg --configure
prompts about as a conffile change) or running their own postinst
sed-edit (which I had to do during the incident — fragile).
Proposal:
/etc/smartbotic-database/config.json first./etc/smartbotic-database/conf.d/*.json in
lexicographic order, deep-merging each over the running config./etc/smartbotic-database/conf.d/00-defaults.json — optional;
smartbotic-database project's conservative defaults, separate
from config.json so users can edit config.json without
losing them./etc/smartbotic-database/conf.d/50-shadowman.json — shipped
by shadowman-cpp's deb postinst with shadowman-tuned values
(memory cap 2 GB, eviction tuned for the workload, etc.)./etc/smartbotic-database/conf.d/99-local.json — optional;
operator overrides anything above.cat+mental-merge to debug what's
actually in effect.Implementation sketch in service/src/config/config_loader.cpp
(create if needed): nlohmann::json deep-merge, single new function
Config::loadFromDirectory(base_path). Existing Config struct
parsing stays unchanged.
Acceptance: shadowman-cpp's deb postinst can drop a file under
conf.d/ with {"storage":{"memory":{"max_memory_mb":2048}}} and
the database picks it up on next start with no manual ops.
The eviction loop's "free everything down to target in one pass" is the proximate cause of the incident. Replace with a steady, low-rate background drip:
evictionChunkSize documents per
pass (default 1000). After a chunk, sleep for
evictionChunkPauseMs (default 50 ms) before the next.maxEvictionPassesPerTrigger is hit (default
20 — protects against runaway loops).updatedAt is within the
last hotWriteFloorMs (default 30 000 ms) are unevictable. This
protects in-flight LLM streaming turns where messages get
upserted-and-then-not-touched-for-seconds, exactly the case that
caused the data loss.(its share of total bytes) × 1.5 to a single chunk. Stops
one big collection (e.g. metrics_events) from monopolising the
drop and starving everything else.Config additions (under storage.memory):
{
"eviction_chunk_size": 1000,
"eviction_chunk_pause_ms": 50,
"max_eviction_passes_per_trigger": 20,
"hot_write_floor_ms": 30000
}
This alone would have made the incident a non-event: 137k → 1k×N chunks with 50 ms pauses, total ~7 s of background work spread out, no individual collection-mutex hold longer than it takes to drop 1000 docs (~ms).
Today the model is "stay under maxMemoryBytes; when we cross 80%, evict to 60%". Two problems:
Proposal — three thresholds:
MEMORY_PRESSURE_HIGH event subscribers can
observe.RESOURCE_
EXHAUSTED until below hard. Caller (shadowman) can retry.The RESOURCE_EXHAUSTED path is what protects against the OOM the
eviction was originally added for. Same gRPC status code the existing
gRPC ResourceQuota path uses, so the shadowman-cpp client already
knows how to surface it.
When find falls through to getEvictedDocumentIds + WAL scan, the
collection mutex is held for the duration of the disk read. That's
the contention amplifier in the incident.
Proposal:
documents map on first
re-access, marking them as recent → unevictable for the next pass.
Same effect as Postgres's "buffer pin" — a doc you just paged in
shouldn't be the next one to go.The current config has pinned: true/false per collection — pinned
collections never evict (memory_store.cpp:2414). Useful but binary.
Generalise to per-collection budgets:
pinned: true keeps the doc-always-resident behaviour.memory_priority: "high" | "normal" | "low" weights eviction
selection. messages is high; metrics_events is low. The
selection sort uses priority × access_recency instead of pure
recency.This is a downstream consumer concern (shadowman-cpp picks the
priorities for its collections), so wire-compatibility-wise it lives
in CollectionOptions. Existing collections default to normal →
no behaviour change for callers that don't opt in.
Today the operator sees one info log per minute:
Memory check: 311 MB estimated, 409 MB threshold, 15043 docs, 137310 evicted
That's lossy. We see "evicted" went from 0 to 137k between two minutely samples, but we don't see the per-collection breakdown, the hot/cold split, or any rate.
Add:
/metrics endpoint or gRPC GetMemoryStats RPC (the existing
metrics_events collection isn't suitable — it's downstream data,
not infra metrics). Returns per-collection doc count, est. bytes,
evicted-stub count, last-eviction-event timestamp + size.SubscribeEvents get a MEMORY_PRESSURE_HIGH event
when crossing soft→hard, and MEMORY_EVICTION_BURST after a
trigger that touched > N docs (configurable, default 10k). A
shadowman-side metrics rollup can then show "eviction storms"
alongside its other graphs.evictDocument honour an outstanding-write quiesceTiny win, low priority: before evicting a doc, check if there's an
outstanding update / patch in flight against it. If yes, skip
this round. Avoids a tiny race where eviction lands between an
upsert's "find existing" and "write new version" phases.
RESOURCE_EXHAUSTED
retry path before we turn the strict admission on by default.estimateDocumentSize is an
approximation; actual heap footprint can drift. Not the cause of
this incident, but a separate cleanup.referenced)
messages->save / appendContent / appendThinking in a
3-attempt retry-with-backoff specifically on Deadline Exceeded
and RESOURCE_EXHAUSTED. Writes are upsert-by-id, idempotent.conf.d/50-shadowman.json from the shadowman-cpp deb
postinst with the workload-tuned memory cap (default 2 GB) +
thresholds.These are tracked in shadowman-cpp's repo, not here.