# Plan — Eviction-induced data loss, conf.d, and resilient cache management > Authored after the 2026-04-22 Zoe incident where smartbotic-database's > in-memory eviction blocked a shadowman-cpp LLM streaming turn long > enough that several `messages` row writes failed with > `Deadline Exceeded` and never landed. The user-visible symptom was > "missing assistant text" after refresh. ## 1. The incident, in detail ### 1.1 Timeline ``` 23:18:32 MemoryStore loaded snapshot — 148,233 docs / 404 MB / cap 512 MB / threshold 80%=409 MB 23:18:32–23:23:32 Steady warm-up; +600 docs/min from normal LLM activity 23:23:52 estimatedMemoryBytes crossed 409 MB threshold → "Memory at 409 MB (80%), starting eviction to free 102 MB" 23:23:55 "Eviction complete: 137,310 documents evicted" ← 92% of dataset in 3 s 23:23:55+ shadowman-server log: "Client::insert failed: Deadline Exceeded" repeating for ~60 s as the WAL-fallback path serves cold reads 23:35:18 Same insert deadline still tripping intermittently ~ user "wth, now I can access (before ERR_CONNECTION_CLOSED), but now I see the onboarding modal" ~ user "missing LLM messages" after refresh ``` ### 1.2 What was lost vs. what was evicted These are two different things and the distinction is the whole point of this plan. - **Evicted** (no data loss): `evictDocuments()` removes the doc from the in-memory index. The doc is still on disk in the WAL + most recent snapshot. A subsequent `get` / `find` falls through to the WAL-fallback path (`memory_store.cpp:808 // Phase 2: Search evicted documents from WAL`) and re-loads. - **Lost** (real data loss): a write that was issued during the eviction window timed out at the client's gRPC deadline (5 s) before the server could service it. The shadowman-server's `messages->appendContent(...)` raised, the surrounding `try { } catch (...) { SPDLOG_ERROR("DB write error: ..."); }` swallowed it, the WS chunk had already streamed to the WebUI, and on refresh the row was missing. The eviction in itself was working as designed. The data loss was a **second-order effect**: writes were starved during the eviction window because (a) the WAL-fallback path serializes find traffic behind disk I/O, and (b) shadowman-cpp's client deadline didn't allow for a single slow read. ### 1.3 Root causes (current code paths) **Cause A — bulk eviction step is too aggressive.** `memory_store.cpp:2369 evictDocuments()` selects every document needed to drop from `threshold (80%)` all the way to `target (60%)` in one pass. With a 512 MB cap, 80→60 means freeing 102 MB at once, which maps to 137k docs at the observed average size. There is no cooldown or chunking between batches. **Cause B — selection holds shared locks across the whole dataset.** `selectDocumentsForEviction()` (line 2404) takes `globalMutex_` shared and iterates every collection's `documents` map sorting by `lastAccessedAt`. For 148k docs this is a multi-second sort on the critical path of the evictionLoop. Concurrent reads work, but any write that needs `globalMutex_` exclusively (e.g. `createCollection`, collection metadata changes, view writes) blocks for the duration. **Cause C — WAL-fallback path is slow and serializes.** After the eviction, `find` and `get` for any of the 137k evicted IDs trips `getEvictedDocumentIds()` (line 2603) which scans the collection's eviction stubs. Each fallback then reads from WAL/disk. The per-collection `coll->mutex` is held during the WAL scan. Multiple concurrent finds against the same collection serialize. With shadowman-server polling messages + the agent loop reading history during a turn, this saturates instantly. **Cause D — eviction policy is too coarse.** The policy is "sort by lastAccessedAt across all eligible docs, take N from the cold end." It's a true LRU, but: - It treats hot-write collections (`messages`, `agent_task_messages`) the same as cold archive collections (`metrics_events` after rollup). - A doc written 2 seconds ago can be evicted if its `lastAccessedAt` hasn't been touched since (e.g. fire-and-forget messages — no read between insert and eviction). - Vector-bearing docs (`document_chunks`, `kb_chunks`) cost ~16 KB per row vs ~1 KB for a plain doc. They're proportionally over- represented in the "evict to free X bytes" target, taking a disproportionate share of the bulk drop. **Cause E — shadowman-cpp's 5 s gRPC deadline doesn't allow for a cold read.** That's a consumer fix (separate plan), but worth noting: even a perfect eviction policy will occasionally need to page in a cold doc, and a 5 s ceiling is too tight when the data path is on spinning rust or under contention. The DB shouldn't assume callers have generous deadlines, but it should also not require them. ## 2. Proposals ### 2.1 [Must] Drop-in config — `/etc/smartbotic-database/conf.d/*.json` Today the database reads exactly one file: `/etc/smartbotic-database/ config.json`. Consumer projects (shadowman-cpp, callerai-storage's historical successor, anything else multi-tenant) can't tune the DB without either editing the upstream config (which `dpkg --configure` prompts about as a conffile change) or running their own postinst sed-edit (which I had to do during the incident — fragile). **Proposal:** - Database reads `/etc/smartbotic-database/config.json` first. - Then iterates `/etc/smartbotic-database/conf.d/*.json` in lexicographic order, deep-merging each over the running config. - Each drop-in is owned by the consumer that ships it. - `/etc/smartbotic-database/conf.d/00-defaults.json` — optional; smartbotic-database project's conservative defaults, separate from `config.json` so users can edit `config.json` without losing them. - `/etc/smartbotic-database/conf.d/50-shadowman.json` — shipped by shadowman-cpp's deb postinst with shadowman-tuned values (memory cap 2 GB, eviction tuned for the workload, etc.). - `/etc/smartbotic-database/conf.d/99-local.json` — optional; operator overrides anything above. - Numeric prefix gives deterministic merge order. Last write wins. - On startup, the database logs which files it merged and the final effective config — easier than `cat`+mental-merge to debug what's actually in effect. Implementation sketch in `service/src/config/config_loader.cpp` (create if needed): `nlohmann::json` deep-merge, single new function `Config::loadFromDirectory(base_path)`. Existing `Config` struct parsing stays unchanged. Acceptance: shadowman-cpp's deb postinst can drop a file under `conf.d/` with `{"storage":{"memory":{"max_memory_mb":2048}}}` and the database picks it up on next start with no manual ops. ### 2.2 [Must] Chunked, throttled eviction with hot-write protection The eviction loop's "free everything down to target in one pass" is the proximate cause of the incident. Replace with a steady, low-rate background drip: - **Chunk size.** Evict at most `evictionChunkSize` documents per pass (default 1000). After a chunk, sleep for `evictionChunkPauseMs` (default 50 ms) before the next. - **Multi-pass to target.** Keep evicting chunks until either the target is reached or `maxEvictionPassesPerTrigger` is hit (default 20 — protects against runaway loops). - **Hot-write floor.** Documents whose `updatedAt` is within the last `hotWriteFloorMs` (default 30 000 ms) are unevictable. This protects in-flight LLM streaming turns where messages get upserted-and-then-not-touched-for-seconds, exactly the case that caused the data loss. - **Per-collection budget.** No single collection contributes more than `(its share of total bytes) × 1.5` to a single chunk. Stops one big collection (e.g. `metrics_events`) from monopolising the drop and starving everything else. Config additions (under `storage.memory`): ```json { "eviction_chunk_size": 1000, "eviction_chunk_pause_ms": 50, "max_eviction_passes_per_trigger": 20, "hot_write_floor_ms": 30000 } ``` This alone would have made the incident a non-event: 137k → 1k×N chunks with 50 ms pauses, total ~7 s of background work spread out, no individual collection-mutex hold longer than it takes to drop 1000 docs (~ms). ### 2.3 [Should] Soft cap + admission control instead of hard target Today the model is "stay under maxMemoryBytes; when we cross 80%, evict to 60%". Two problems: - The 80→60 gap is wide enough to be disruptive. - The DB never refuses or back-pressures; it always accepts the next write, even if accepting it means tomorrow's eviction will be huge. Proposal — three thresholds: - **soft (default 70%)**: start of trickle eviction (1 chunk per loop tick, no rush). - **hard (default 85%)**: aggressive eviction (multiple chunks per tick) + emit a `MEMORY_PRESSURE_HIGH` event subscribers can observe. - **emergency (default 95%)**: refuse new writes with `RESOURCE_ EXHAUSTED` until below `hard`. Caller (shadowman) can retry. The `RESOURCE_EXHAUSTED` path is what protects against the OOM the eviction was originally added for. Same gRPC status code the existing `gRPC ResourceQuota` path uses, so the shadowman-cpp client already knows how to surface it. ### 2.4 [Should] WAL-fallback async + cached When `find` falls through to `getEvictedDocumentIds` + WAL scan, the collection mutex is held for the duration of the disk read. That's the contention amplifier in the incident. Proposal: - WAL-fallback runs without holding the per-collection write lock (separate "evicted-stubs" map with its own RWLock). - Hot-load evicted docs back into the live `documents` map on first re-access, marking them as recent → unevictable for the next pass. Same effect as Postgres's "buffer pin" — a doc you just paged in shouldn't be the next one to go. - Optional: an LRU "page-in cache" at the WAL layer (configurable size) so multi-find scans of the same evicted doc don't re-read the WAL N times. ### 2.5 [Could] Per-collection memory budgets ("pinned" extended) The current config has `pinned: true/false` per collection — pinned collections never evict (`memory_store.cpp:2414`). Useful but binary. Generalise to **per-collection budgets**: - `pinned: true` keeps the doc-always-resident behaviour. - `memory_priority: "high" | "normal" | "low"` weights eviction selection. `messages` is `high`; `metrics_events` is `low`. The selection sort uses `priority × access_recency` instead of pure recency. This is a downstream consumer concern (shadowman-cpp picks the priorities for its collections), so wire-compatibility-wise it lives in `CollectionOptions`. Existing collections default to `normal` → no behaviour change for callers that don't opt in. ### 2.6 [Could] Better observability for cache pressure Today the operator sees one info log per minute: ``` Memory check: 311 MB estimated, 409 MB threshold, 15043 docs, 137310 evicted ``` That's lossy. We see "evicted" went from 0 to 137k between two minutely samples, but we don't see the per-collection breakdown, the hot/cold split, or any rate. Add: - A `/metrics` endpoint or gRPC `GetMemoryStats` RPC (the existing `metrics_events` collection isn't suitable — it's downstream data, not infra metrics). Returns per-collection doc count, est. bytes, evicted-stub count, last-eviction-event timestamp + size. - Subscribers via `SubscribeEvents` get a `MEMORY_PRESSURE_HIGH` event when crossing soft→hard, and `MEMORY_EVICTION_BURST` after a trigger that touched > N docs (configurable, default 10k). A shadowman-side metrics rollup can then show "eviction storms" alongside its other graphs. ### 2.7 [Could] Make `evictDocument` honour an outstanding-write quiesce Tiny win, low priority: before evicting a doc, check if there's an outstanding `update` / `patch` in flight against it. If yes, skip this round. Avoids a tiny race where eviction lands between an upsert's "find existing" and "write new version" phases. ## 3. Recommended sequencing 1. **conf.d** — landable in a day. Makes 2.1 / 2.2 / 2.3 tunable from shadowman-cpp's deb without touching the DB project for every tweak. Ship in the next smartbotic-database minor. 2. **Chunked eviction (2.2)** — the actual fix for the incident. Roll into the same minor as conf.d so consumers can land both together. 3. **Soft cap + admission control (2.3)** — separate minor. Needs shadowman-cpp client side to handle the `RESOURCE_EXHAUSTED` retry path before we turn the strict admission on by default. 4. **WAL-fallback unlocking (2.4)** — separate minor; involves touching the WAL layer's locking model. Higher risk, deserves its own release. 5. **Per-collection budgets, observability** — when needed. ## 4. What this plan does NOT cover - Memory accounting accuracy. `estimateDocumentSize` is an approximation; actual heap footprint can drift. Not the cause of this incident, but a separate cleanup. - Snapshot loader behaviour. The eviction triggered after warm-up from snapshot, not during it. No changes proposed there. - Replication. Replication lag during eviction wasn't relevant to this incident (single-node deployment); revisit if multi-node is used in production. ## 5. Companion changes in shadowman-cpp (out of scope here, but referenced) - Raise the gRPC client deadline for **write** RPCs from 5 s → 15 s. Read RPCs stay short. - Wrap `messages->save / appendContent / appendThinking` in a 3-attempt retry-with-backoff specifically on `Deadline Exceeded` and `RESOURCE_EXHAUSTED`. Writes are upsert-by-id, idempotent. - Don't swallow the DB-write exception silently in the agent-loop callback — at least log the (turn, message-id, delta-size) so a missing message is correlatable to a DB error in the journal. - Drop a `conf.d/50-shadowman.json` from the shadowman-cpp deb postinst with the workload-tuned memory cap (default 2 GB) + thresholds. These are tracked in shadowman-cpp's repo, not here.