plan.md 14 KB

Plan — Eviction-induced data loss, conf.d, and resilient cache management

Authored after the 2026-04-22 Zoe incident where smartbotic-database's in-memory eviction blocked a shadowman-cpp LLM streaming turn long enough that several messages row writes failed with Deadline Exceeded and never landed. The user-visible symptom was "missing assistant text" after refresh.

1. The incident, in detail

1.1 Timeline

23:18:32  MemoryStore loaded snapshot — 148,233 docs / 404 MB / cap 512 MB / threshold 80%=409 MB
23:18:32–23:23:32  Steady warm-up; +600 docs/min from normal LLM activity
23:23:52  estimatedMemoryBytes crossed 409 MB threshold
          → "Memory at 409 MB (80%), starting eviction to free 102 MB"
23:23:55  "Eviction complete: 137,310 documents evicted"  ← 92% of dataset in 3 s
23:23:55+ shadowman-server log: "Client::insert failed: Deadline Exceeded"
          repeating for ~60 s as the WAL-fallback path serves cold reads
23:35:18  Same insert deadline still tripping intermittently
~ user   "wth, now I can access (before ERR_CONNECTION_CLOSED), but
         now I see the onboarding modal"
~ user   "missing LLM messages" after refresh

1.2 What was lost vs. what was evicted

These are two different things and the distinction is the whole point of this plan.

  • Evicted (no data loss): evictDocuments() removes the doc from the in-memory index. The doc is still on disk in the WAL + most recent snapshot. A subsequent get / find falls through to the WAL-fallback path (memory_store.cpp:808 // Phase 2: Search evicted documents from WAL) and re-loads.
  • Lost (real data loss): a write that was issued during the eviction window timed out at the client's gRPC deadline (5 s) before the server could service it. The shadowman-server's messages->appendContent(...) raised, the surrounding try { } catch (...) { SPDLOG_ERROR("DB write error: ..."); } swallowed it, the WS chunk had already streamed to the WebUI, and on refresh the row was missing.

The eviction in itself was working as designed. The data loss was a second-order effect: writes were starved during the eviction window because (a) the WAL-fallback path serializes find traffic behind disk I/O, and (b) shadowman-cpp's client deadline didn't allow for a single slow read.

1.3 Root causes (current code paths)

Cause A — bulk eviction step is too aggressive. memory_store.cpp:2369 evictDocuments() selects every document needed to drop from threshold (80%) all the way to target (60%) in one pass. With a 512 MB cap, 80→60 means freeing 102 MB at once, which maps to 137k docs at the observed average size. There is no cooldown or chunking between batches.

Cause B — selection holds shared locks across the whole dataset. selectDocumentsForEviction() (line 2404) takes globalMutex_ shared and iterates every collection's documents map sorting by lastAccessedAt. For 148k docs this is a multi-second sort on the critical path of the evictionLoop. Concurrent reads work, but any write that needs globalMutex_ exclusively (e.g. createCollection, collection metadata changes, view writes) blocks for the duration.

Cause C — WAL-fallback path is slow and serializes. After the eviction, find and get for any of the 137k evicted IDs trips getEvictedDocumentIds() (line 2603) which scans the collection's eviction stubs. Each fallback then reads from WAL/disk. The per-collection coll->mutex is held during the WAL scan. Multiple concurrent finds against the same collection serialize. With shadowman-server polling messages + the agent loop reading history during a turn, this saturates instantly.

Cause D — eviction policy is too coarse. The policy is "sort by lastAccessedAt across all eligible docs, take N from the cold end." It's a true LRU, but:

  • It treats hot-write collections (messages, agent_task_messages) the same as cold archive collections (metrics_events after rollup).
  • A doc written 2 seconds ago can be evicted if its lastAccessedAt hasn't been touched since (e.g. fire-and-forget messages — no read between insert and eviction).
  • Vector-bearing docs (document_chunks, kb_chunks) cost ~16 KB per row vs ~1 KB for a plain doc. They're proportionally over- represented in the "evict to free X bytes" target, taking a disproportionate share of the bulk drop.

Cause E — shadowman-cpp's 5 s gRPC deadline doesn't allow for a cold read. That's a consumer fix (separate plan), but worth noting: even a perfect eviction policy will occasionally need to page in a cold doc, and a 5 s ceiling is too tight when the data path is on spinning rust or under contention. The DB shouldn't assume callers have generous deadlines, but it should also not require them.

2. Proposals

2.1 [Must] Drop-in config — /etc/smartbotic-database/conf.d/*.json

Today the database reads exactly one file: /etc/smartbotic-database/ config.json. Consumer projects (shadowman-cpp, callerai-storage's historical successor, anything else multi-tenant) can't tune the DB without either editing the upstream config (which dpkg --configure prompts about as a conffile change) or running their own postinst sed-edit (which I had to do during the incident — fragile).

Proposal:

  • Database reads /etc/smartbotic-database/config.json first.
  • Then iterates /etc/smartbotic-database/conf.d/*.json in lexicographic order, deep-merging each over the running config.
  • Each drop-in is owned by the consumer that ships it.
    • /etc/smartbotic-database/conf.d/00-defaults.json — optional; smartbotic-database project's conservative defaults, separate from config.json so users can edit config.json without losing them.
    • /etc/smartbotic-database/conf.d/50-shadowman.json — shipped by shadowman-cpp's deb postinst with shadowman-tuned values (memory cap 2 GB, eviction tuned for the workload, etc.).
    • /etc/smartbotic-database/conf.d/99-local.json — optional; operator overrides anything above.
  • Numeric prefix gives deterministic merge order. Last write wins.
  • On startup, the database logs which files it merged and the final effective config — easier than cat+mental-merge to debug what's actually in effect.

Implementation sketch in service/src/config/config_loader.cpp (create if needed): nlohmann::json deep-merge, single new function Config::loadFromDirectory(base_path). Existing Config struct parsing stays unchanged.

Acceptance: shadowman-cpp's deb postinst can drop a file under conf.d/ with {"storage":{"memory":{"max_memory_mb":2048}}} and the database picks it up on next start with no manual ops.

2.2 [Must] Chunked, throttled eviction with hot-write protection

The eviction loop's "free everything down to target in one pass" is the proximate cause of the incident. Replace with a steady, low-rate background drip:

  • Chunk size. Evict at most evictionChunkSize documents per pass (default 1000). After a chunk, sleep for evictionChunkPauseMs (default 50 ms) before the next.
  • Multi-pass to target. Keep evicting chunks until either the target is reached or maxEvictionPassesPerTrigger is hit (default 20 — protects against runaway loops).
  • Hot-write floor. Documents whose updatedAt is within the last hotWriteFloorMs (default 30 000 ms) are unevictable. This protects in-flight LLM streaming turns where messages get upserted-and-then-not-touched-for-seconds, exactly the case that caused the data loss.
  • Per-collection budget. No single collection contributes more than (its share of total bytes) × 1.5 to a single chunk. Stops one big collection (e.g. metrics_events) from monopolising the drop and starving everything else.

Config additions (under storage.memory):

{
  "eviction_chunk_size":            1000,
  "eviction_chunk_pause_ms":          50,
  "max_eviction_passes_per_trigger":  20,
  "hot_write_floor_ms":            30000
}

This alone would have made the incident a non-event: 137k → 1k×N chunks with 50 ms pauses, total ~7 s of background work spread out, no individual collection-mutex hold longer than it takes to drop 1000 docs (~ms).

2.3 [Should] Soft cap + admission control instead of hard target

Today the model is "stay under maxMemoryBytes; when we cross 80%, evict to 60%". Two problems:

  • The 80→60 gap is wide enough to be disruptive.
  • The DB never refuses or back-pressures; it always accepts the next write, even if accepting it means tomorrow's eviction will be huge.

Proposal — three thresholds:

  • soft (default 70%): start of trickle eviction (1 chunk per loop tick, no rush).
  • hard (default 85%): aggressive eviction (multiple chunks per tick) + emit a MEMORY_PRESSURE_HIGH event subscribers can observe.
  • emergency (default 95%): refuse new writes with RESOURCE_ EXHAUSTED until below hard. Caller (shadowman) can retry.

The RESOURCE_EXHAUSTED path is what protects against the OOM the eviction was originally added for. Same gRPC status code the existing gRPC ResourceQuota path uses, so the shadowman-cpp client already knows how to surface it.

2.4 [Should] WAL-fallback async + cached

When find falls through to getEvictedDocumentIds + WAL scan, the collection mutex is held for the duration of the disk read. That's the contention amplifier in the incident.

Proposal:

  • WAL-fallback runs without holding the per-collection write lock (separate "evicted-stubs" map with its own RWLock).
  • Hot-load evicted docs back into the live documents map on first re-access, marking them as recent → unevictable for the next pass. Same effect as Postgres's "buffer pin" — a doc you just paged in shouldn't be the next one to go.
  • Optional: an LRU "page-in cache" at the WAL layer (configurable size) so multi-find scans of the same evicted doc don't re-read the WAL N times.

2.5 [Could] Per-collection memory budgets ("pinned" extended)

The current config has pinned: true/false per collection — pinned collections never evict (memory_store.cpp:2414). Useful but binary.

Generalise to per-collection budgets:

  • pinned: true keeps the doc-always-resident behaviour.
  • memory_priority: "high" | "normal" | "low" weights eviction selection. messages is high; metrics_events is low. The selection sort uses priority × access_recency instead of pure recency.

This is a downstream consumer concern (shadowman-cpp picks the priorities for its collections), so wire-compatibility-wise it lives in CollectionOptions. Existing collections default to normal → no behaviour change for callers that don't opt in.

2.6 [Could] Better observability for cache pressure

Today the operator sees one info log per minute:

Memory check: 311 MB estimated, 409 MB threshold, 15043 docs, 137310 evicted

That's lossy. We see "evicted" went from 0 to 137k between two minutely samples, but we don't see the per-collection breakdown, the hot/cold split, or any rate.

Add:

  • A /metrics endpoint or gRPC GetMemoryStats RPC (the existing metrics_events collection isn't suitable — it's downstream data, not infra metrics). Returns per-collection doc count, est. bytes, evicted-stub count, last-eviction-event timestamp + size.
  • Subscribers via SubscribeEvents get a MEMORY_PRESSURE_HIGH event when crossing soft→hard, and MEMORY_EVICTION_BURST after a trigger that touched > N docs (configurable, default 10k). A shadowman-side metrics rollup can then show "eviction storms" alongside its other graphs.

2.7 [Could] Make evictDocument honour an outstanding-write quiesce

Tiny win, low priority: before evicting a doc, check if there's an outstanding update / patch in flight against it. If yes, skip this round. Avoids a tiny race where eviction lands between an upsert's "find existing" and "write new version" phases.

3. Recommended sequencing

  1. conf.d — landable in a day. Makes 2.1 / 2.2 / 2.3 tunable from shadowman-cpp's deb without touching the DB project for every tweak. Ship in the next smartbotic-database minor.
  2. Chunked eviction (2.2) — the actual fix for the incident. Roll into the same minor as conf.d so consumers can land both together.
  3. Soft cap + admission control (2.3) — separate minor. Needs shadowman-cpp client side to handle the RESOURCE_EXHAUSTED retry path before we turn the strict admission on by default.
  4. WAL-fallback unlocking (2.4) — separate minor; involves touching the WAL layer's locking model. Higher risk, deserves its own release.
  5. Per-collection budgets, observability — when needed.

4. What this plan does NOT cover

  • Memory accounting accuracy. estimateDocumentSize is an approximation; actual heap footprint can drift. Not the cause of this incident, but a separate cleanup.
  • Snapshot loader behaviour. The eviction triggered after warm-up from snapshot, not during it. No changes proposed there.
  • Replication. Replication lag during eviction wasn't relevant to this incident (single-node deployment); revisit if multi-node is used in production.

5. Companion changes in shadowman-cpp (out of scope here, but

referenced)

  • Raise the gRPC client deadline for write RPCs from 5 s → 15 s. Read RPCs stay short.
  • Wrap messages->save / appendContent / appendThinking in a 3-attempt retry-with-backoff specifically on Deadline Exceeded and RESOURCE_EXHAUSTED. Writes are upsert-by-id, idempotent.
  • Don't swallow the DB-write exception silently in the agent-loop callback — at least log the (turn, message-id, delta-size) so a missing message is correlatable to a DB error in the journal.
  • Drop a conf.d/50-shadowman.json from the shadowman-cpp deb postinst with the workload-tuned memory cap (default 2 GB) + thresholds.

These are tracked in shadowman-cpp's repo, not here.