Browse Source

release(v2.4.3): LMDB handle lifetime + eviction drain cap + live assertions

Three fixes from an incident where a production instance appeared to
lose all data. Nothing was deleted; two faults combined to make the
serving path return 0 docs for every collection.

1. LMDB dbi handle lifetime.

try_open_for_read() called mdb_dbi_open inside a ReadTxn and cached the
handle. LMDB keeps a handle private to the opening transaction until it
COMMITS, and closes it if that transaction ABORTS - and ReadTxn always
aborts. The cached MDB_dbi then dangled, so every later mdb_put and
mdb_cursor_open on that collection failed EINVAL for the life of the
process. It hit whichever collection was READ before it was WRITTEN
after a restart; peers were fine because open_for_write commits, which
promotes the handle into the env's shared table.

The read path no longer caches. The handle is valid inside its own txn,
which is all the caller needs, and writes still populate the cache.

That failed mirror write flipped mirror_healthy_=false, which sent every
read to MemoryStore - where fault 2 was waiting.

2. Eviction drained the store.

Eviction assumes estimatedMemoryBytes_ falls as docs leave. It did not
(347 -> 358 MB while evicting 3000 docs; RSS peaked at 1.9 GB against a
512 MB budget), so current <= targetBytes was never satisfied and it
trimmed one chunk per tick for 30 minutes until 5996 of ~5990 docs were
gone.

Two guards, both ERROR because hitting them means the estimate is wrong
rather than the store being too large:
  - eviction_max_episode_percent (default 50) caps what one pressure
    episode may evict, enforced PER CHUNK. A Hard/Emergency tick runs up
    to 20 chunks, enough to empty a store between tick-level checks.
  - an episode re-arms only after 3 consecutive Normal ticks, because
    eviction transiently dips the estimate and a per-tick reset drains
    the store in repeated 50% stages.
  - a no-progress detector pauses eviction after 3 consecutive evicting
    ticks that fail to lower the estimate.

3. 104 assertions were no-ops.

CMAKE_BUILD_TYPE=Release sets -DNDEBUG, compiling out every assert() in
test_views (24), test_snapshot_durability (30), test_timestamp_precision
(26), test_config_dropins (14) and test_eviction (10). Those suites ran,
printed PASS and verified nothing - part of why the v2.3 views bug lived
for two releases. Test targets now build with -UNDEBUG. Restoring the
assertions surfaced no further failures.

Tests: test_read_first_collection_survives_repeated_access reproduces
the EINVAL (the ENV must be reopened, not just the store - handles live
in the env's shared table, so reusing one LmdbEnv hides the bug).
test_eviction_never_drains_the_store holds the cap at 200/400.
ctest 14/14 with live assertions, both e2e suites green.
fszontagh 1 month ago
parent
commit
da39f2835c

File diff suppressed because it is too large
+ 0 - 0
CLAUDE.md


+ 1 - 1
VERSION

@@ -1 +1 @@
-2.4.2
+2.4.3

+ 30 - 0
docs/integration-guide.md

@@ -883,6 +883,7 @@ The server configuration file is at `/etc/smartbotic-database/config.json`.
 | `rpc_port` | `9004` | gRPC listen port |
 | `data_directory` | `/var/lib/smartbotic-database` | WAL, snapshots, files stored here |
 | `memory.max_memory_mb` | `512` | Maximum memory for document cache. Evicts LRU when exceeded |
+| `memory.eviction_max_episode_percent` | `50` | (2.4.3+) Ceiling on the share of the resident set one pressure episode may evict. `0` disables. See the drain-cap note below |
 | `persistence.snapshot_interval_sec` | `3600` | Full snapshot every N seconds (WAL truncated after) |
 | `encryption.enabled` | `true` | Field-level AES-256-GCM encryption. Key auto-generated on first start |
 | `migrations.enabled` | `false` | Enable to auto-apply JSON migrations on startup |
@@ -990,6 +991,35 @@ The server tracks four pressure levels based on current usage vs. `max_memory_mb
 
 Query current state via `smartbotic-db-cli status` or the `GetMemoryStats` RPC.
 
+#### Drain cap (2.4.3+)
+
+Eviction assumes the memory estimate falls as documents leave. When memory is
+held by something eviction cannot free, that assumption breaks: the target is
+never reachable, and eviction keeps trimming until the cache is empty. On
+2026-08-03 a production instance did exactly this — 5996 of ~5990 documents
+evicted over 30 minutes while the estimate *rose* (347 → 358 MB). Every
+collection then reported 0 docs, which looks indistinguishable from data loss.
+No data was actually lost; it was all still in LMDB.
+
+Two guards now bound this:
+
+- **Episode cap** — one pressure episode may evict at most
+  `memory.eviction_max_episode_percent` (default 50%) of the resident set it
+  started with, enforced per chunk. An episode re-arms only after pressure has
+  been `normal` for 3 consecutive ticks.
+- **No-progress detector** — after 3 consecutive ticks that evict documents
+  without lowering the estimate, eviction pauses until pressure clears.
+
+Both log at **ERROR**. Hitting either means the memory estimate is wrong, not
+that the store is oversized — treat it as an alert, raise
+`storage.memory.max_memory_mb`, and check what is holding memory. Set
+`eviction_max_episode_percent: 0` to disable the cap (not recommended).
+
+> **Note for v2.0+.** MemoryStore is a cache in front of LMDB; the durable copy
+> lives in `<dataDir>/projects/<name>/env/`. Eviction draining the cache only
+> becomes visible if the LMDB mirror is also unhealthy, which is what made the
+> above incident look total. Watch for `v2.0 mirror failed` in the log.
+
 #### Chunked eviction
 
 Eviction never stops the world. Each tick (default every `eviction_check_interval_ms` = 5000ms, but reactive to pressure):

+ 4 - 0
service/src/database_service.cpp

@@ -645,6 +645,9 @@ DatabaseService::Config DatabaseService::parseConfig(const nlohmann::json& json)
 
             // v1.7.0 T10 — eviction burst event threshold
             config.evictionBurstThreshold = memory.value("eviction_burst_threshold", config.evictionBurstThreshold);
+
+            // v2.4.3 — eviction drain cap (0 disables)
+            config.evictionMaxEpisodePercent = memory.value("eviction_max_episode_percent", config.evictionMaxEpisodePercent);
         }
 
         // Persistence settings
@@ -804,6 +807,7 @@ void DatabaseService::setupComponents() {
     storeConfig.memoryHardPercent = config_.memoryHardPercent;
     storeConfig.memoryEmergencyPercent = config_.memoryEmergencyPercent;
     storeConfig.evictionBurstThreshold = config_.evictionBurstThreshold;
+    storeConfig.evictionMaxEpisodePercent = config_.evictionMaxEpisodePercent;
     store_ = std::make_unique<MemoryStore>(storeConfig);
 
     // Create view manager (cache loaded in initialize() after persistence recovery)

+ 4 - 0
service/src/database_service.hpp

@@ -122,6 +122,10 @@ public:
         // v1.7.0 T10 — eviction burst event threshold (docs per tick).
         uint32_t evictionBurstThreshold = 10000;
 
+        // v2.4.3 — cap on the share of the resident set one pressure episode
+        // may evict. 0 disables. See MemoryStore::Config for the rationale.
+        uint32_t evictionMaxEpisodePercent = 50;
+
         // Persistence settings
         uint32_t walSyncIntervalMs = 100;
         uint32_t snapshotIntervalSec = 3600;

+ 103 - 0
service/src/memory_store.cpp

@@ -1,3 +1,4 @@
+#include <limits>
 #include "memory_store.hpp"
 
 #include "config/collection_config_manager.hpp"
@@ -2536,8 +2537,74 @@ void MemoryStore::evictionLoop() {
         lastObservedPressure_.store(p, std::memory_order_relaxed);
 
         if (p == MemoryPressure::Normal) {
+            // Re-arm only after pressure has been Normal for a few ticks in a
+            // row. Eviction transiently pushes the estimate under the
+            // threshold, so re-arming on the first Normal tick lets the drain
+            // cap restart repeatedly and empty the store in 50% stages.
+            if (++evictionNormalTicks_ >= kEvictionRearmNormalTicks) {
+                evictionEpisodeStartDocs_ = 0;
+                evictionEpisodeEvicted_ = 0;
+                evictionDrainCapLogged_ = false;
+                evictionNoProgressTicks_ = 0;
+            }
             continue;  // nothing to do
         }
+        evictionNormalTicks_ = 0;
+
+        // v2.4.3 — start of a pressure episode: remember how big the resident
+        // set was, so the drain cap below is measured against it.
+        if (evictionEpisodeStartDocs_ == 0) {
+            evictionEpisodeStartDocs_ = getStats().totalDocuments;
+            evictionEpisodeEvicted_ = 0;
+            evictionDrainCapLogged_ = false;
+            evictionNoProgressTicks_ = 0;
+        }
+
+        // No-progress detector — the production signature. Eviction assumes
+        // estimatedMemoryBytes_ falls as docs leave. When memory is held by
+        // something eviction cannot free, the estimate stays flat (it rose,
+        // 347 -> 358 MB, in the incident) and the target is never reachable.
+        // Evicting harder cannot help, so stop rather than drain.
+        if (evictionNoProgressTicks_ >= kEvictionNoProgressLimit) {
+            if (!evictionDrainCapLogged_) {
+                evictionDrainCapLogged_ = true;
+                spdlog::error(
+                    "Eviction made no headway for {} consecutive ticks (pressure={}, "
+                    "estimate={} bytes): evicting documents is not lowering the memory "
+                    "estimate, so the estimate is not tracking what eviction can free. "
+                    "Eviction is PAUSED until pressure returns to normal; raise "
+                    "storage.memory.max_memory_mb or investigate the estimator.",
+                    evictionNoProgressTicks_, memoryPressureToString(p),
+                    estimatedMemoryBytes_.load(std::memory_order_relaxed));
+            }
+            continue;
+        }
+
+        // Drain cap: refuse to evict more than evictionMaxEpisodePercent of
+        // the set this episode started with. Reaching it means evicting is not
+        // relieving pressure, i.e. estimatedMemoryBytes_ is not tracking what
+        // eviction can actually free. Draining the rest of the store will not
+        // help and costs every read a fallback.
+        if (config_.evictionMaxEpisodePercent > 0 && evictionEpisodeStartDocs_ > 0) {
+            const uint64_t cap = evictionEpisodeStartDocs_
+                               * config_.evictionMaxEpisodePercent / 100;
+            if (evictionEpisodeEvicted_ >= cap) {
+                if (!evictionDrainCapLogged_) {
+                    evictionDrainCapLogged_ = true;
+                    spdlog::error(
+                        "Eviction drain cap hit: evicted {} of {} docs this episode "
+                        "({}% cap) and pressure is still {}. The memory estimate is "
+                        "not falling as docs are evicted — suspecting the estimate is "
+                        "wrong rather than the store being too large. Eviction is "
+                        "PAUSED until pressure returns to normal; investigate "
+                        "estimatedMemoryBytes_ / raise storage.memory.max_memory_mb.",
+                        evictionEpisodeEvicted_, evictionEpisodeStartDocs_,
+                        config_.evictionMaxEpisodePercent,
+                        memoryPressureToString(p));
+                }
+                continue;  // stay paused for the rest of the episode
+            }
+        }
 
         // Pick how aggressive we're willing to be this tick. SOFT trickles
         // one chunk at a time so steady-state background noise doesn't
@@ -2555,17 +2622,39 @@ void MemoryStore::evictionLoop() {
         uint64_t targetBytes = static_cast<uint64_t>(config_.maxMemoryBytes)
                              * config_.evictionTargetPercent / 100;
 
+        const uint64_t estimateBeforeTick =
+            estimatedMemoryBytes_.load(std::memory_order_relaxed);
+
         uint64_t totalFreed = 0;
         uint64_t totalDocs = 0;
         uint32_t passesRun = 0;
+        const uint64_t episodeCap =
+            (config_.evictionMaxEpisodePercent > 0 && evictionEpisodeStartDocs_ > 0)
+                ? evictionEpisodeStartDocs_ * config_.evictionMaxEpisodePercent / 100
+                : std::numeric_limits<uint64_t>::max();
+
         for (uint32_t pass = 0; pass < maxPassesThisTick; ++pass) {
             if (!running_.load(std::memory_order_acquire)) break;
 
+            // The cap has to be enforced per PASS, not just per tick: a Hard or
+            // Emergency tick runs up to maxEvictionPassesPerTrigger chunks, which
+            // is enough to empty the store before the next tick's check.
+            if (evictionEpisodeEvicted_ >= episodeCap) break;
+
             uint64_t current = estimatedMemoryBytes_.load(std::memory_order_relaxed);
             if (current <= targetBytes) break;
 
             uint64_t freed = evictOneChunk();
             passesRun++;
+            if (freed > 0) {
+                // evictOneChunk() reports bytes; the doc count lands in stats_.
+                // Only read it when this chunk actually evicted something —
+                // stats_.lastEvictionDocs is left at its previous value on a
+                // no-op chunk, so reading it unconditionally double-counts.
+                std::lock_guard<std::mutex> statsLock(statsMutex_);
+                evictionEpisodeEvicted_ += stats_.lastEvictionDocs;
+                totalDocs += stats_.lastEvictionDocs;
+            }
             if (freed == 0) {
                 // Nothing evictable (all pinned / hot-write / empty). Don't
                 // spin — wait for the next tick.
@@ -2582,6 +2671,20 @@ void MemoryStore::evictionLoop() {
             }
         }
 
+        // Did this tick's eviction actually lower the estimate? Concurrent
+        // writes can mask a real drop, hence "consecutive ticks" rather than
+        // reacting to a single one.
+        if (totalDocs > 0) {
+            const uint64_t estimateAfterTick =
+                estimatedMemoryBytes_.load(std::memory_order_relaxed);
+            if (estimateAfterTick >= estimateBeforeTick) {
+                ++evictionNoProgressTicks_;
+            } else {
+                evictionNoProgressTicks_ = 0;
+            }
+            evictionLastEstimate_ = estimateAfterTick;
+        }
+
         if (passesRun > 0) {
             spdlog::info("Eviction tick complete: {} passes, {} MB freed (pressure={})",
                         passesRun, totalFreed / (1024 * 1024),

+ 34 - 0
service/src/memory_store.hpp

@@ -97,6 +97,21 @@ public:
         // fired so observability tooling can annotate "eviction storms".
         uint32_t evictionBurstThreshold = 10000;
 
+        // v2.4.3 — eviction drain cap. Ceiling on the share of the resident
+        // set a SINGLE pressure episode may evict, as a percent of the live
+        // doc count when that episode began. 0 disables the cap.
+        //
+        // Exists because eviction trusts estimatedMemoryBytes_ to fall as docs
+        // leave. When that estimate is wrong - held up by memory eviction
+        // cannot free - the target is unreachable and the loop keeps trimming
+        // until the store is empty. Production hit exactly this: 5996 of ~5990
+        // docs evicted over 30 minutes while the estimate went UP (347 -> 358
+        // MB), and reads fell back to an empty MemoryStore.
+        //
+        // Eviction is a cache trim, not a store drain. Hitting this cap means
+        // the estimator is wrong, so it is an ERROR, not a warning.
+        uint32_t evictionMaxEpisodePercent = 50;
+
         Config()
             : maxMemoryBytes(800ULL * 1024 * 1024)   // 800 MB default
             , expirationCheckIntervalMs(1000)         // Check TTL every second
@@ -828,6 +843,25 @@ private:
     // mutable because pressure() is const but may update this as a side effect
     mutable std::atomic<MemoryPressure> lastObservedPressure_{MemoryPressure::Normal};
 
+    // v2.4.3 eviction drain cap state. Touched only by the eviction thread,
+    // reset every time pressure returns to Normal (one "episode" = one
+    // continuous stretch above the soft threshold).
+    uint64_t evictionEpisodeStartDocs_ = 0;   // live docs when the episode began
+    uint64_t evictionEpisodeEvicted_ = 0;     // docs evicted during this episode
+    bool     evictionDrainCapLogged_ = false; // ERROR is logged once per episode
+    uint32_t evictionNormalTicks_ = 0;        // consecutive Normal-pressure ticks
+    uint64_t evictionLastEstimate_ = 0;       // estimate at the end of last tick
+    uint32_t evictionNoProgressTicks_ = 0;    // ticks that evicted without relief
+
+    // An episode re-arms only after pressure has been Normal for this many
+    // consecutive ticks. A single tick is not enough: eviction itself briefly
+    // drops the estimate below the threshold, so a per-tick reset lets the
+    // cap restart over and over and still drain the store in stages.
+    static constexpr uint32_t kEvictionRearmNormalTicks = 3;
+    // Consecutive evicting ticks with no fall in the estimate before we
+    // conclude the estimate is not tracking what eviction frees.
+    static constexpr uint32_t kEvictionNoProgressLimit = 3;
+
     // ===== LRU Eviction =====
 
     // Evicted documents: collection -> docId -> stub

+ 15 - 4
service/src/storage/document_store_lmdb.cpp

@@ -233,10 +233,21 @@ LmdbDocumentStore::try_open_for_read(ReadTxn& rtxn,
     int rc = mdb_dbi_open(rtxn.raw(), name.c_str(), 0, &raw_dbi);
     if (rc == MDB_NOTFOUND) return std::nullopt;
     if (rc != MDB_SUCCESS) throw_mdb(rc, "dbi_open (collection read)");
-    {
-        std::lock_guard<std::mutex> lock(cache_mutex_);
-        dbi_cache_[key] = raw_dbi;
-    }
+
+    // Deliberately NOT cached. LMDB keeps a handle private to the opening
+    // transaction until that transaction COMMITS; if it aborts, the handle is
+    // closed. ReadTxn always aborts (ReadTxn::~ReadTxn), so caching this
+    // handle would hand out a dangling MDB_dbi to every later caller, and
+    // every mdb_put / mdb_cursor_open using it fails with EINVAL - for the
+    // life of the process, for that one collection.
+    //
+    // In production this hit whichever collection happened to be READ before
+    // it was WRITTEN after a restart: peers were fine because open_for_write
+    // commits, which promotes the handle into the env's shared table.
+    //
+    // The handle IS valid inside `rtxn`, which is all the caller needs. The
+    // cost is one mdb_dbi_open per read on collections this process has never
+    // written; once a write happens, open_for_write caches it permanently.
     return raw_dbi;
 }
 

+ 14 - 0
tests/CMakeLists.txt

@@ -471,6 +471,20 @@ endif()
 find_package(Threads REQUIRED)
 target_link_libraries(test_dual_write_mirror PRIVATE Threads::Threads)
 
+# Five test files assert via assert(), which CMAKE_BUILD_TYPE=Release compiles
+# out via -DNDEBUG. That silently reduced 104 assertions across test_views,
+# test_eviction, test_snapshot_durability, test_timestamp_precision and
+# test_config_dropins to no-ops: the suites ran, printed PASS, and checked
+# nothing. Strip NDEBUG from test targets so assertions are live in every
+# build type. Production targets are unaffected.
+foreach(_t
+    test_vector_storage test_views test_timestamp_precision
+    test_snapshot_durability test_config_dropins test_eviction
+    test_json_parse test_doc_binary test_lmdb_env test_document_store
+    test_migrate_v1_to_v2 test_dual_write_mirror)
+  target_compile_options(${_t} PRIVATE -UNDEBUG)
+endforeach()
+
 # Register with CTest
 enable_testing()
 add_test(NAME vector_storage COMMAND test_vector_storage)

+ 56 - 0
tests/load_test/config_crash.json

@@ -0,0 +1,56 @@
+{
+  "log_level": "info",
+  "storage": {
+    "bind_address": "127.0.0.1",
+    "rpc_port": 9005,
+    "node_id": "loadtest",
+    "data_directory": "/tmp/smartbotic-loadtest-crash/data",
+    "memory": {
+      "max_memory_mb": 32,
+      "eviction_threshold_percent": 80,
+      "eviction_target_percent": 50,
+      "eviction_check_interval_ms": 200,
+      "eviction_chunk_size": 500,
+      "eviction_chunk_pause_ms": 30,
+      "max_eviction_passes_per_trigger": 15,
+      "hot_write_floor_ms": 2000,
+      "memory_soft_percent": 60,
+      "memory_hard_percent": 80,
+      "memory_emergency_percent": 92,
+      "eviction_burst_threshold": 2000
+    },
+    "persistence": {
+      "wal_sync_interval_ms": 100,
+      "snapshot_interval_sec": 3600,
+      "compression": "lz4",
+      "snapshots": {
+        "validate_after_write": true,
+        "cleanup_only_if_verified": true
+      },
+      "recovery": {
+        "mode": "normal",
+        "auto_escalate": true,
+        "allow_empty_on_fresh_install": true
+      }
+    },
+    "encryption": {
+      "enabled": false
+    },
+    "migrations": {
+      "enabled": false
+    },
+    "files": {
+      "max_file_size_mb": 100
+    },
+    "replication": {
+      "enabled": false
+    },
+    "grpc": {
+      "max_receive_message_size_mb": 100,
+      "max_send_message_size_mb": 100,
+      "resource_quota_memory_mb": 64,
+      "max_concurrent_subscribe_streams": 50,
+      "max_concurrent_file_streams": 10
+    }
+  }
+}

BIN
tests/load_test/load_test_crash_insert


BIN
tests/load_test/load_test_crash_verify


BIN
tests/load_test/load_test_integrity


BIN
tests/load_test/load_test_mixed


BIN
tests/load_test/load_test_mixed_v23


BIN
tests/load_test/load_test_multi_project


BIN
tests/load_test/load_test_multi_project_replication


BIN
tests/load_test/load_test_pinned


BIN
tests/load_test/load_test_replica_eviction


BIN
tests/load_test/load_test_replication


BIN
tests/load_test/load_test_replication_v23


+ 66 - 0
tests/test_document_store.cpp

@@ -741,6 +741,69 @@ void test_scan_total_matched_and_has_more_counters() {
     check(r.has_more, "has_more=true (4 remain)");
 }
 
+// --- dbi handle lifetime -------------------------------------------------
+//
+// LMDB closes any sub-db handle opened by a transaction that ABORTS. Read
+// transactions always abort (ReadTxn::~ReadTxn), so a handle first opened on
+// the read path and cached in dbi_cache_ dangles the moment that scan
+// returns. Every later use of the cached handle fails with EINVAL.
+//
+// Production symptom: after a restart, one collection - whichever happened to
+// be READ before it was WRITTEN in that process - failed permanently with
+// "LMDB cursor_open: Invalid argument", while its peers worked. Peers were
+// fine because their handle was first opened inside a committing write txn,
+// which makes it valid process-wide.
+//
+// The second store instance below is the restart: same env on disk, empty
+// handle cache, collection reached by read first.
+void test_read_first_collection_survives_repeated_access() {
+    // NOTE: the env must be CLOSED and reopened, not just the store. MDB_dbi
+    // handles live in the environment's shared table, so a handle opened by a
+    // committing write txn stays valid for the life of the env - reusing one
+    // LmdbEnv hides the bug entirely. Only a fresh env reproduces a restart.
+    const std::string path = make_tmpdir("dbi-read-first");
+    const LmdbEnvOpts opts{path, 64ULL << 20, 256, 126, false};
+
+    // Process 1: create the collection through the write path, then close.
+    {
+        LmdbEnv env(opts);
+        LmdbDocumentStore writer(env);
+        writer.put("sessions", "s-1", make_doc("s-1", "sessions",
+                                               nlohmann::json{{"user", "alice"}}));
+    }
+
+    // Process 2: fresh env, empty handle table and empty dbi cache. It READS
+    // "sessions" before ever writing it, so the handle is first opened inside
+    // a ReadTxn - which aborts.
+    LmdbEnv env2(opts);
+    LmdbDocumentStore reader(env2);
+    struct Cleanup {
+        const std::string& p;
+        ~Cleanup() { std::error_code ec; fs::remove_all(p, ec); }
+    } cleanup{path};
+
+    Query q;
+    q.limit = 100;
+
+    auto first = reader.scan("sessions", q);
+    check(first.documents.size() == 1, "read-first scan sees the doc");
+
+    // Second access reuses the cached handle. Before the fix this threw
+    // "LMDB cursor_open: Invalid argument".
+    auto second = reader.scan("sessions", q);
+    check(second.documents.size() == 1, "cached handle still valid on re-scan");
+
+    // get() and count() go through the same cache.
+    auto got = reader.get("sessions", "s-1");
+    check(got.has_value(), "get works after a read-first scan");
+    check(reader.count("sessions") == 1, "count works after a read-first scan");
+
+    // Writing through the same cached handle must work too.
+    reader.put("sessions", "s-2", make_doc("s-2", "sessions",
+                                           nlohmann::json{{"user", "bob"}}));
+    check(reader.count("sessions") == 2, "put works through a read-first handle");
+}
+
 }  // namespace
 
 int main() {
@@ -783,6 +846,9 @@ int main() {
     test_scan_limit_offset_pagination();
     test_scan_total_matched_and_has_more_counters();
 
+    // dbi handle lifetime.
+    test_read_first_collection_survives_repeated_access();
+
     std::cout << "test_document_store: all passed\n";
     return 0;
 }

+ 58 - 0
tests/test_eviction.cpp

@@ -196,11 +196,69 @@ void test_eviction_chunk_size_honored() {
               << " docs in ~1 tick, chunkSize=" << cfg.evictionChunkSize << ")\n";
 }
 
+// v2.4.3 — eviction must never drain the store.
+//
+// Regression for the production incident: the memory estimate stopped
+// tracking what eviction could free, so `current <= targetBytes` was never
+// satisfied and the loop trimmed one chunk per tick for 30 minutes until 5996
+// of ~5990 docs were gone. Reads then fell back to an empty MemoryStore and
+// every collection reported 0 docs, which looked exactly like data loss.
+//
+// Here the target is made unreachable on purpose (evictionTargetPercent = 0,
+// so targetBytes = 0 and no amount of eviction satisfies it). Before the drain
+// cap this emptied the store; now it must stop at the cap.
+void test_eviction_never_drains_the_store() {
+    auto cfg = evictionTestConfig(1);
+    cfg.evictionTargetPercent = 0;        // unreachable target
+    cfg.memorySoftPercent = 1;            // stay under pressure permanently
+    cfg.memoryHardPercent = 2;
+    cfg.memoryEmergencyPercent = 99;      // don't trip admission control
+    cfg.hotWriteFloorMs = 0;              // everything is evictable
+    cfg.evictionChunkSize = 50;
+    cfg.evictionCheckIntervalMs = 20;
+    cfg.evictionMaxEpisodePercent = 50;   // cap under test
+    MemoryStore store(cfg);
+    store.start();
+    store.createCollection("docs", CollectionOptions{});
+
+    const uint64_t seeded = 400;
+    fillStore(store, "docs", seeded, 2048);
+
+    auto liveDocs = [&] {
+        uint64_t n = 0;
+        for (const auto& c : store.getMemoryStatsSnapshot().collections) {
+            n += c.documentCount;
+        }
+        return n;
+    };
+
+    const uint64_t before = liveDocs();
+    assert(before > 0);
+
+    // Give the eviction thread many ticks — far more than it would need to
+    // empty the store one 50-doc chunk at a time.
+    std::this_thread::sleep_for(std::chrono::milliseconds(1200));
+
+    const uint64_t after = liveDocs();
+    const uint64_t evicted = (before > after) ? (before - after) : 0;
+
+    // The cap is 50% of the episode's starting set. Allow one chunk of
+    // overshoot, since the check runs per tick rather than per doc.
+    const uint64_t cap = before / 2 + cfg.evictionChunkSize;
+    assert(evicted <= cap);
+    assert(after > 0);  // the store must never be drained
+
+    store.stop();
+    std::cout << "PASS: eviction drain cap held (" << evicted << " of " << before
+              << " evicted, " << after << " docs still resident)\n";
+}
+
 int main() {
     test_pressure_levels();
     test_hot_write_floor_protects_recent_writes();
     test_priority_low_evicted_first();
     test_eviction_chunk_size_honored();
+    test_eviction_never_drains_the_store();
     std::cout << "\nAll eviction tests PASSED!\n";
     return 0;
 }

Some files were not shown because too many files changed in this diff