Prechádzať zdrojové kódy

fix(relations): close-out review round 3 - lock inversion, gauge surface, TTL config

Five findings from the round-3 review. One (the inversion) was introduced by my
own round-2 diff.

1. MUST FIX - lock-order inversion on ttlBlockedMutex_. Phase 2 called
   clearTtlBlocked() from inside the globalMutex_+coll->mutex scope while phase 1
   holds ttlBlockedMutex_ across the acquisition of both: two sweepers deadlock,
   and because both then pin globalMutex_ shared, the next getOrCreateCollection
   for a new collection blocks on the exclusive acquire and the whole store hangs
   (restart only). Not reachable through the shipped service - expirationLoop is
   the only caller of expireDocuments - but reachable from the new tests, and live
   the moment a second caller exists. The pre-check scope now records an outcome
   and clears after the scope closes; the header comment (which claimed
   "INNERMOST", the opposite of what phase 1 does) states the real order and names
   the deadlock. ttlSweepCounter_/ttlBlockedSummarySweep_ are now atomic - they
   are touched outside that mutex, so they were a data race on the retry schedule.

2. A dropped collection leaked its blocked-set entries. The proposed clear in the
   sweeper cannot fix it - phase 1 builds candidates by walking collections_, so
   once the collection is gone no candidate is produced for those ids again and
   the branch is unreachable for exactly the leaking case (the test proved this).
   Fixed in dropCollection() instead, after the globalMutex_ scope closes: after
   the drop so a concurrent sweep cannot re-add, outside the lock for reason 1.

3. MUST FIX - the gauge was exposed to nobody, so the WARN pointed operators at a
   C++ method they cannot call. Surfaced as GetMemoryStats.ttl_blocked_documents
   (gauge) and .ttl_expiry_blocked_events (counter), and the log text now names
   them. Client::MemoryStats deliberately untouched - growing a public struct
   under an unchanged soname is the webserver restart-loop bug.

4. The finding-1 pre-check is now gated on a hook being installed. Without one
   every candidate takes Proceed, whose re-check is airtight, so the pre-check was
   a second EXCLUSIVE collection lock per candidate guarding a window that path
   does not have. No test: the gate is behaviour-neutral by construction.

5. The four TTL sweep thresholds are wired into the config loader under
   `storage.ttl`, rather than dropping the word "tunable" from the header.

Residual window: accepted for this release, not built. Recorded in the report that
the versioned-delete option does NOT close it (no version-checked delete exists,
and the cascade never goes through remove()), and that the claim flag is the one
that would - so it is not restarted from the wrong end. The offered
re-read-before-writeCascadeWal improvement is skipped: it needs a TTL-only
precondition threaded through the cascade entry point shared with request-path
Delete.

test_relation_enforcement 364 -> 370; relation_manager 70, document_ttl 22,
subdb_identity 264, dual_write_mirror 73, document_store green. E2E
test_relations.sh passes with two new phases. VERSION untouched.
fszontagh 1 mesiac pred
rodič
commit
f75239c14f

+ 18 - 0
proto/database.proto

@@ -1262,4 +1262,22 @@ message GetMemoryStatsResponse {
     uint64 last_eviction_timestamp = 6;  // ms since epoch, 0 if none
     uint64 last_eviction_docs = 7;
     uint64 last_eviction_bytes_freed = 8;
+
+    // v2.11.0 — TTL expiries a relation is refusing.
+    //
+    // ttl_blocked_documents is a GAUGE: how many documents are past their TTL
+    // right now and NOT being expired because deleting them would break a
+    // `restrict` relation (or because the cascade path is refusing while the
+    // LMDB mirror is unhealthy/drifted). It falls when a block lifts. Nonzero
+    // means those documents are outliving their retention policy - the
+    // deliberate alternative to silently orphaning their children.
+    //
+    // ttl_expiry_blocked_events is a monotonic COUNTER of refusal events, so it
+    // keeps rising while a block persists (each blocked document is retried on a
+    // slow cadence). Use the gauge for "how bad is it now", the counter for
+    // "is this still happening".
+    //
+    // Both are zero on any install that declares no relations.
+    uint64 ttl_blocked_documents = 9;
+    uint64 ttl_expiry_blocked_events = 10;
 }

+ 12 - 0
service/src/database_grpc_impl.cpp

@@ -2496,6 +2496,18 @@ grpc::Status DatabaseGrpcImpl::GetMemoryStats(
     response->set_last_eviction_docs(snap.lastEvictionDocs);
     response->set_last_eviction_bytes_freed(snap.lastEvictionBytesFreed);
 
+    // v2.11.0 close-out (round-3 review, item 3) — the ONLY operator-reachable
+    // surface for a document silently outliving its TTL. Before this the WARN in
+    // relations/relation_cascade.cpp pointed at MemoryStore::ttlBlockedDocumentCount(),
+    // a C++ method no operator can call, and Stats::ttlExpiryBlockedByRelation
+    // was in no proto response at all - a signal nobody can read is not a signal.
+    // Deliberately NOT added to Client::MemoryStats: that is a public struct and
+    // growing it under an unchanged soname is what crashed the webserver in a
+    // restart loop (see CLAUDE.md's ABI note). GetMemoryStats is the operator
+    // path; the C++ client keeps its current struct byte-for-byte.
+    response->set_ttl_blocked_documents(store_.ttlBlockedDocumentCount());
+    response->set_ttl_expiry_blocked_events(store_.getStats().ttlExpiryBlockedByRelation);
+
     auto toProtoPriority = [](MemoryPriority p) -> pb::MemoryPriority {
         switch (p) {
             case MemoryPriority::Low:    return pb::MEMORY_PRIORITY_LOW;

+ 19 - 0
service/src/database_service.cpp

@@ -1141,6 +1141,20 @@ DatabaseService::Config DatabaseService::parseConfig(const nlohmann::json& json)
             config.evictionMaxEpisodePercent = memory.value("eviction_max_episode_percent", config.evictionMaxEpisodePercent);
         }
 
+        // v2.11.0 close-out — TTL sweep budgets. Its own block rather than
+        // `memory`, because these bound expiry work per sweep, not memory.
+        if (db.contains("ttl")) {
+            auto& ttl = db["ttl"];
+            config.ttlMaxCandidatesPerSweep =
+                ttl.value("max_candidates_per_sweep", config.ttlMaxCandidatesPerSweep);
+            config.ttlMaxBlockedRetriesPerSweep =
+                ttl.value("max_blocked_retries_per_sweep", config.ttlMaxBlockedRetriesPerSweep);
+            config.ttlBlockedRetrySweeps =
+                ttl.value("blocked_retry_sweeps", config.ttlBlockedRetrySweeps);
+            config.ttlBlockedSummarySweeps =
+                ttl.value("blocked_summary_sweeps", config.ttlBlockedSummarySweeps);
+        }
+
         // Persistence settings
         if (db.contains("persistence")) {
             auto& persistence = db["persistence"];
@@ -1325,6 +1339,11 @@ void DatabaseService::setupComponents() {
     storeConfig.memoryEmergencyPercent = config_.memoryEmergencyPercent;
     storeConfig.evictionBurstThreshold = config_.evictionBurstThreshold;
     storeConfig.evictionMaxEpisodePercent = config_.evictionMaxEpisodePercent;
+    // v2.11.0 close-out — TTL sweep budgets (storage.ttl).
+    storeConfig.ttlMaxCandidatesPerSweep = config_.ttlMaxCandidatesPerSweep;
+    storeConfig.ttlMaxBlockedRetriesPerSweep = config_.ttlMaxBlockedRetriesPerSweep;
+    storeConfig.ttlBlockedRetrySweeps = config_.ttlBlockedRetrySweeps;
+    storeConfig.ttlBlockedSummarySweeps = config_.ttlBlockedSummarySweeps;
     store_ = std::make_unique<MemoryStore>(storeConfig);
 
     // Create view manager (cache loaded in initialize() after persistence recovery)

+ 13 - 0
service/src/database_service.hpp

@@ -151,6 +151,19 @@ public:
         // may evict. 0 disables. See MemoryStore::Config for the rationale.
         uint32_t evictionMaxEpisodePercent = 50;
 
+        // v2.11.0 close-out (round-3 review, item 5) — TTL sweep budgets, under
+        // `storage.ttl` in config.json. Mirrors of MemoryStore::Config's fields
+        // of the same name; see there for what each one bounds and why the
+        // blocked-document budget is separate from the fresh one. Present here
+        // because a header calling them tunable while nothing read them from
+        // config meant production always ran the defaults and an operator facing
+        // a blocked-document flood could not change the cadence without a
+        // rebuild.
+        uint32_t ttlMaxCandidatesPerSweep = 10000;
+        uint32_t ttlMaxBlockedRetriesPerSweep = 100;
+        uint32_t ttlBlockedRetrySweeps = 60;
+        uint32_t ttlBlockedSummarySweeps = 300;
+
         // Persistence settings
         uint32_t walSyncIntervalMs = 100;
         uint32_t snapshotIntervalSec = 3600;

+ 82 - 19
service/src/memory_store.cpp

@@ -180,6 +180,7 @@ bool MemoryStore::createCollection(const std::string& name, const CollectionOpti
 }
 
 bool MemoryStore::dropCollection(const std::string& name) {
+    {
     std::unique_lock<std::shared_mutex> lock(globalMutex_);
 
     auto it = collections_.find(name);
@@ -225,6 +226,25 @@ bool MemoryStore::dropCollection(const std::string& name) {
 
     // Update memory tracking atomically
     estimatedMemoryBytes_.fetch_sub(totalSize, std::memory_order_relaxed);
+    }   // globalMutex_ released here — see below.
+
+    // v2.11.0 close-out (round-3 review, item 2) — release any TTL blocked-set
+    // entries for this collection.
+    //
+    // Necessary because expireDocuments() cannot clean these up itself: phase 1
+    // builds candidates by walking `collections_`, so once the collection is gone
+    // no candidate is ever produced for those ids again and phase 2 never sees
+    // them. They would sit in ttlBlocked_ for the process lifetime, inflating
+    // GetMemoryStats.ttl_blocked_documents with documents that no longer exist.
+    //
+    // ⚠ AFTER the globalMutex_ scope, and both halves of that matter.
+    // AFTER the drop, so a concurrent sweep cannot re-add an entry between the
+    // clear and the erase (once the collection is gone nothing can re-add).
+    // OUTSIDE the lock, because ttlBlockedMutex_ is ordered BEFORE globalMutex_
+    // (phase 1 holds it across that acquisition) and taking it while holding
+    // globalMutex_ exclusively would close a deadlock cycle - the same inversion
+    // this review found in phase 2.
+    clearTtlBlockedForCollection(name);
 
     return true;
 }
@@ -2212,7 +2232,7 @@ void MemoryStore::setTtlExpiryRelationHook(TtlExpiryRelationHook hook) {
 uint64_t MemoryStore::expireDocuments() {
     uint64_t expired = 0;
     auto now = currentTimeMs();
-    const uint64_t sweep = ++ttlSweepCounter_;
+    const uint64_t sweep = ttlSweepCounter_.fetch_add(1, std::memory_order_relaxed) + 1;
 
     // ---- Phase 1: collect candidates under the locks, mutate nothing. -----
     // See the header comment on expireDocuments() for why the phases are split:
@@ -2294,8 +2314,9 @@ uint64_t MemoryStore::expireDocuments() {
     // Periodic summary instead of a per-document WARN per sweep. A flood is not
     // a signal, and at the 1s default a persistent block was ~one WARN per
     // blocked document per second, indefinitely.
-    if (blockedNotDue > 0 && sweep - ttlBlockedSummarySweep_ >= kBlockedSummarySweeps) {
-        ttlBlockedSummarySweep_ = sweep;
+    if (blockedNotDue > 0 &&
+        sweep - ttlBlockedSummarySweep_.load(std::memory_order_relaxed) >= kBlockedSummarySweeps) {
+        ttlBlockedSummarySweep_.store(sweep, std::memory_order_relaxed);
         spdlog::warn("TTL expiry: {} document(s) are still past their TTL and cannot be "
                      "expired because a relation blocks the delete (each was logged once "
                      "when first blocked; they are retried every {} sweeps). They will "
@@ -2320,26 +2341,55 @@ uint64_t MemoryStore::expireDocuments() {
         // See the header comment for the RESIDUAL: a write landing during the
         // hook's own cascade (a WAL fsync plus an LMDB commit) is still
         // possible. Narrowing, not eliminating.
-        uint64_t currentExpiresAt = 0;
-        {
+        //
+        // ⚠ SKIPPED ENTIRELY WITH NO HOOK INSTALLED (round-3 review, item 4).
+        // Without a hook every candidate takes the Proceed path, whose own
+        // re-check shares one critical section with its erase and is therefore
+        // airtight - so this pre-check would be a second EXCLUSIVE collection
+        // lock per candidate (up to ttlMaxCandidatesPerSweep per sweep) to guard
+        // a window that path does not have. Every install that declares no
+        // relation stays exactly as fast as before v2.11.0.
+        //
+        // ⚠ LOCK ORDER (round-3 review, item 1): clearTtlBlocked() MUST NOT be
+        // called from inside this scope. ttlBlockedMutex_ is ordered BEFORE
+        // globalMutex_/coll->mutex (phase 1 holds it across both), so taking it
+        // while holding either closes a cycle: one sweeper holding
+        // ttlBlockedMutex_ and waiting for coll->mutex against another holding
+        // coll->mutex and waiting for ttlBlockedMutex_ - and both then pin
+        // globalMutex_ shared, so the next getOrCreateCollection() for a new
+        // collection blocks on the exclusive acquire and the whole store hangs.
+        // The outcome is therefore recorded and acted on AFTER the scope closes.
+        enum class PreCheck { Ok, Vanished, NoLongerExpired };
+        PreCheck pre = PreCheck::Ok;
+        if (ttlExpiryRelationHook_) {
             std::shared_lock<std::shared_mutex> globalLock(globalMutex_);
             auto collIt = collections_.find(cand.collection);
-            if (collIt == collections_.end()) continue;   // collection dropped
-            auto* coll = collIt->second.get();
-            std::unique_lock<std::shared_mutex> collLock(coll->mutex);
-            auto docIt = coll->documents.find(cand.id);
-            if (docIt == coll->documents.end()) {
-                // Gone (deleted, or evicted). Drop the stale index entry so it
-                // is not re-examined every sweep forever.
-                removeFromExpirationIndex(*coll, cand.id, cand.expiresAt);
-                clearTtlBlocked(cand.collection, cand.id);
-                continue;
+            if (collIt == collections_.end()) {
+                // Collection dropped. Nothing to unindex (it went with the
+                // collection), but the blocked entry has to go or it inflates
+                // ttlBlockedDocumentCount() for the process lifetime - phase 1
+                // never revisits a key whose collection is gone, and
+                // dropCollection() knows nothing about ttlBlocked_ (round-3
+                // review, item 2).
+                pre = PreCheck::Vanished;
+            } else {
+                auto* coll = collIt->second.get();
+                std::unique_lock<std::shared_mutex> collLock(coll->mutex);
+                auto docIt = coll->documents.find(cand.id);
+                if (docIt == coll->documents.end()) {
+                    // Gone (deleted, or evicted). Drop the stale index entry so
+                    // it is not re-examined every sweep forever.
+                    removeFromExpirationIndex(*coll, cand.id, cand.expiresAt);
+                    pre = PreCheck::Vanished;
+                } else if (docIt->second.expiresAt == 0 || docIt->second.expiresAt > now) {
+                    pre = PreCheck::NoLongerExpired;
+                }
             }
-            currentExpiresAt = docIt->second.expiresAt;
         }
-        if (currentExpiresAt == 0 || currentExpiresAt > now) {
-            // The TTL was extended or cleared while this sweep was running. Not
-            // expired any more - and no longer blocked either, if it was.
+        if (pre != PreCheck::Ok) {
+            // Vanished, or its TTL was extended/cleared while this sweep was
+            // running. Either way it is not expiring now, and it is no longer
+            // blocked either. Called with NO other lock held - see above.
             clearTtlBlocked(cand.collection, cand.id);
             continue;
         }
@@ -2485,6 +2535,19 @@ void MemoryStore::clearTtlBlocked(const std::string& collection, const std::stri
     ttlBlocked_.erase(ttlBlockedKey(collection, id));
 }
 
+void MemoryStore::clearTtlBlockedForCollection(const std::string& collection) {
+    std::lock_guard<std::mutex> lock(ttlBlockedMutex_);
+    if (ttlBlocked_.empty()) return;
+    // Keys are "<collection>\0<id>" (see ttlBlockedKey), so a prefix match on
+    // the collection plus the separator is exact - it cannot match a different
+    // collection whose name merely starts with this one.
+    const std::string prefix = collection + std::string(1, '\0');
+    for (auto it = ttlBlocked_.begin(); it != ttlBlocked_.end();) {
+        it = (it->first.compare(0, prefix.size(), prefix) == 0) ? ttlBlocked_.erase(it)
+                                                               : std::next(it);
+    }
+}
+
 size_t MemoryStore::ttlBlockedDocumentCount() const {
     std::lock_guard<std::mutex> lock(ttlBlockedMutex_);
     return ttlBlocked_.size();

+ 39 - 7
service/src/memory_store.hpp

@@ -127,8 +127,9 @@ public:
         // remembered, skipped for free until `ttlBlockedRetrySweeps` sweeps have
         // passed, and then retried against their own small budget.
         //
-        // Tunable mainly so the behaviour is testable at small numbers; the
-        // defaults are what production runs.
+        // Wired into the config loader under `storage.ttl` (see parseConfig), so
+        // an operator facing a blocked-document flood can change the cadence
+        // without a rebuild - which is what makes calling them tunable true.
         uint32_t ttlMaxCandidatesPerSweep = 10000;
         uint32_t ttlMaxBlockedRetriesPerSweep = 100;
         uint32_t ttlBlockedRetrySweeps = 60;      // ~1 min at the 1s default
@@ -797,6 +798,11 @@ public:
      * NOT being expired because a relation blocks the delete. A GAUGE (it falls
      * when a block lifts), unlike Stats::ttlExpiryBlockedByRelation, which
      * counts refusal events. Zero on any install that declares no relations.
+     *
+     * Surfaced to operators as GetMemoryStats.ttl_blocked_documents (with the
+     * event counter as ttl_expiry_blocked_events). It is the only visibility
+     * there is for a document outliving its retention policy, so if a caller is
+     * added here it must reach an RPC too.
      */
     [[nodiscard]] size_t ttlBlockedDocumentCount() const;
 
@@ -1332,16 +1338,42 @@ private:
     // it in `collections_` iteration order from being swept at all.
     //
     // Guarded by its own mutex rather than a collection lock: it is keyed across
-    // collections, and it is read in phase 1 while globalMutex_ is held shared,
-    // so it must be the INNERMOST lock. Nothing taken under it takes any other
-    // lock.
+    // collections, so no single collection lock covers it.
+    //
+    // ⚠ LOCK ORDER, and it was WRONG in the first cut (round-3 review, item 1):
+    // ttlBlockedMutex_ is ordered **BEFORE** globalMutex_ and coll->mutex, not
+    // after. Phase 1 of expireDocuments() holds it across the acquisition of
+    // both, so the ONLY legal edge is
+    //
+    //     ttlBlockedMutex_  ->  globalMutex_  ->  coll->mutex
+    //
+    // and it must NEVER be acquired while either of the others is held. The
+    // first cut called clearTtlBlocked() from inside phase 2's
+    // globalMutex_+coll->mutex scope, which closed the cycle: sweeper A holding
+    // ttlBlockedMutex_ and waiting for coll->mutex, sweeper B holding
+    // coll->mutex and waiting for ttlBlockedMutex_ - and because both then pin
+    // globalMutex_ shared forever, the next getOrCreateCollection() for a new
+    // collection blocks on the exclusive acquire and the WHOLE STORE hangs
+    // (restart only). Every other call site takes it with no other lock held,
+    // which is also legal. Nothing taken under it takes any other lock.
+    //
+    // Not reachable through the shipped service today - expirationLoop() is the
+    // only caller of expireDocuments() and no RPC triggers a sweep - but a
+    // second caller, or the obvious next improvement (clearing a block from
+    // update()/dropRelation() so a lifted block retries at once), makes it live.
     mutable std::mutex ttlBlockedMutex_;
     std::unordered_map<std::string, uint64_t> ttlBlocked_;
-    uint64_t ttlSweepCounter_ = 0;
-    uint64_t ttlBlockedSummarySweep_ = 0;
+    // Atomic because they are read and written OUTSIDE ttlBlockedMutex_ (the
+    // sweep number is taken before phase 1 acquires anything), so with two
+    // concurrent sweepers a plain member is a data race on the retry schedule.
+    std::atomic<uint64_t> ttlSweepCounter_{0};
+    std::atomic<uint64_t> ttlBlockedSummarySweep_{0};
 
     static std::string ttlBlockedKey(const std::string& collection, const std::string& id);
     void clearTtlBlocked(const std::string& collection, const std::string& id);
+    // Called by dropCollection(), from OUTSIDE globalMutex_ - see the lock-order
+    // note above and the call site.
+    void clearTtlBlockedForCollection(const std::string& collection);
 
     std::atomic<bool>* mirrorHealthy_ = nullptr;
     std::atomic<uint64_t>* mirrorDriftCount_ = nullptr;

+ 4 - 3
service/src/relations/relation_cascade.cpp

@@ -435,9 +435,10 @@ MemoryStore::TtlExpiryAction ttlExpiryDecision(
             // children, and the alternative is what this change removes.
             logBlock(fmt::format(
                 "TTL expiry of '{}/{}' is BLOCKED by a restrict relation, so the document "
-                "remains past its TTL and will be retried on a slower cadence "
-                "(MemoryStore::ttlBlockedDocumentCount() is how many are stuck right now): "
-                "{}", qualifiedParentCollection, parentId, err));
+                "remains past its TTL and will be retried on a slower cadence. How many are "
+                "stuck right now: GetMemoryStats.ttl_blocked_documents (a periodic summary "
+                "line also reports it). Reason: {}",
+                qualifiedParentCollection, parentId, err));
             return Action::Skip;
         }
 

+ 24 - 0
tests/load_test/test_relations.sh

@@ -71,6 +71,12 @@ cat > "$WORK/config.json" <<EOF
     "rpc_port": $PORT,
     "encryption": { "enabled": false, "key_file": "$WORK/data/storage.key" },
     "migrations": { "enabled": true, "directory": "$WORK/migrations" },
+    "ttl": {
+      "max_candidates_per_sweep": 5000,
+      "max_blocked_retries_per_sweep": 50,
+      "blocked_retry_sweeps": 30,
+      "blocked_summary_sweeps": 120
+    },
     "replication": { "enabled": false }
   }
 }
@@ -266,6 +272,24 @@ grep -E "rebuilt 'mig_rel'" "$WORK/boot1.log" \
     && fail "the expected path was logged as the 'sub-db was absent' FAULT"
 echo "  boot log: $(grep 'a migration declared this relation' "$WORK/boot1.log" | tail -1 | sed 's/.*\] //')"
 
+echo
+echo "=== phase: the TTL blocked-document gauge is reachable over the wire ==="
+# The gauge is the ONLY operator visibility for a document outliving its TTL
+# because a relation blocks the delete. Before the close-out review it existed
+# only as a C++ method with no RPC, which is not a signal. This asserts the field
+# is part of GetMemoryStats and the RPC still answers after the proto change; the
+# nonzero path is unit-tested (test_relation_enforcement's restrict/starvation
+# tests assert the gauge directly).
+OUT="$(rpc GetMemoryStats '{}' || true)"
+grep -q "pressureLevel" <<<"$OUT" \
+    || { echo "$OUT"; fail "GetMemoryStats did not answer after the proto change"; }
+if ! "$GRPCURL" -plaintext -import-path "$REPO/proto" -proto database.proto \
+        -msg-template describe smartbotic.databasepb.GetMemoryStatsResponse 2>&1 \
+        | grep -q "ttlBlockedDocuments"; then
+    fail "ttl_blocked_documents is not on GetMemoryStatsResponse - the gauge is unreachable"
+fi
+echo "  GetMemoryStatsResponse carries ttlBlockedDocuments / ttlExpiryBlockedEvents"
+
 echo
 echo "=== phase: DropProject must clean up its relation declarations (close-out) ==="
 # Declarations live in the GLOBAL _relations collection, so dropping a project

+ 37 - 0
tests/test_relation_enforcement.cpp

@@ -2743,6 +2743,42 @@ void test_ttl_blocked_documents_do_not_starve_the_sweep() {
     mstore.stop();
 }
 
+// Round-3 review, item 2 — dropping a blocked document's collection must not
+// leak its blocked-set entry. Phase 1 never revisits a key whose collection is
+// gone and dropCollection() knows nothing about the blocked set, so a leaked
+// entry would inflate ttlBlockedDocumentCount() - the gauge an operator reads -
+// for the rest of the process's life.
+void test_ttl_blocked_entry_is_released_when_the_collection_is_dropped() {
+    TmpEnv t("ttl-blocked-drop");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-blocked-drop-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore::Config cfg;
+    cfg.ttlBlockedRetrySweeps = 1;   // due again on the very next sweep
+    MemoryStore mstore(cfg);
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    seedTtlParentAndChild(mstore, store, rm, OnDelete::Restrict);
+    installTtlHook(mstore, rm, store, p.pm, cfgManager);
+
+    check(mstore.expireDocuments() == 0, "the restrict-protected parent is blocked");
+    check(mstore.ttlBlockedDocumentCount() == 1, "and is counted as stuck");
+
+    // The collection goes away underneath it.
+    mstore.dropCollection("default:workflows");
+    check(mstore.expireDocuments() == 0, "nothing left to expire");
+    check(mstore.ttlBlockedDocumentCount() == 0,
+          "the blocked entry was released with the collection - the gauge does not "
+          "keep counting a document that no longer exists");
+
+    p.pm.stop();
+    mstore.stop();
+}
+
 // =========================================================================
 // v2.11.0 close-out — BOOT-PASS DRIFT MUST NOT DISABLE DESTRUCTIVE POLICIES.
 //
@@ -2891,6 +2927,7 @@ int main() {
     test_ttl_cascade_wal_is_durable_before_the_lmdb_commit();
     test_ttl_concurrent_ttl_extension_is_not_expired();
     test_ttl_blocked_documents_do_not_starve_the_sweep();
+    test_ttl_blocked_entry_is_released_when_the_collection_is_dropped();
     test_boot_pass_drift_does_not_disable_the_cascade();
     test_pending_remirror_list_is_released_after_the_pass();