ソースを参照

fix(relations): close-out review round 2 - TTL race, sweep starvation, 3 minors

Two Important findings, three Minors, plus one defect the new e2e phase exposed.
Each fix reverted in isolation and re-run to confirm the test fails without it.

1. IMPORTANT - a concurrent TTL change could have its document AND its children
   deleted by an in-flight sweep. Phase 2 handed the candidate to the hook with
   no re-check, and `Handled` cannot re-check the way `Proceed` does (the hook
   has no view of expiresAt), so an update()/patch() extending or clearing the
   TTL between phase 1 and the hook lost the parent and cascaded its children.
   The pre-v2.11.0 loop held the collection lock across the whole erase and had
   no such window, so this was a cost of the two-phase split. Now re-checked
   under the collection lock immediately before the hook. RESIDUAL, stated in
   the report: a write landing during the cascade's own WAL fsync + LMDB commit
   still loses; closing that needs an expiry claim or routing expiry through the
   versioned Delete path, which is a design decision, not a close-out fix.

2. IMPORTANT - a persistent `restrict` block starved expiry and flooded the log.
   A blocked document is never erased from the expiration index (deliberately -
   it must stay armed), so it refilled the per-sweep cap on every sweep and,
   `collections_` order being stable, silently stopped every collection after it
   from being swept at all; each blocked document also cost one LMDB cursor scan
   plus one WARN per sweep, i.e. N per second forever at the 1s default. Blocked
   ids are now remembered with a retry sweep number, skipped for free until due,
   drawn against their own budget, and logged once per document plus a periodic
   summary. The four thresholds moved to MemoryStore::Config.

Minors: a cascade-expiry no longer counts as both an expiry and a delete; the T7
self-heal log no longer reports the expected post-migration index build as the
"sub-db was absent" fault (applyRelationDeclarations gained an afterMigrations
flag affecting logging only); and Stats::ttlExpiryBlockedByRelation is documented
as the refusal-EVENT rate it is, with a new ttlBlockedDocumentCount() gauge for
"how many are stuck right now".

Found while writing finding 2's e2e phase: docStore() is projects_->get(), which
does not create an env, so a create_relation migration in a project never written
to was skipped by the post-migration re-arm and enforced nothing until the next
restart - the exact state the op was added to avoid, narrowed to new projects.
The post-migration pass now uses getOrCreate(); the boot pass keeps get() so it
cannot resurrect a stale project's env.

test_relation_enforcement 336 -> 364; relation_manager 70, document_ttl 22,
subdb_identity 264, dual_write_mirror 73, document_store green. E2E
test_relations.sh passes with a new create_relation-migration phase. VERSION
untouched. Report appended: closeout-report.md
fszontagh 1 ヶ月 前
親
コミット
9e28219757

+ 47 - 11
service/src/database_service.cpp

@@ -208,8 +208,9 @@ bool DatabaseService::initialize() {
         // pre-v2.11.0 (which is also what every unit fixture that never
         // installs the hook keeps doing).
         store_->setTtlExpiryRelationHook(
-            [this](const std::string& qualifiedCollection, const std::string& id) {
-                return ttlExpiryRelationDecision(qualifiedCollection, id);
+            [this](const std::string& qualifiedCollection, const std::string& id,
+                   bool firstAttempt) {
+                return ttlExpiryRelationDecision(qualifiedCollection, id, firstAttempt);
             });
 
         // v2.9.0 — re-apply persisted index declarations to each project's LMDB
@@ -487,7 +488,7 @@ bool DatabaseService::runMigrations() {
     // identical list, and the bootstrap loop skips any relation whose index
     // sub-db already exists. Runs even when the migration run reported failure,
     // because a partially-applied run may still have created a relation.
-    applyRelationDeclarations();
+    applyRelationDeclarations(/*afterMigrations=*/true);
     return ok;
 }
 
@@ -543,7 +544,7 @@ void DatabaseService::applyIndexDeclarations() {
 }
 
 MemoryStore::TtlExpiryAction DatabaseService::ttlExpiryRelationDecision(
-    const std::string& qualifiedCollection, const std::string& id) {
+    const std::string& qualifiedCollection, const std::string& id, bool firstAttempt) {
     using Action = MemoryStore::TtlExpiryAction;
 
     // Thin binder over relations/relation_cascade.cpp's ttlExpiryDecision() -
@@ -574,9 +575,13 @@ MemoryStore::TtlExpiryAction DatabaseService::ttlExpiryRelationDecision(
             notifyReplicationAndEvents(coll, docId, doc, eventType);
         };
 
+        // firstAttempt -> the refusal (if any) is logged at WARN; a retry of an
+        // already-reported document logs at DEBUG. The sweeper owns that
+        // decision because it is the only thing that knows the document has
+        // been refused before.
         return smartbotic::database::ttlExpiryDecision(
             *relation_manager_, *lmdb, *persistence_, *store_, *config_manager_,
-            qualifiedCollection, id, notify);
+            qualifiedCollection, id, notify, firstAttempt);
     } catch (const std::exception& e) {
         // Never expire a parent whose children could not be evaluated.
         spdlog::error("TTL expiry of '{}/{}': could not resolve its storage ({}) - the "
@@ -586,7 +591,7 @@ MemoryStore::TtlExpiryAction DatabaseService::ttlExpiryRelationDecision(
     }
 }
 
-void DatabaseService::applyRelationDeclarations() {
+void DatabaseService::applyRelationDeclarations(bool afterMigrations) {
     // Group by (project, bare child collection) — set_relations() is a
     // per-collection call on that project's LmdbDocumentStore and replaces
     // whatever was declared before, so every relation sharing a child
@@ -626,7 +631,22 @@ void DatabaseService::applyRelationDeclarations() {
     for (const auto& [key, refs] : byChild) {
         const auto& [project, collection] = keyToProjectCollection[key];
         try {
-            auto* ds = docStore(project);
+            // ⚠ v2.11.0 close-out — getOrCreate, not get(), on the POST-MIGRATION
+            // pass. docStore() is projects_->get(), which does NOT create an env,
+            // and a `create_relation` migration in a project that has never been
+            // written to has no env yet: the declaration persisted, this pass
+            // skipped it, and the relation maintained no reverse index and
+            // enforced NOTHING until the next restart - the exact
+            // "declared but not enforcing" state the op was added to avoid, just
+            // narrowed to new projects. A migration declaring a relation IS a
+            // statement that the project exists, and the mirror resolver would
+            // getOrCreate the same env on the first write anyway.
+            //
+            // The BOOT pass deliberately keeps get(): creating envs there would
+            // resurrect the env of a project whose declarations are merely stale.
+            auto* ds = (afterMigrations && projects_)
+                           ? projects_->getOrCreate(project)
+                           : docStore(project);
             auto* lmdb = dynamic_cast<smartbotic::db::storage::LmdbDocumentStore*>(ds);
             if (lmdb == nullptr) continue;
             lmdb->set_relations(collection, refs);
@@ -644,14 +664,30 @@ void DatabaseService::applyRelationDeclarations() {
             // is no dropRelation-only-the-index tool). CreateRelation's own
             // bootstrap scan (see database_grpc_impl.cpp) covers the normal
             // declare-time path; this covers everything else.
+            //
+            // ⚠ v2.11.0 close-out (review minor 2) — the SAME code path is the
+            // NORMAL, expected one for a relation a migration just declared:
+            // migrations run after this pass, so runMigrations() calls it again
+            // and the brand-new relation legitimately has no sub-db yet. Logging
+            // "declared but its sub-db was absent (restored snapshot or a manual
+            // removal)" at WARN there describes a fault that did not happen.
+            // `afterMigrations` distinguishes the two so the text and the level
+            // match the event.
             for (const auto& ref : refs) {
                 if (lmdb->relation_index_exists(ref.name)) continue;
                 const uint64_t rows =
                     lmdb->build_relation_index(ref.name, collection, ref.childField);
-                spdlog::warn("v2.11 relations: rebuilt '{}' over {} row(s) in '{}:{}' - the "
-                             "relation was declared but its reverse index sub-db was absent "
-                             "(restored snapshot predating it, or a manual removal)",
-                             ref.name, rows, project, collection);
+                if (afterMigrations) {
+                    spdlog::info("v2.11 relations: built the reverse index for '{}' over {} "
+                                 "existing row(s) in '{}:{}' - a migration declared this "
+                                 "relation on this boot",
+                                 ref.name, rows, project, collection);
+                } else {
+                    spdlog::warn("v2.11 relations: rebuilt '{}' over {} row(s) in '{}:{}' - the "
+                                 "relation was declared but its reverse index sub-db was absent "
+                                 "(restored snapshot predating it, or a manual removal)",
+                                 ref.name, rows, project, collection);
+                }
             }
         } catch (const std::exception& e) {
             spdlog::error("v2.11 relations: could not apply declarations for '{}:{}': {}",

+ 6 - 2
service/src/database_service.hpp

@@ -368,7 +368,7 @@ private:
     // must never throw (the sweeper catches, but a throw per document would
     // stop expiry making progress).
     MemoryStore::TtlExpiryAction ttlExpiryRelationDecision(
-        const std::string& qualifiedCollection, const std::string& id);
+        const std::string& qualifiedCollection, const std::string& id, bool firstAttempt);
 
     // v2.11.0 T6a — arm each project's LmdbDocumentStore with the relation
     // declarations loaded by relation_manager_->loadFromStore(), grouped by
@@ -378,7 +378,11 @@ private:
     // boot. Task 6b's createRelation/dropRelation RPCs must call
     // set_relations() again on live mutation; this only covers what was
     // already persisted at startup.
-    void applyRelationDeclarations();
+    // `afterMigrations` only affects LOGGING (see the self-heal block): a
+    // relation whose reverse index sub-db is absent is a fault at boot and the
+    // expected state for a relation a migration just declared, and the same code
+    // handles both.
+    void applyRelationDeclarations(bool afterMigrations = false);
 
     // v2.3 Stage C — atomic rename of <dataDir>/env/ into
     // <dataDir>/projects/default/env/ when the v2.2 layout is detected

+ 144 - 17
service/src/memory_store.cpp

@@ -2212,6 +2212,7 @@ void MemoryStore::setTtlExpiryRelationHook(TtlExpiryRelationHook hook) {
 uint64_t MemoryStore::expireDocuments() {
     uint64_t expired = 0;
     auto now = currentTimeMs();
+    const uint64_t sweep = ++ttlSweepCounter_;
 
     // ---- Phase 1: collect candidates under the locks, mutate nothing. -----
     // See the header comment on expireDocuments() for why the phases are split:
@@ -2226,43 +2227,136 @@ uint64_t MemoryStore::expireDocuments() {
         std::string collection;
         std::string id;
         uint64_t expiresAt;
+        bool isBlockedRetry = false;
     };
-    // ⚠ BOUNDED PER SWEEP. Collecting the whole expired set first is what makes
-    // the lock-free second phase possible, but an unbounded vector here is the
-    // same retention class the v2.11.0 re-mirror finding was about: a backlog of
-    // millions of expired documents (a long outage, or a TTL applied
-    // retroactively) would be ~80-100 bytes each, resident at once, on a
-    // background thread. Whatever this sweep does not reach stays armed in the
-    // expiration index and is picked up by the next one, so the only cost of the
-    // cap is latency on a backlog that is already late.
-    constexpr size_t kMaxCandidatesPerSweep = 10000;
+    // ⚠ BOUNDED PER SWEEP, WITH A SEPARATE BUDGET FOR RETRIES OF BLOCKED
+    // DOCUMENTS. Collecting the expired set up front is what makes the
+    // lock-free second phase possible, but an unbounded vector is the same
+    // retention class the v2.11.0 re-mirror finding was about: a backlog of
+    // millions of expired rows at ~80-100 bytes each, resident at once, on a
+    // background thread.
+    //
+    // ⚠ AND THE TWO BUDGETS ARE SEPARATE FOR A REASON (close-out review,
+    // finding 2). A document a `restrict` relation blocks is never erased from
+    // the expiration index - that is deliberate, it has to stay armed so the
+    // block can lift - so with ONE budget those documents refill the cap on
+    // every sweep, and because `collections_` iteration order is stable, every
+    // collection after them in the map SILENTLY STOPS BEING SWEPT. TTL'd
+    // parents under `restrict` is precisely the combination this feature
+    // creates, so that is the expected steady state, not an edge case. Blocked
+    // ids therefore live in ttlBlocked_ with a next-retry sweep number, are
+    // skipped entirely until due, and when due are drawn against their OWN
+    // small budget - so fresh expiries can never be starved by them.
+    // Budgets come from Config so the behaviour is testable at small numbers;
+    // a zero is coerced to 1 so a misconfiguration cannot stop expiry entirely.
+    const size_t kMaxCandidatesPerSweep =
+        std::max<size_t>(1, config_.ttlMaxCandidatesPerSweep);
+    const size_t kMaxBlockedRetriesPerSweep =
+        std::max<size_t>(1, config_.ttlMaxBlockedRetriesPerSweep);
+    const uint64_t kBlockedRetrySweeps =
+        std::max<uint64_t>(1, config_.ttlBlockedRetrySweeps);
+    const uint64_t kBlockedSummarySweeps =
+        std::max<uint64_t>(1, config_.ttlBlockedSummarySweeps);
+
     std::vector<Candidate> candidates;
+    size_t fresh = 0;
+    size_t retries = 0;
+    size_t blockedNotDue = 0;
     {
+        std::lock_guard<std::mutex> blockedLock(ttlBlockedMutex_);
         std::shared_lock<std::shared_mutex> globalLock(globalMutex_);
         for (auto& [collName, coll] : collections_) {
-            if (candidates.size() >= kMaxCandidatesPerSweep) break;
+            if (fresh >= kMaxCandidatesPerSweep) break;
             std::shared_lock<std::shared_mutex> collLock(coll->mutex);
             for (auto it = coll->expirationIndex.begin();
                  it != coll->expirationIndex.end() && it->first <= now; ++it) {
                 for (const auto& id : it->second) {
-                    candidates.push_back(Candidate{collName, id, it->first});
+                    const auto blockedIt = ttlBlocked_.find(ttlBlockedKey(collName, id));
+                    if (blockedIt != ttlBlocked_.end()) {
+                        // Known blocked. Costs NOTHING (no LMDB read txn, no
+                        // cursor scan, no log line) until its retry is due, and
+                        // never consumes the fresh budget.
+                        if (blockedIt->second > sweep) { ++blockedNotDue; continue; }
+                        if (retries >= kMaxBlockedRetriesPerSweep) continue;
+                        ++retries;
+                        candidates.push_back(Candidate{collName, id, it->first, true});
+                        continue;
+                    }
+                    if (fresh >= kMaxCandidatesPerSweep) break;
+                    ++fresh;
+                    candidates.push_back(Candidate{collName, id, it->first, false});
                 }
-                if (candidates.size() >= kMaxCandidatesPerSweep) break;
+                if (fresh >= kMaxCandidatesPerSweep) break;
             }
         }
     }
+
+    // Periodic summary instead of a per-document WARN per sweep. A flood is not
+    // a signal, and at the 1s default a persistent block was ~one WARN per
+    // blocked document per second, indefinitely.
+    if (blockedNotDue > 0 && sweep - ttlBlockedSummarySweep_ >= kBlockedSummarySweeps) {
+        ttlBlockedSummarySweep_ = sweep;
+        spdlog::warn("TTL expiry: {} document(s) are still past their TTL and cannot be "
+                     "expired because a relation blocks the delete (each was logged once "
+                     "when first blocked; they are retried every {} sweeps). They will "
+                     "remain until the referencing children are removed or re-pointed.",
+                     blockedNotDue, kBlockedRetrySweeps);
+    }
+
     if (candidates.empty()) return 0;
 
     // ---- Phase 2: one document at a time, no lock held on entry. ----------
     for (const auto& cand : candidates) {
+        // ⚠ RE-CHECK BEFORE THE HOOK (close-out review, finding 1). The hook may
+        // DELETE the document and cascade its children, and it has no view of
+        // expiresAt, so `Handled` cannot re-check the way `Proceed` does below.
+        // A concurrent update()/patch() that extends or clears the TTL between
+        // phase 1 and here would otherwise have the parent AND its children
+        // deleted despite the document no longer being expired - the
+        // pre-v2.11.0 loop held the collection lock across the whole erase and
+        // had no such window, so this is a cost of the two-phase split and it
+        // is paid here. Cheap: one lock, one hash lookup, no document copy.
+        //
+        // See the header comment for the RESIDUAL: a write landing during the
+        // hook's own cascade (a WAL fsync plus an LMDB commit) is still
+        // possible. Narrowing, not eliminating.
+        uint64_t currentExpiresAt = 0;
+        {
+            std::shared_lock<std::shared_mutex> globalLock(globalMutex_);
+            auto collIt = collections_.find(cand.collection);
+            if (collIt == collections_.end()) continue;   // collection dropped
+            auto* coll = collIt->second.get();
+            std::unique_lock<std::shared_mutex> collLock(coll->mutex);
+            auto docIt = coll->documents.find(cand.id);
+            if (docIt == coll->documents.end()) {
+                // Gone (deleted, or evicted). Drop the stale index entry so it
+                // is not re-examined every sweep forever.
+                removeFromExpirationIndex(*coll, cand.id, cand.expiresAt);
+                clearTtlBlocked(cand.collection, cand.id);
+                continue;
+            }
+            currentExpiresAt = docIt->second.expiresAt;
+        }
+        if (currentExpiresAt == 0 || currentExpiresAt > now) {
+            // The TTL was extended or cleared while this sweep was running. Not
+            // expired any more - and no longer blocked either, if it was.
+            clearTtlBlocked(cand.collection, cand.id);
+            continue;
+        }
+
         // The hook decides what this document's children require. Unset (no
         // relations wired at all) is the pre-v2.11.0 path, unchanged; the hook
         // itself early-returns Proceed when the declaration map holds nothing
         // for this collection, so the common case costs one map lookup.
+        //
+        // `firstAttempt` is false on a retry of an already-blocked document, so
+        // the hook can log the refusal ONCE per document rather than once per
+        // document per sweep.
         TtlExpiryAction action = TtlExpiryAction::Proceed;
         if (ttlExpiryRelationHook_) {
             try {
-                action = ttlExpiryRelationHook_(cand.collection, cand.id);
+                action = ttlExpiryRelationHook_(cand.collection, cand.id,
+                                                !cand.isBlockedRetry);
             } catch (const std::exception& e) {
                 // The sweeper is a background thread: an escaping exception
                 // would kill it and stop ALL expiry for the process. Skip this
@@ -2277,12 +2371,20 @@ uint64_t MemoryStore::expireDocuments() {
 
         if (action == TtlExpiryAction::Skip) {
             // Nothing to undo: phase 1 erased nothing, so the expiration index
-            // entry is still armed and the next sweep sees this candidate again.
+            // entry is still armed. Record the block so the next sweeps skip it
+            // for free until its retry is due.
+            {
+                std::lock_guard<std::mutex> blockedLock(ttlBlockedMutex_);
+                ttlBlocked_[ttlBlockedKey(cand.collection, cand.id)] =
+                    sweep + kBlockedRetrySweeps;
+            }
             std::lock_guard<std::mutex> statsLock(statsMutex_);
             stats_.ttlExpiryBlockedByRelation++;
             continue;
         }
 
+        clearTtlBlocked(cand.collection, cand.id);
+
         if (action == TtlExpiryAction::Handled) {
             // The hook ran the whole cascade: WAL, LMDB commit, MemoryStore
             // apply and per-mutation replication/events. The document is gone
@@ -2293,6 +2395,12 @@ uint64_t MemoryStore::expireDocuments() {
             {
                 std::lock_guard<std::mutex> statsLock(statsMutex_);
                 stats_.expiredCount++;
+                // The cascade removed the parent through unloadDocument(), which
+                // counted it as a DELETE. It is an expiry, not a delete, and
+                // counting it as both made a cascade-expiry double-counted
+                // (close-out review, minor 1). The children's deletes stay
+                // counted - those really are deletes.
+                if (stats_.deleteCount > 0) stats_.deleteCount--;
             }
             // Defensive: applyCascadeToMemory -> unloadDocument drops the
             // expiration index entry along with the document, but only if the
@@ -2330,9 +2438,11 @@ uint64_t MemoryStore::expireDocuments() {
                 removeFromExpirationIndex(*coll, cand.id, cand.expiresAt);
                 continue;
             }
-            // Re-check the expiry under the lock: between phase 1 and here an
-            // update may have extended or cleared the TTL, and expiring on a
-            // stale reading would delete a document that is no longer expired.
+            // Re-check the expiry under the lock: between the pre-check above
+            // and here an update may have extended or cleared the TTL, and
+            // expiring on a stale reading would delete a document that is no
+            // longer expired. This one IS airtight - the check and the erase
+            // share a single critical section.
             if (docIt->second.expiresAt == 0 || docIt->second.expiresAt > now) {
                 continue;
             }
@@ -2363,6 +2473,23 @@ uint64_t MemoryStore::expireDocuments() {
     return expired;
 }
 
+std::string MemoryStore::ttlBlockedKey(const std::string& collection, const std::string& id) {
+    // NUL separator: collection names cannot contain one, so the key cannot be
+    // ambiguous between ("a", "b:c") and ("a:b", "c").
+    return collection + std::string(1, '\0') + id;
+}
+
+void MemoryStore::clearTtlBlocked(const std::string& collection, const std::string& id) {
+    std::lock_guard<std::mutex> lock(ttlBlockedMutex_);
+    if (ttlBlocked_.empty()) return;
+    ttlBlocked_.erase(ttlBlockedKey(collection, id));
+}
+
+size_t MemoryStore::ttlBlockedDocumentCount() const {
+    std::lock_guard<std::mutex> lock(ttlBlockedMutex_);
+    return ttlBlocked_.size();
+}
+
 // ===== Statistics =====
 
 // Estimate JSON object size without serializing (fast approximation)

+ 87 - 11
service/src/memory_store.hpp

@@ -112,6 +112,28 @@ public:
         // the estimator is wrong, so it is an ERROR, not a warning.
         uint32_t evictionMaxEpisodePercent = 50;
 
+        // v2.11.0 close-out (review finding 2) — TTL sweep budgets.
+        //
+        // `ttlMaxCandidatesPerSweep` bounds how many freshly-expired documents
+        // one sweep collects, because the two-phase sweeper materialises the
+        // candidate list before releasing its locks and an unbounded list is a
+        // retention hazard on a large backlog.
+        //
+        // The other two exist because a document a relation refuses to let
+        // expire is NEVER erased from the expiration index - it has to stay
+        // armed so the block can lift - so without separate accounting those
+        // documents refill the fresh budget on every sweep and silently starve
+        // every collection after them in iteration order. Blocked ids are
+        // remembered, skipped for free until `ttlBlockedRetrySweeps` sweeps have
+        // passed, and then retried against their own small budget.
+        //
+        // Tunable mainly so the behaviour is testable at small numbers; the
+        // defaults are what production runs.
+        uint32_t ttlMaxCandidatesPerSweep = 10000;
+        uint32_t ttlMaxBlockedRetriesPerSweep = 100;
+        uint32_t ttlBlockedRetrySweeps = 60;      // ~1 min at the 1s default
+        uint32_t ttlBlockedSummarySweeps = 300;   // ~5 min at the 1s default
+
         Config()
             : maxMemoryBytes(800ULL * 1024 * 1024)   // 800 MB default
             , expirationCheckIntervalMs(1000)         // Check TTL every second
@@ -700,11 +722,15 @@ public:
         // The expiry must NOT happen: a `restrict` relation blocks it (a manual
         // delete would have failed with FAILED_PRECONDITION), or the cascade
         // machinery refused because the LMDB mirror is unhealthy/drifted.
-        // ⚠ The document is left in place WITH ITS EXPIRY RE-ARMED, so the next
+        // ⚠ The document is left in place WITH ITS EXPIRY STILL ARMED, so a later
         // sweep retries - which means it OUTLIVES ITS TTL for as long as the
-        // block lasts. That is a retention-policy surprise, so each skip logs a
-        // WARN naming the relation and the blocking child count, and
-        // Stats::ttlExpiryBlockedByRelation counts them.
+        // block lasts. That is a retention-policy surprise, so the first refusal
+        // per document logs a WARN naming the relation and the blocking child
+        // count, a periodic summary line reports how many are still stuck, and
+        // Stats::ttlExpiryBlockedByRelation counts refusal events.
+        // The id is remembered (expireDocuments()' ttlBlocked_) so subsequent
+        // sweeps skip it for free - no LMDB read, no log - until its retry is
+        // due, and so it cannot consume the sweep's fresh-candidate budget.
         Skip,
         // The hook already performed the whole delete through the ordinary
         // cascade machinery (relations/relation_cascade.cpp: WAL for the parent
@@ -714,9 +740,16 @@ public:
         // would double-log the delete to the WAL and re-mirror it.
         Handled
     };
+    // `firstAttempt` is false when this document has already been refused by a
+    // previous sweep and is being retried (see expireDocuments()' blocked set).
+    // It exists so the hook can log its refusal ONCE per document instead of
+    // once per document per sweep: at the 1s default sweep interval, a
+    // persistent `restrict` block on N parents was N WARN lines per second,
+    // indefinitely, and a flood is not a signal.
     using TtlExpiryRelationHook =
         std::function<TtlExpiryAction(const std::string& qualifiedCollection,
-                                      const std::string& id)>;
+                                      const std::string& id,
+                                      bool firstAttempt)>;
 
     /**
      * Install the hook above. DatabaseService does this once, at boot, after
@@ -745,10 +778,28 @@ public:
      * is not upgradable - and would deadlock outright against a request thread
      * whenever the child collection and the parent collection differed.
      *
+     * ⚠ THE RESIDUAL RACE, stated because it is not fully closed (close-out
+     * review, finding 1). Phase 2 re-checks residency and `expiresAt` under the
+     * collection lock IMMEDIATELY before invoking the hook, and the `Proceed`
+     * path re-checks again in the same critical section as its erase (that one
+     * is airtight). A `Handled` cascade cannot: it spans a WAL fsync and an LMDB
+     * commit and the hook has no view of `expiresAt`, so a write that extends or
+     * clears the TTL DURING the cascade still loses. Closing that needs the
+     * expiry decision and the cascade to share one atomic unit, which nothing in
+     * the current design provides - see the close-out report.
+     *
      * @return Number of documents actually expired (Skip does not count)
      */
     uint64_t expireDocuments();
 
+    /**
+     * v2.11.0 close-out — how many documents are currently past their TTL and
+     * NOT being expired because a relation blocks the delete. A GAUGE (it falls
+     * when a block lifts), unlike Stats::ttlExpiryBlockedByRelation, which
+     * counts refusal events. Zero on any install that declares no relations.
+     */
+    [[nodiscard]] size_t ttlBlockedDocumentCount() const;
+
     // ===== Statistics =====
 
     struct Stats {
@@ -756,12 +807,18 @@ public:
         uint64_t totalCollections = 0;
         uint64_t estimatedMemoryBytes = 0;
         uint64_t expiredCount = 0;
-        // v2.11.0 close-out — TTL expiries the sweeper REFUSED because a
-        // relation blocked them (a `restrict` relation with live children, or
-        // the cascade path refusing while the LMDB mirror is unhealthy/drifted).
-        // Each one means a document is still present PAST its TTL and will be
-        // retried on the next sweep. Nonzero is an operator signal, not an
-        // error: the alternative was silently orphaning the children.
+        // v2.11.0 close-out — TTL expiry REFUSAL EVENTS, because a relation
+        // blocked the delete (a `restrict` relation with live children, or the
+        // cascade path refusing while the LMDB mirror is unhealthy/drifted).
+        //
+        // ⚠ EVENTS, NOT DISTINCT DOCUMENTS - it is a rate, and it keeps growing
+        // while a block persists, because a blocked document is retried
+        // (every ~60 sweeps; see expireDocuments()). For "how many documents are
+        // stuck past their TTL right now", which is the gauge an operator
+        // actually wants, use MemoryStore::ttlBlockedDocumentCount().
+        //
+        // Nonzero is an operator signal, not an error: the alternative was
+        // silently orphaning the children.
         uint64_t ttlExpiryBlockedByRelation = 0;
         uint64_t insertCount = 0;
         uint64_t updateCount = 0;
@@ -1267,6 +1324,25 @@ private:
     // from the cleanup thread only.
     TtlExpiryRelationHook ttlExpiryRelationHook_;
 
+    // v2.11.0 close-out (review finding 2) — documents a relation is currently
+    // refusing to let expire, mapped to the sweep number at which to retry.
+    // Exists so a persistent `restrict` block costs nothing per sweep (no LMDB
+    // read txn, no cursor scan, no log line) and, crucially, cannot refill the
+    // per-sweep candidate cap - which would silently stop every collection after
+    // it in `collections_` iteration order from being swept at all.
+    //
+    // Guarded by its own mutex rather than a collection lock: it is keyed across
+    // collections, and it is read in phase 1 while globalMutex_ is held shared,
+    // so it must be the INNERMOST lock. Nothing taken under it takes any other
+    // lock.
+    mutable std::mutex ttlBlockedMutex_;
+    std::unordered_map<std::string, uint64_t> ttlBlocked_;
+    uint64_t ttlSweepCounter_ = 0;
+    uint64_t ttlBlockedSummarySweep_ = 0;
+
+    static std::string ttlBlockedKey(const std::string& collection, const std::string& id);
+    void clearTtlBlocked(const std::string& collection, const std::string& id);
+
     std::atomic<bool>* mirrorHealthy_ = nullptr;
     std::atomic<uint64_t>* mirrorDriftCount_ = nullptr;
     // v2.11.0 close-out — drift already accrued when the boot path finished.

+ 19 - 8
service/src/relations/relation_cascade.cpp

@@ -5,6 +5,7 @@
 #include <unordered_map>
 
 #include <spdlog/spdlog.h>
+#include <spdlog/fmt/fmt.h>
 
 #include "../config/collection_config_manager.hpp"
 #include "../memory_store.hpp"
@@ -395,9 +396,17 @@ MemoryStore::TtlExpiryAction ttlExpiryDecision(
     CollectionConfigManager& configManager,
     const std::string& qualifiedParentCollection,
     const std::string& parentId,
-    const CascadeNotifyFn& notify) {
+    const CascadeNotifyFn& notify,
+    bool logBlockAtWarn) {
     using Action = MemoryStore::TtlExpiryAction;
 
+    // One helper for both refusal sites, so the two cannot drift in either text
+    // or level. See the header: the level is a volume decision only.
+    const auto logBlock = [logBlockAtWarn](const std::string& msg) {
+        if (logBlockAtWarn) spdlog::warn("{}", msg);
+        else spdlog::debug("{}", msg);
+    };
+
     // ---- Scope: the no-relations case must cost exactly one map lookup. ----
     // Same early-out shape the write path uses. Nothing below this line runs
     // for a collection that is nobody's parent, which is every collection on
@@ -424,10 +433,11 @@ MemoryStore::TtlExpiryAction ttlExpiryDecision(
             // OUTLIVES ITS TTL, indefinitely, for as long as a child keeps
             // referencing it. That is the accepted price of not orphaning the
             // children, and the alternative is what this change removes.
-            spdlog::warn("TTL expiry of '{}/{}' is BLOCKED by a restrict relation, so the "
-                         "document remains past its TTL and will be retried on the next "
-                         "sweep (MemoryStore stat ttlExpiryBlockedByRelation counts these): "
-                         "{}", qualifiedParentCollection, parentId, err);
+            logBlock(fmt::format(
+                "TTL expiry of '{}/{}' is BLOCKED by a restrict relation, so the document "
+                "remains past its TTL and will be retried on a slower cadence "
+                "(MemoryStore::ttlBlockedDocumentCount() is how many are stuck right now): "
+                "{}", qualifiedParentCollection, parentId, err));
             return Action::Skip;
         }
 
@@ -454,9 +464,10 @@ MemoryStore::TtlExpiryAction ttlExpiryDecision(
             // (both are thrown before writeCascadeWal()), so skipping is a
             // clean no-op and the next sweep retries. Expiring the parent
             // anyway is exactly the orphaning this function exists to remove.
-            spdlog::warn("TTL expiry of '{}/{}' skipped: {} - the document remains past "
-                         "its TTL and will be retried on the next sweep",
-                         qualifiedParentCollection, parentId, e.what());
+            logBlock(fmt::format(
+                "TTL expiry of '{}/{}' skipped: {} - the document remains past its TTL and "
+                "will be retried on a slower cadence",
+                qualifiedParentCollection, parentId, e.what()));
             return Action::Skip;
         }
     } catch (const std::exception& e) {

+ 11 - 1
service/src/relations/relation_cascade.hpp

@@ -403,6 +403,15 @@ void applyCascadeToMemory(MemoryStore& store,
 // (applyCascadeToMemory takes each affected collection's lock, and
 // getOrCreateCollection takes globalMutex_ exclusively). That is exactly what
 // expireDocuments()' two-phase structure guarantees.
+//
+// `logBlockAtWarn` controls the VOLUME of the refusal log, never the decision:
+// true logs the block at WARN with the blocking relation and child count (what
+// the first refusal for a document does), false logs the same text at DEBUG
+// (what a retry of an already-reported document does). At the sweeper's 1s
+// default, a persistent `restrict` block on N parents was N WARN lines per
+// second forever, which buries every other operator signal. The periodic
+// "still stuck" summary comes from expireDocuments() itself, which is the only
+// place that knows the total.
 MemoryStore::TtlExpiryAction ttlExpiryDecision(
     RelationManager& relations,
     smartbotic::db::storage::LmdbDocumentStore& store,
@@ -411,7 +420,8 @@ MemoryStore::TtlExpiryAction ttlExpiryDecision(
     CollectionConfigManager& configManager,
     const std::string& qualifiedParentCollection,
     const std::string& parentId,
-    const CascadeNotifyFn& notify = nullptr);
+    const CascadeNotifyFn& notify = nullptr,
+    bool logBlockAtWarn = true);
 
 bool executeCascade(RelationManager& relations,
                     smartbotic::db::storage::LmdbDocumentStore& store,

+ 67 - 1
tests/load_test/test_relations.sh

@@ -38,6 +38,31 @@ g++ -std=c++20 -O1 -o "$DROP_TOOL" "$REPO/tests/load_test/relidx_drop_tool.cpp"
     || fail "relidx_drop_tool did not compile"
 
 mkdir -p "$WORK/data"
+
+# v2.11.0 close-out — the create_relation MIGRATION OP, over a real boot.
+# Declaring schema in migration files is how consumers ship views, so this is the
+# surface that matters for relations too. Written before the first boot so the op
+# runs on boot 1, which is also the only boot on which the reverse index for it
+# does not yet exist - the case whose log text minor 2 was about.
+MIGPROJ=relproj_m
+mkdir -p "$WORK/migrations"
+cat > "$WORK/migrations/001_relation.json" <<EOF
+{
+  "version": "001",
+  "name": "declare_mig_rel",
+  "operations": [
+    {"type": "create_collection", "collection": "$MIGPROJ:workflows"},
+    {"type": "create_collection", "collection": "$MIGPROJ:executions"},
+    {"type": "create_relation",
+     "name": "$MIGPROJ:mig_rel",
+     "child": "$MIGPROJ:executions",
+     "child_field": "workflowId",
+     "parent": "$MIGPROJ:workflows",
+     "on_delete": "restrict"}
+  ]
+}
+EOF
+
 cat > "$WORK/config.json" <<EOF
 {
   "storage": {
@@ -45,7 +70,7 @@ cat > "$WORK/config.json" <<EOF
     "bind_address": "127.0.0.1",
     "rpc_port": $PORT,
     "encryption": { "enabled": false, "key_file": "$WORK/data/storage.key" },
-    "migrations": { "enabled": false },
+    "migrations": { "enabled": true, "directory": "$WORK/migrations" },
     "replication": { "enabled": false }
   }
 }
@@ -200,6 +225,47 @@ echo
 echo "=== phase: verify the self-heal rebuilt the index and enforcement works again ==="
 "$DRIVER" "127.0.0.1:$PORT" "$PROJECT_A" "$PROJECT_B" verify_selfheal || fail "verify_selfheal phase"
 
+echo
+echo "=== phase: the create_relation MIGRATION OP declared and ARMED a relation ==="
+# The op only persists the declaration; runMigrations() re-arms afterwards,
+# because migrations run AFTER the boot-time arming pass. Without that re-arm the
+# relation would maintain no reverse index and enforce nothing until the NEXT
+# restart - so this phase is the test of the ordering claim, not just the op.
+OUT="$(rpc GetRelationInfo "{\"name\":\"$MIGPROJ:mig_rel\"}" || true)"
+grep -q '"found": true' <<<"$OUT" \
+    || { echo "$OUT"; fail "the create_relation migration op did not declare the relation"; }
+echo "  declared by migration: $MIGPROJ:mig_rel"
+
+# Armed? Insert a parent and a child, then try to delete the parent. A relation
+# that was never armed maintains no posting, so the delete would SUCCEED - which
+# is exactly the silent failure mode.
+#
+# ⚠ This check alone cannot prove it was armed on the boot that DECLARED it: two
+# restarts have happened since, and the ordinary boot-time pass would have armed
+# it on either of them. The boot1.log assertion below is what pins that, and it
+# is the one that fails if the post-migration re-arm is removed.
+M_PARENT_B64="$(b64 '{"name":"wf-m1"}')"
+M_CHILD_B64="$(b64 '{"workflowId":"wf-m1"}')"
+rpc Upsert "{\"collection\":\"$MIGPROJ:workflows\",\"id\":\"wf-m1\",\"data\":\"$M_PARENT_B64\"}" >/dev/null
+rpc Upsert "{\"collection\":\"$MIGPROJ:executions\",\"id\":\"ex-m1\",\"data\":\"$M_CHILD_B64\"}" >/dev/null
+OUT="$(rpc Delete "{\"collection\":\"$MIGPROJ:workflows\",\"id\":\"wf-m1\"}" || true)"
+grep -q "FailedPrecondition" <<<"$OUT" \
+    || { echo "$OUT"; fail "a migration-declared relation is NOT enforcing - the post-migration re-arm did not happen"; }
+grep -q "mig_rel" <<<"$OUT" \
+    || { echo "$OUT"; fail "the refusal does not name the migration-declared relation"; }
+echo "  and it is enforcing: the parent delete was refused by $MIGPROJ:mig_rel"
+
+# minor 2: on the boot that first applies the migration, the reverse index
+# legitimately does not exist yet. That must NOT be logged as the
+# "declared but its sub-db was absent (restored snapshot or manual removal)"
+# fault - the text describes a fault and the event is the normal path.
+grep -q "a migration declared this relation on this boot" "$WORK/boot1.log" \
+    || { grep -iE "mig_rel" "$WORK/boot1.log" | tail -5
+         fail "the migration-declared relation's index build was not logged as the expected path"; }
+grep -E "rebuilt 'mig_rel'" "$WORK/boot1.log" \
+    && fail "the expected path was logged as the 'sub-db was absent' FAULT"
+echo "  boot log: $(grep 'a migration declared this relation' "$WORK/boot1.log" | tail -1 | sed 's/.*\] //')"
+
 echo
 echo "=== phase: DropProject must clean up its relation declarations (close-out) ==="
 # Declarations live in the GLOBAL _relations collection, so dropping a project

+ 262 - 14
tests/test_relation_enforcement.cpp

@@ -2115,8 +2115,10 @@ void installTtlHook(MemoryStore& mstore, RelationManager& rm, LmdbDocumentStore&
                     const smartbotic::database::CascadeNotifyFn& notify = nullptr) {
     mstore.setTtlExpiryRelationHook(
         [&rm, &store, &pm, &mstore, &cfgManager, notify](const std::string& coll,
-                                                          const std::string& id) {
-            return ttlExpiryDecision(rm, store, pm, mstore, cfgManager, coll, id, notify);
+                                                          const std::string& id,
+                                                          bool firstAttempt) {
+            return ttlExpiryDecision(rm, store, pm, mstore, cfgManager, coll, id, notify,
+                                     firstAttempt);
         });
 }
 
@@ -2168,7 +2170,10 @@ void test_ttl_restrict_blocks_the_expiry() {
     TmpPersistence p("ttl-restrict-wal");
     check(p.pm.start(), "persistence manager started");
 
-    MemoryStore mstore(MemoryStore::Config{});
+    // Small retry cadence so the test does not have to run 60 sweeps.
+    MemoryStore::Config cfg;
+    cfg.ttlBlockedRetrySweeps = 3;
+    MemoryStore mstore(cfg);
     mstore.start();
     RelationManager rm(mstore);
     rm.loadFromStore();
@@ -2188,21 +2193,30 @@ void test_ttl_restrict_blocks_the_expiry() {
           "the skip is counted, so an operator can see a document is stuck past its TTL");
     check(mstore.getStats().expiredCount == 0, "and it is not counted as expired");
 
-    // The expiry must stay ARMED so the next sweep retries - a skip that also
-    // dropped the expiration index entry would leave the document permanently
-    // unexpirable even after the last child went away.
+    check(mstore.ttlBlockedDocumentCount() == 1,
+          "and it is tracked as ONE stuck document - the gauge an operator wants");
+
+    // ⚠ The next sweep must NOT re-examine it (close-out review, finding 2). A
+    // blocked document costs one LMDB read txn + cursor scan + one log line every
+    // time it is examined, and at the 1s default that is a permanent flood plus a
+    // permanent index-lookup load. It is skipped for free until its retry is due.
     const uint64_t again = mstore.expireDocuments();
-    check(again == 0, "still blocked on the second sweep");
-    check(mstore.getStats().ttlExpiryBlockedByRelation == 2,
-          "the second sweep re-examined it, so the expiry is still armed");
+    check(again == 0, "still not expired on the next sweep");
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 1,
+          "and the relation was NOT re-consulted - no second refusal event");
+    check(residentInMemory(mstore, "default:workflows", "wf-1"), "the parent is still there");
 
-    // And once the child is gone, the same sweep expires it - the block is the
-    // relation's, not a permanent quarantine.
+    // The expiry must stay ARMED, so once the retry comes due AND the child is
+    // gone, it expires - the block is the relation's, not a permanent quarantine.
     store.del("executions", "ex-1");
     check(mstore.remove("default:executions", "ex-1"), "child removed");
-    const uint64_t third = mstore.expireDocuments();
-    check(third == 1, "with no children left, the retry finally expires the parent");
-    check(!mstore.get("default:workflows", "wf-1").has_value(), "the parent is gone now");
+    uint64_t total = 0;
+    for (uint32_t i = 0; i < cfg.ttlBlockedRetrySweeps + 1; ++i) {
+        total += mstore.expireDocuments();
+    }
+    check(total == 1, "when the retry comes due, the parent finally expires");
+    check(!residentInMemory(mstore, "default:workflows", "wf-1"), "the parent is gone now");
+    check(mstore.ttlBlockedDocumentCount() == 0, "and it is no longer counted as stuck");
 
     p.pm.stop();
     mstore.stop();
@@ -2250,6 +2264,15 @@ void test_ttl_cascade_deletes_children_like_a_manual_delete() {
     check(notifications.size() == 2,
           "replication/events fired once per child plus once for the parent");
 
+    // Accounting (close-out review, minor 1): the parent is ONE expiry, and the
+    // one child is ONE delete. The cascade removes the parent through
+    // unloadDocument(), which counts a DELETE, so without compensating for that
+    // the parent was counted as both an expiry and a delete.
+    const auto st = mstore.getStats();
+    check(st.expiredCount == 1, "the parent counted as exactly one expiry");
+    check(st.deleteCount == 1,
+          "and the delete count covers the CHILD only - the parent is not counted twice");
+
     // The property a hand-rolled sweeper cascade breaks: MemoryStore is rebuilt
     // from snapshot + WAL, NEVER from LMDB, so a cascade whose child deletions
     // exist only in LMDB has them RESURRECTED on the next boot.
@@ -2497,6 +2520,229 @@ void test_ttl_cascade_wal_is_durable_before_the_lmdb_commit() {
     mstore.stop();
 }
 
+
+// =========================================================================
+// v2.11.0 close-out review, FINDING 1 — a concurrent TTL change must not have
+// its document (and its children) deleted by an in-flight sweep.
+//
+// Phase 1 collects candidates and releases every lock; the hook may then DELETE
+// the document and cascade its children, and the hook has no view of
+// `expiresAt`, so `Handled` cannot re-check the way `Proceed` does. An
+// update()/patch() that extends or clears the TTL in that window would
+// otherwise destroy a document that is no longer expired - and its children with
+// it. The pre-v2.11.0 loop held the collection lock across the whole erase and
+// had no such window, so this is a cost of the two-phase split.
+//
+// ⚠ HOW THIS IS INDUCED, and why it is honest rather than staged: a
+// single-threaded test cannot interleave a real writer, so the write is made to
+// land at exactly the wrong moment from INSIDE the sweep. Two documents expire
+// in the same sweep, wf-early before wf-late (the expiration index is keyed by
+// expiry time and walked in ascending order, so the order is deterministic).
+// The hook, while handling wf-early, extends wf-late's TTL - which is precisely
+// "a write landed after phase 1 collected wf-late and before the sweeper got to
+// it". Everything about wf-late's path after that point is the production path.
+void test_ttl_concurrent_ttl_extension_is_not_expired() {
+    TmpEnv t("ttl-race");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-race-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    RelationInfo rel;
+    rel.name = "default:exec_wf";
+    rel.child = "default:executions";
+    rel.childField = "workflowId";
+    rel.parent = "default:workflows";
+    rel.onDelete = OnDelete::Cascade;
+    std::string mgrErr;
+    check(rm.createRelation(rel, mgrErr), "declared the cascade relation");
+    store.set_relations("executions",
+                        {RelationRef{"exec_wf", "workflowId", "workflows", false, true}});
+
+    const uint64_t farFuture = static_cast<uint64_t>(1) << 62;
+
+    auto seedParent = [&](const std::string& id, uint64_t expiresAt) {
+        Document d; d.id = id; d.collection = "workflows";
+        d.set_data({{"name", id}});
+        d.expiresAt = expiresAt;
+        store.put("workflows", id, d);
+        mstore.loadDocument("default:workflows", d);
+    };
+    auto seedChild = [&](const std::string& id, const std::string& parentId) {
+        Document d; d.id = id; d.collection = "executions";
+        d.set_data({{"workflowId", parentId}});
+        store.put("executions", id, d);
+        mstore.loadDocument("default:executions", d);
+    };
+
+    // expiresAt 1 sorts before 2, so wf-early is handled first in the sweep.
+    seedParent("wf-early", 1);
+    seedParent("wf-late", 2);
+    seedChild("ex-early", "wf-early");
+    seedChild("ex-late", "wf-late");
+
+    // The injected "concurrent" write: while the sweeper is busy expiring
+    // wf-early (a real cascade - WAL fsync plus an LMDB commit, which is what
+    // makes the window wide), wf-late's TTL is extended.
+    bool injected = false;
+    mstore.setTtlExpiryRelationHook(
+        [&](const std::string& coll, const std::string& id, bool firstAttempt) {
+            if (!injected && id == "wf-early") {
+                injected = true;
+                Document renewed;
+                renewed.id = "wf-late";
+                renewed.collection = "workflows";
+                renewed.set_data({{"name", "wf-late"}});
+                renewed.expiresAt = farFuture;   // TTL extended
+                mstore.loadDocument("default:workflows", renewed);
+            }
+            return ttlExpiryDecision(rm, store, p.pm, mstore, cfgManager, coll, id,
+                                     nullptr, firstAttempt);
+        });
+
+    const uint64_t expired = mstore.expireDocuments();
+
+    check(injected, "the injected concurrent TTL extension actually ran");
+    check(expired == 1, "exactly ONE document expired - wf-early only");
+
+    // wf-early: genuinely expired, cascaded as normal. Proves the sweep worked.
+    check(!residentInMemory(mstore, "default:workflows", "wf-early"), "wf-early expired");
+    check(!store.get("executions", "ex-early").has_value(), "its child cascaded away");
+
+    // wf-late: the whole point. Not expired, and - the data-loss half - its
+    // child was not cascaded either.
+    check(residentInMemory(mstore, "default:workflows", "wf-late"),
+          "wf-late was NOT expired - its TTL was extended after phase 1 collected it");
+    check(store.get("workflows", "wf-late").has_value(), "and it is intact in LMDB");
+    check(store.get("executions", "ex-late").has_value(),
+          "and ITS CHILD was not cascaded - this is the data-loss half of the race");
+    check(store.relation_index_child_count("exec_wf", "wf-late") == 1,
+          "the reverse-index posting survived too");
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 0,
+          "and it was not recorded as relation-blocked - it simply is not expired");
+
+    p.pm.stop();
+    mstore.stop();
+}
+
+// =========================================================================
+// v2.11.0 close-out review, FINDING 2 — a relation-blocked document must not
+// consume the sweep's candidate budget.
+//
+// Phase 1 never erases a blocked document's expiration index entry (deliberately
+// - it has to stay armed so the block can lift), so with one shared budget those
+// documents refill the cap on every sweep. `collections_` iteration order is
+// stable, so every collection after them SILENTLY STOPS BEING SWEPT - and TTL'd
+// parents under `restrict` is exactly the combination this feature creates, so
+// that is the expected steady state, not an edge case.
+//
+// Budgets come from Config so this can be shown at 3 documents instead of 10001.
+void test_ttl_blocked_documents_do_not_starve_the_sweep() {
+    TmpEnv t("ttl-starve");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-starve-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore::Config cfg;
+    cfg.ttlMaxCandidatesPerSweep = 3;     // exactly filled by the blocked parents
+    cfg.ttlBlockedRetrySweeps = 100;      // far enough away not to interfere
+    MemoryStore mstore(cfg);
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    std::atomic<bool> healthy{true};
+    std::atomic<uint64_t> drift{0};
+    mstore.setDocumentStoreMirror(
+        [&store](std::string_view) -> smartbotic::db::storage::DocumentStore* { return &store; },
+        &healthy, &drift);
+
+    RelationInfo rel;
+    rel.name = "default:exec_wf";
+    rel.child = "default:executions";
+    rel.childField = "workflowId";
+    rel.parent = "default:workflows";
+    rel.onDelete = OnDelete::Restrict;
+    std::string mgrErr;
+    check(rm.createRelation(rel, mgrErr), "declared the restrict relation");
+    store.set_relations("executions",
+                        {RelationRef{"exec_wf", "workflowId", "workflows", false, true}});
+
+    // Three restrict-blocked TTL'd parents. Their expiry times are the earliest,
+    // so they are collected first and fill the budget on the first sweep.
+    for (int i = 0; i < 3; ++i) {
+        const std::string pid = "wf-" + std::to_string(i);
+        Document parent; parent.id = pid; parent.collection = "workflows";
+        parent.set_data({{"name", pid}});
+        parent.expiresAt = 1 + static_cast<uint64_t>(i);
+        store.put("workflows", pid, parent);
+        mstore.loadDocument("default:workflows", parent);
+
+        Document child; child.id = "ex-" + std::to_string(i);
+        child.collection = "executions";
+        child.set_data({{"workflowId", pid}});
+        store.put("executions", child.id, child);
+        mstore.loadDocument("default:executions", child);
+    }
+
+    // The victim: a TTL'd document with NO children, expiring after the three
+    // blocked parents. Deliberately in the SAME collection, because that makes
+    // the ordering deterministic - `collections_` is an unordered_map, but a
+    // collection's expirationIndex is a std::map walked in ascending expiry
+    // order, so wf-0/1/2 (1, 2, 3) are always collected before wf-victim (100).
+    // Starvation inside one collection is the same defect as starvation across
+    // collections, and this way the test cannot pass or fail on hash order.
+    Document victim;
+    victim.id = "wf-victim";
+    victim.collection = "workflows";
+    victim.set_data({{"name", "wf-victim"}});
+    victim.expiresAt = 100;
+    store.put("workflows", "wf-victim", victim);
+    mstore.loadDocument("default:workflows", victim);
+
+    installTtlHook(mstore, rm, store, p.pm, cfgManager);
+
+    // Sweep 1: the three blocked parents are fresh, so they legitimately take
+    // the whole budget and the victim is not reached at all.
+    const uint64_t first = mstore.expireDocuments();
+    check(first == 0, "sweep 1 expired nothing - the budget went to the blocked parents");
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 3, "all three parents were refused");
+    check(mstore.ttlBlockedDocumentCount() == 3, "and all three are tracked as stuck");
+    check(residentInMemory(mstore, "default:workflows", "wf-victim"),
+          "and the victim has not been reached yet");
+
+    // Sweep 2: the blocked parents are known and not due, so they must consume
+    // NOTHING and the victim finally gets in. With one shared budget it never
+    // would - the three blocked parents refill the cap on every sweep, forever.
+    uint64_t total = first;
+    total += mstore.expireDocuments();
+    check(total == 1,
+          "the childless document was expired despite three blocked parents filling "
+          "the budget - blocked documents no longer starve the sweep");
+    check(!residentInMemory(mstore, "default:workflows", "wf-victim"), "the victim is gone");
+    check(!store.get("workflows", "wf-victim").has_value(), "and its DELETE reached LMDB");
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 3,
+          "and the blocked parents were not re-consulted - still three refusal events, "
+          "not three more per sweep");
+
+    // Ten more sweeps: still nothing re-examined, so the steady-state cost of a
+    // persistent block is zero rather than one LMDB scan + one WARN per document
+    // per second.
+    for (int i = 0; i < 10; ++i) mstore.expireDocuments();
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 3,
+          "ten further sweeps consulted the relation zero times");
+    check(mstore.ttlBlockedDocumentCount() == 3, "and the three are still tracked as stuck");
+
+    p.pm.stop();
+    mstore.stop();
+}
+
 // =========================================================================
 // v2.11.0 close-out — BOOT-PASS DRIFT MUST NOT DISABLE DESTRUCTIVE POLICIES.
 //
@@ -2643,6 +2889,8 @@ int main() {
     test_ttl_no_action_expires_and_leaves_the_reference_dangling();
     test_ttl_with_no_relations_expires_exactly_as_before();
     test_ttl_cascade_wal_is_durable_before_the_lmdb_commit();
+    test_ttl_concurrent_ttl_extension_is_not_expired();
+    test_ttl_blocked_documents_do_not_starve_the_sweep();
     test_boot_pass_drift_does_not_disable_the_cascade();
     test_pending_remirror_list_is_released_after_the_pass();