Przeglądaj źródła

fix(relations): close-out - TTL/relations, drift baseline, DropProject, migration op

Five remaining items before merge, each with a test confirmed by reverting the
fix and re-running.

1. Two verified one-liners. The undo-guard ERROR log in patchDocument/setAdd/
   setRemove used '{{}}/{{}}', which fmt escapes, so the sole diagnostic for a
   state the code calls unreachable printed '{}/{}' and dropped the collection
   and id. And runPendingRemirror() now clear()+shrink_to_fit()s
   pendingRemirror: RecoveryOutcome lives for the process lifetime, the pass
   runs once, and nothing reads the list after it (~100B x millions under
   --recovery-mode=wal_only).

2. Boot-pass drift no longer disables destructive relation policies. Drift is
   never reset and the post-replay re-mirror pass bumps it for every row it
   cannot repair - on every boot, since the restart replays the same failing
   row - so ONE unrepairable row permanently refused every cascade/set_null
   delete and every CreateRelation, on a condition runPendingRemirror()
   deliberately treats as advisory, with an error message telling the operator
   to restart. MemoryStore::markMirrorDriftBaseline() is now called as
   initialize()'s last act and the cascade, CreateRelation and
   CreateIndex(unique) gates test drift accrued SINCE READY. Every read gate
   keeps using the raw counter - a stale row really does mean MemoryStore is
   ahead. Residual per-row risk documented on the accessor.

3. TTL expiry handles children exactly as a manual delete does. expireDocuments()
   erased the parent and mirrored a DELETE with no enforcement at all, so a
   restrict-protected parent silently orphaned its children. Now: restrict
   SKIPS the expiry (document outlives its TTL, WARN per skip plus
   Stats::ttlExpiryBlockedByRelation), cascade and set_null run through
   relation_cascade.cpp's existing WAL-before-LMDB machinery (array references
   pull-and-keep), no_action proceeds. expireDocuments() is now two-phase so the
   decision runs with NO MemoryStore lock held - applyCascadeToMemory takes
   collection locks and getOrCreateCollection takes globalMutex_ exclusively, so
   calling it under phase 1's shared globalMutex_ would self-deadlock.
   Replication/Subscribe events wired per mutation. Also: a stale expiration
   index entry left by an EVICTED cascaded parent would have re-run the whole
   cascade every sweep forever, and the candidate list is capped per sweep.

4. DropProject sweeps the relation declarations it invalidates and un-arms the
   child. _views/_policies/_collection_meta have the same gap; not a shared fix,
   recorded at the call site.

5. create_relation migration op, so consumers that declare schema in migration
   files can declare relations at all. Modelled on create_view; goes through
   RelationManager::createRelation so the same-project and on_delete validation
   apply, and an unrecognised on_delete fails the migration rather than silently
   arming restrict. runMigrations() re-arms afterwards, since migrations run
   after the boot arming pass.

test_relation_enforcement 252 -> 336, test_relation_manager 50 -> 70; index 43,
subdb_identity 264, dual_write_mirror 73, secondary_index_keys 41,
document_store and document_ttl green. E2E: relations (with a new DropProject
phase), relations_client_e2e 33, policy_enforcement all passed. VERSION
untouched. Report: closeout-report.md
fszontagh 1 miesiąc temu
rodzic
commit
0d3ea91132

+ 27 - 10
service/src/database_grpc_impl.cpp

@@ -3094,15 +3094,23 @@ grpc::Status DatabaseGrpcImpl::CreateRelation(
     // declared to prevent. Refused up front rather than half-built: the
     // declaration would look successful and `rows_indexed` would even report
     // a plausible number.
-    if (!service_.mirrorHealthy() || service_.mirrorDriftCount() != 0) {
+    //
+    // ⚠ v2.11.0 close-out — mirrorDriftSinceReady(), NOT mirrorDriftCount().
+    // The raw counter is never reset and the boot path's re-mirror pass bumps
+    // it for every row it cannot repair, on every boot, so one unrepairable
+    // row made CreateRelation permanently impossible - and the old text below
+    // told the operator that a restart would clear it, which is false in
+    // exactly that case. See MemoryStore::markMirrorDriftBaseline().
+    if (!service_.mirrorHealthy() || service_.mirrorDriftSinceReady() != 0) {
         response->set_success(false);
         response->set_error("cannot declare a relation while the LMDB mirror is unhealthy "
-                            "or has drifted: the bootstrap scan that indexes existing rows "
-                            "reads LMDB only, and MemoryStore may be ahead of it right now, "
-                            "so the reverse index would come out incomplete and `restrict` "
-                            "would permit deletes it should refuse. Check the mirror drift "
-                            "count in the service log; a restart clears drift once the "
-                            "underlying cause is fixed.");
+                            "or has drifted since startup: the bootstrap scan that indexes "
+                            "existing rows reads LMDB only, and MemoryStore may be ahead of "
+                            "it right now, so the reverse index would come out incomplete "
+                            "and `restrict` would permit deletes it should refuse. Check the "
+                            "ERROR lines naming the failed collection and id in the service "
+                            "log; drift accrued under live traffic latches for this process, "
+                            "so a restart clears it ONLY if the underlying cause is gone.");
         return grpc::Status::OK;
     }
 
@@ -3618,12 +3626,21 @@ grpc::Status DatabaseGrpcImpl::CreateIndex(
     // the process to MemoryStore for life over one legacy row - the v2.8.1
     // fault). Health-only here meant that after such a boot, a duplicate
     // hiding in a stale LMDB row would be accepted silently.
+    //
+    // ⚠ v2.11.0 close-out — and it is mirrorDriftSinceReady(), not the raw
+    // counter, for the same reason CreateRelation's sibling gate is: the raw
+    // counter is never reset and the boot path's re-mirror pass bumps it on
+    // every boot for every row it cannot repair, which made declaring a unique
+    // constraint permanently impossible after one such row. The residual risk
+    // this accepts (a duplicate hiding in a row the boot pass left stale) is
+    // the same one documented on MemoryStore::markMirrorDriftBaseline().
     if (request->unique() &&
-        (!service_.mirrorHealthy() || service_.mirrorDriftCount() != 0)) {
+        (!service_.mirrorHealthy() || service_.mirrorDriftSinceReady() != 0)) {
         response->set_success(false);
         response->set_error("cannot declare a unique constraint while the LMDB mirror is "
-                            "unhealthy or has drifted: the duplicate check only consults "
-                            "LMDB, and MemoryStore may be ahead of it right now");
+                            "unhealthy or has drifted since startup: the duplicate check "
+                            "only consults LMDB, and MemoryStore may be ahead of it right "
+                            "now");
         return grpc::Status::OK;
     }
 

+ 176 - 3
service/src/database_service.cpp

@@ -4,6 +4,8 @@
 #include "storage/document_store.hpp"
 #include "storage/document_store_lmdb.hpp"
 #include "storage/dual_write_mirror.hpp"
+#include "relations/relation_cascade.hpp"
+#include "relations/relation_enforcement.hpp"
 #include "auth/auth_interceptor.hpp"
 #include "auth/principal.hpp"
 #include "tls/cert_generator.hpp"
@@ -198,6 +200,18 @@ bool DatabaseService::initialize() {
         relation_manager_->loadFromStore();
         applyRelationDeclarations();
 
+        // v2.11.0 close-out — teach the TTL sweeper about relations. Installed
+        // HERE, after loadFromStore() and the arming pass, deliberately: the
+        // hook consults the declaration cache and the reverse index, and an
+        // expiry that fired before either was ready would have decided from an
+        // empty cache. Until this line the sweeper behaves exactly as it did
+        // pre-v2.11.0 (which is also what every unit fixture that never
+        // installs the hook keeps doing).
+        store_->setTtlExpiryRelationHook(
+            [this](const std::string& qualifiedCollection, const std::string& id) {
+                return ttlExpiryRelationDecision(qualifiedCollection, id);
+            });
+
         // v2.9.0 — re-apply persisted index declarations to each project's LMDB
         // store. This is load-bearing, not bookkeeping: the declaration is what
         // makes the write path maintain an index, and the planner consults an
@@ -273,6 +287,30 @@ bool DatabaseService::initialize() {
             }
         }
 
+        // v2.11.0 close-out — LAST ACT OF THE BOOT PATH, and the position is
+        // load-bearing. Everything above that can bump the drift counter (the
+        // post-replay re-mirror pass, backfillIntoDocStore(), a malformed
+        // legacy collection key) is now behind the baseline, so
+        // mirrorDriftSinceReady() reports only drift a LIVE write caused.
+        // Destructive relation policies gate on that, not on the raw counter:
+        // see MemoryStore::markMirrorDriftBaseline() for the trap this closes
+        // (one unrepairable row disabling every cascade/set_null delete and
+        // every CreateRelation for the process's life, with an error message
+        // telling the operator to restart, which re-incurs it). Reads keep
+        // gating on the RAW counter - a stale row genuinely means MemoryStore
+        // is ahead, and that fallback must stay.
+        store_->markMirrorDriftBaseline();
+        if (store_->mirrorDriftBaseline() > 0) {
+            spdlog::warn("mirror drift at READY: {} - accrued on the boot path (see the "
+                         "ERROR lines above for the exact rows). Reads for this process "
+                         "will be served from MemoryStore rather than LMDB. Destructive "
+                         "relation policies are NOT disabled by this, but the rows named "
+                         "above are still stale in LMDB until each is rewritten through "
+                         "an ordinary write; a restart re-attempts and, if the cause "
+                         "persists, re-incurs the same count.",
+                         store_->mirrorDriftBaseline());
+        }
+
         spdlog::info("Database service initialized successfully");
         return true;
 
@@ -345,7 +383,83 @@ bool DatabaseService::dropProject(const std::string& name, std::string& error) {
         error = "project registry not initialized";
         return false;
     }
-    return projects_->drop(name, error);
+
+    // v2.11.0 close-out — capture the relation declarations belonging to this
+    // project BEFORE the drop. Relations are validated same-project at
+    // creation (no LMDB transaction spans two envs), so `child` and `parent`
+    // are always in the same project as `name` itself; listRelations(name)
+    // therefore names every declaration the drop invalidates. Captured first
+    // because listRelations needs the project to still be nameable and, more
+    // importantly, because un-arming needs the store to still exist.
+    std::vector<RelationInfo> doomed;
+    if (relation_manager_) {
+        try {
+            doomed = relation_manager_->listRelations(name);
+        } catch (const std::exception& e) {
+            spdlog::warn("v2.11 relations: could not enumerate relations of project '{}' "
+                         "before dropping it ({}); its declarations may be left behind",
+                         name, e.what());
+        }
+    }
+
+    if (!projects_->drop(name, error)) return false;
+
+    // ---- The drop succeeded; the declarations now point at a namespace that
+    // no longer exists. Sweep them out of the global `_relations` collection
+    // and un-arm each affected child collection.
+    //
+    // Order (drop first, sweep second) is deliberate: the registry is the one
+    // that refuses "default" and rejects unknown names, so sweeping first would
+    // destroy declarations for a project whose drop then failed. The cost of
+    // this order is that the per-project store is usually already gone by the
+    // time we try to un-arm, which makes the un-arm a no-op - correct, since a
+    // store that no longer exists cannot be maintaining a reverse index. It
+    // still runs, because ProjectStoreRegistry may hold a live
+    // LmdbDocumentStore for a re-created project of the same name.
+    //
+    // ⚠ Advisory, never fatal: the project IS dropped at this point, and a
+    // leftover declaration is litter, not corruption (its child and parent
+    // collections are both gone, so nothing can be enforced against them). A
+    // failure here must not report the drop as failed.
+    for (const auto& r : doomed) {
+        try {
+            std::string err;
+            if (!relation_manager_->dropRelation(r.name, err)) {
+                spdlog::warn("v2.11 relations: could not drop declaration '{}' left by "
+                             "dropped project '{}': {}", r.name, name, err);
+                continue;
+            }
+            const auto rc = resolveCollection(r.child);
+            auto* lmdb = dynamic_cast<smartbotic::db::storage::LmdbDocumentStore*>(
+                docStore(rc.project));
+            if (lmdb != nullptr) {
+                // Empty list = "this collection has no relations", which is what
+                // set_relations replaces the previous list with. Every relation
+                // of a dropped project goes, so clearing the child wholesale is
+                // exactly right here (unlike DropRelation, which must re-arm the
+                // survivors).
+                lmdb->set_relations(rc.collection, {});
+            }
+            spdlog::info("v2.11 relations: dropped declaration '{}' with the project '{}'",
+                         r.name, name);
+        } catch (const std::exception& e) {
+            spdlog::warn("v2.11 relations: could not clean up declaration '{}' after "
+                         "dropping project '{}': {}", r.name, name, e.what());
+        }
+    }
+
+    // ⚠ KNOWN, NOT FIXED HERE (v2.11.0 close-out): _views, _policies and
+    // _collection_meta have the SAME gap - ProjectStoreRegistry::drop() removes
+    // the project's env and nothing else, so a dropped project leaves its view
+    // definitions, access policies and per-collection configs behind in those
+    // global system collections too. Not a shared fix: each manager keys and
+    // caches differently (ViewManager keys `<project>:<name>`, PolicyManager
+    // `<project>:<principal>`, CollectionConfigManager `<project>:<collection>`
+    // and canonicalises the cache key only), so each needs its own sweep, and
+    // PolicyManager's has a security dimension this one does not (a re-created
+    // project inheriting a stale policy). Scoped out deliberately; recorded so
+    // it is not rediscovered as new.
+    return true;
 }
 
 bool DatabaseService::runMigrations() {
@@ -358,8 +472,23 @@ bool DatabaseService::runMigrations() {
     migrationConfig.autoApply = config_.migrations.autoApply;
     migrationConfig.failOnError = config_.migrations.failOnError;
 
-    migrationRunner_ = std::make_unique<MigrationRunner>(*store_, *view_manager_, migrationConfig);
-    return migrationRunner_->runMigrations();
+    migrationRunner_ = std::make_unique<MigrationRunner>(
+        *store_, *view_manager_, *relation_manager_, migrationConfig);
+    const bool ok = migrationRunner_->runMigrations();
+
+    // v2.11.0 close-out — a `create_relation` migration op only PERSISTS the
+    // declaration. Arming the write path and building the reverse index over
+    // rows that already exist is what applyRelationDeclarations() does, and it
+    // has already run for this boot (initialize() calls it before migrations),
+    // so without this a migration-declared relation would maintain no reverse
+    // index and enforce nothing until the NEXT restart - the "declared but not
+    // enforcing" state this feature's self-heal exists to make unreachable.
+    // Idempotent and cheap: re-arming an already-armed child replaces an
+    // identical list, and the bootstrap loop skips any relation whose index
+    // sub-db already exists. Runs even when the migration run reported failure,
+    // because a partially-applied run may still have created a relation.
+    applyRelationDeclarations();
+    return ok;
 }
 
 void DatabaseService::applyIndexDeclarations() {
@@ -413,6 +542,50 @@ void DatabaseService::applyIndexDeclarations() {
     }
 }
 
+MemoryStore::TtlExpiryAction DatabaseService::ttlExpiryRelationDecision(
+    const std::string& qualifiedCollection, const std::string& id) {
+    using Action = MemoryStore::TtlExpiryAction;
+
+    // Thin binder over relations/relation_cascade.cpp's ttlExpiryDecision() -
+    // the policy itself lives there so it is unit-testable against the same
+    // fixtures the cascade tests use, without a DatabaseService. All this does
+    // is resolve the per-project LMDB store and supply the replication/events
+    // notifier.
+    if (!relation_manager_ || !config_manager_ || !store_ || !persistence_) {
+        return Action::Proceed;
+    }
+    if (qualifiedCollection.empty() || qualifiedCollection[0] == '_') return Action::Proceed;
+
+    try {
+        const auto rc = resolveCollection(qualifiedCollection);
+        auto* lmdb = dynamic_cast<smartbotic::db::storage::LmdbDocumentStore*>(
+            docStore(rc.project));
+        // No LMDB substrate means no reverse index to consult, so the
+        // pre-v2.11.0 behaviour is all that is available.
+        if (lmdb == nullptr) return Action::Proceed;
+
+        // Replication + Subscribe events, driven per mutation - the same wiring
+        // DatabaseGrpcImpl::performRelationCascade() uses. Without it a
+        // TTL-driven cascade would be invisible to followers and subscribers,
+        // and a follower that never saw the child deletions diverges
+        // permanently (the v2.3.1 class of bug).
+        auto notify = [this](const std::string& coll, const std::string& docId,
+                             const std::optional<Document>& doc, EventType eventType) {
+            notifyReplicationAndEvents(coll, docId, doc, eventType);
+        };
+
+        return smartbotic::database::ttlExpiryDecision(
+            *relation_manager_, *lmdb, *persistence_, *store_, *config_manager_,
+            qualifiedCollection, id, notify);
+    } catch (const std::exception& e) {
+        // Never expire a parent whose children could not be evaluated.
+        spdlog::error("TTL expiry of '{}/{}': could not resolve its storage ({}) - the "
+                      "document is left in place and will be retried",
+                      qualifiedCollection, id, e.what());
+        return Action::Skip;
+    }
+}
+
 void DatabaseService::applyRelationDeclarations() {
     // Group by (project, bare child collection) — set_relations() is a
     // per-collection call on that project's LmdbDocumentStore and replaces

+ 26 - 0
service/src/database_service.hpp

@@ -290,6 +290,15 @@ public:
     uint64_t mirrorDriftCount() const noexcept {
         return mirror_drift_count_.load(std::memory_order_relaxed);
     }
+    // v2.11.0 close-out — drift accrued AFTER the boot path finished. What
+    // destructive relation policies gate on; every READ gate keeps using
+    // mirrorDriftCount() above. See MemoryStore::markMirrorDriftBaseline().
+    uint64_t mirrorDriftSinceReady() const noexcept {
+        return store_ == nullptr ? 0 : store_->mirrorDriftSinceBaseline();
+    }
+    uint64_t mirrorDriftAtReady() const noexcept {
+        return store_ == nullptr ? 0 : store_->mirrorDriftBaseline();
+    }
 
     // v2.11.0 T12 review (C2) — replication queueing + Subscribe event
     // publication, extracted out of the MemoryStore persistCallback_
@@ -344,6 +353,23 @@ private:
     // persisted in _collection_meta. Must run at boot, after config load.
     void applyIndexDeclarations();
 
+    // v2.11.0 close-out — MemoryStore::TtlExpiryRelationHook. Decides what the
+    // TTL sweeper must do about an expiring parent's children, and PERFORMS the
+    // cascade itself when one is required, by calling the ordinary
+    // relations/relation_cascade.cpp machinery (WAL-before-LMDB) rather than a
+    // second, sweeper-local implementation - a hand-rolled cascade that wrote
+    // only LMDB would have its child deletions resurrected by the next boot's
+    // WAL replay, since MemoryStore is rebuilt from snapshot + WAL and NOT from
+    // LMDB.
+    //
+    // Runs on the cleanup thread with NO MemoryStore lock held (that is
+    // MemoryStore::expireDocuments()' phase-2 contract, and it is what makes
+    // calling back into MemoryStore here safe). Installed in setupComponents();
+    // must never throw (the sweeper catches, but a throw per document would
+    // stop expiry making progress).
+    MemoryStore::TtlExpiryAction ttlExpiryRelationDecision(
+        const std::string& qualifiedCollection, const std::string& id);
+
     // v2.11.0 T6a — arm each project's LmdbDocumentStore with the relation
     // declarations loaded by relation_manager_->loadFromStore(), grouped by
     // CHILD collection. LmdbDocumentStore has no knowledge of RelationManager

+ 142 - 43
service/src/memory_store.cpp

@@ -1424,7 +1424,7 @@ uint64_t MemoryStore::patchDocument(const std::string& collection, const std::st
         // behaviour and the lesser harm, and it is unreachable anyway: put()
         // cannot reject without a declaration.
         if (!wantUndo) {
-            spdlog::error("mirror rejected a write to '{{}}/{{}}' but no undo snapshot "
+            spdlog::error("mirror rejected a write to '{}/{}' but no undo snapshot "
                           "was taken (no unique field or validate_on_write relation "
                           "is declared for this collection) - the in-memory row is "
                           "left as written", collection, id);
@@ -1769,7 +1769,7 @@ bool MemoryStore::setAdd(const std::string& collection, const std::string& setId
         // behaviour and the lesser harm, and it is unreachable anyway: put()
         // cannot reject without a declaration.
         if (!wantUndo) {
-            spdlog::error("mirror rejected a write to '{{}}/{{}}' but no undo snapshot "
+            spdlog::error("mirror rejected a write to '{}/{}' but no undo snapshot "
                           "was taken (no unique field or validate_on_write relation "
                           "is declared for this collection) - the in-memory row is "
                           "left as written", collection, setId);
@@ -1849,7 +1849,7 @@ bool MemoryStore::setRemove(const std::string& collection, const std::string& se
         // behaviour and the lesser harm, and it is unreachable anyway: put()
         // cannot reject without a declaration.
         if (!wantUndo) {
-            spdlog::error("mirror rejected a write to '{{}}/{{}}' but no undo snapshot "
+            spdlog::error("mirror rejected a write to '{}/{}' but no undo snapshot "
                           "was taken (no unique field or validate_on_write relation "
                           "is declared for this collection) - the in-memory row is "
                           "left as written", collection, setId);
@@ -2205,60 +2205,159 @@ void MemoryStore::loadVector(const std::string& collection, const std::string& d
 
 // ===== TTL Management =====
 
+void MemoryStore::setTtlExpiryRelationHook(TtlExpiryRelationHook hook) {
+    ttlExpiryRelationHook_ = std::move(hook);
+}
+
 uint64_t MemoryStore::expireDocuments() {
     uint64_t expired = 0;
     auto now = currentTimeMs();
 
-    std::shared_lock<std::shared_mutex> globalLock(globalMutex_);
-
-    for (auto& [collName, coll] : collections_) {
-        std::unique_lock<std::shared_mutex> collLock(coll->mutex);
-
-        // Find expired documents
-        std::vector<std::string> toExpire;
-        auto it = coll->expirationIndex.begin();
-        while (it != coll->expirationIndex.end() && it->first <= now) {
-            for (const auto& id : it->second) {
-                toExpire.push_back(id);
+    // ---- Phase 1: collect candidates under the locks, mutate nothing. -----
+    // See the header comment on expireDocuments() for why the phases are split:
+    // the relation hook re-enters MemoryStore (a cascade takes each affected
+    // collection's lock, and getOrCreateCollection takes globalMutex_
+    // EXCLUSIVELY), so it can only be called with no lock held. Nothing is
+    // erased here, and - unlike the pre-v2.11.0 loop - the expiration index
+    // entry is NOT dropped here either: an expiry a `restrict` relation blocks
+    // has to stay armed so the next sweep retries it, and re-inserting an entry
+    // we had already erased would be a second chance to get it wrong.
+    struct Candidate {
+        std::string collection;
+        std::string id;
+        uint64_t expiresAt;
+    };
+    // ⚠ BOUNDED PER SWEEP. Collecting the whole expired set first is what makes
+    // the lock-free second phase possible, but an unbounded vector here is the
+    // same retention class the v2.11.0 re-mirror finding was about: a backlog of
+    // millions of expired documents (a long outage, or a TTL applied
+    // retroactively) would be ~80-100 bytes each, resident at once, on a
+    // background thread. Whatever this sweep does not reach stays armed in the
+    // expiration index and is picked up by the next one, so the only cost of the
+    // cap is latency on a backlog that is already late.
+    constexpr size_t kMaxCandidatesPerSweep = 10000;
+    std::vector<Candidate> candidates;
+    {
+        std::shared_lock<std::shared_mutex> globalLock(globalMutex_);
+        for (auto& [collName, coll] : collections_) {
+            if (candidates.size() >= kMaxCandidatesPerSweep) break;
+            std::shared_lock<std::shared_mutex> collLock(coll->mutex);
+            for (auto it = coll->expirationIndex.begin();
+                 it != coll->expirationIndex.end() && it->first <= now; ++it) {
+                for (const auto& id : it->second) {
+                    candidates.push_back(Candidate{collName, id, it->first});
+                }
+                if (candidates.size() >= kMaxCandidatesPerSweep) break;
             }
-            it = coll->expirationIndex.erase(it);
         }
+    }
+    if (candidates.empty()) return 0;
 
-        // Remove expired documents
-        for (const auto& id : toExpire) {
-            auto docIt = coll->documents.find(id);
-            if (docIt != coll->documents.end()) {
-                // Track memory before removal
-                uint64_t docSize = estimateDocumentSize(docIt->second);
-
-                coll->documents.erase(docIt);
-                expired++;
-
-                // v2.0 dual-write under lock — expire is semantically a delete
-                // for the substrate. Mirror it as DELETE so LMDB doesn't keep
-                // an orphan doc whose MemoryStore version has expired. Same
-                // for the vector sub-db; del_vector is a no-op if no vector.
-                mirrorWriteToDocStore(collName, id, std::nullopt, EventType::DELETE);
-                mirrorVectorToDocStore(collName, id, nullptr, EventType::DELETE);
+    // ---- Phase 2: one document at a time, no lock held on entry. ----------
+    for (const auto& cand : candidates) {
+        // The hook decides what this document's children require. Unset (no
+        // relations wired at all) is the pre-v2.11.0 path, unchanged; the hook
+        // itself early-returns Proceed when the declaration map holds nothing
+        // for this collection, so the common case costs one map lookup.
+        TtlExpiryAction action = TtlExpiryAction::Proceed;
+        if (ttlExpiryRelationHook_) {
+            try {
+                action = ttlExpiryRelationHook_(cand.collection, cand.id);
+            } catch (const std::exception& e) {
+                // The sweeper is a background thread: an escaping exception
+                // would kill it and stop ALL expiry for the process. Skip this
+                // document and retry it next sweep - never expire a parent
+                // whose children we failed to evaluate.
+                spdlog::error("TTL expiry: relation check for '{}/{}' threw ({}) - "
+                              "the document is left in place and will be retried",
+                              cand.collection, cand.id, e.what());
+                action = TtlExpiryAction::Skip;
+            }
+        }
 
-                // Emit callbacks (unlock first to avoid deadlock)
-                collLock.unlock();
+        if (action == TtlExpiryAction::Skip) {
+            // Nothing to undo: phase 1 erased nothing, so the expiration index
+            // entry is still armed and the next sweep sees this candidate again.
+            std::lock_guard<std::mutex> statsLock(statsMutex_);
+            stats_.ttlExpiryBlockedByRelation++;
+            continue;
+        }
 
-                {
-                    std::lock_guard<std::mutex> statsLock(statsMutex_);
-                    stats_.totalDocuments--;
-                    stats_.expiredCount++;
+        if (action == TtlExpiryAction::Handled) {
+            // The hook ran the whole cascade: WAL, LMDB commit, MemoryStore
+            // apply and per-mutation replication/events. The document is gone
+            // from both substrates and its expiration index entry went with it
+            // (applyCascadeToMemory -> unloadDocument). Touching it here would
+            // double-log the delete.
+            ++expired;
+            {
+                std::lock_guard<std::mutex> statsLock(statsMutex_);
+                stats_.expiredCount++;
+            }
+            // Defensive: applyCascadeToMemory -> unloadDocument drops the
+            // expiration index entry along with the document, but only if the
+            // document was actually resident. A parent that had been evicted
+            // leaves the entry behind, and an entry with no document would be
+            // re-collected on EVERY sweep and re-run the whole cascade each
+            // time - WAL writes and all - forever.
+            {
+                std::shared_lock<std::shared_mutex> globalLock(globalMutex_);
+                auto collIt = collections_.find(cand.collection);
+                if (collIt != collections_.end()) {
+                    auto* coll = collIt->second.get();
+                    std::unique_lock<std::shared_mutex> collLock(coll->mutex);
+                    if (coll->documents.find(cand.id) == coll->documents.end()) {
+                        removeFromExpirationIndex(*coll, cand.id, cand.expiresAt);
+                    }
                 }
+            }
+            continue;
+        }
 
-                // Update memory tracking atomically
-                estimatedMemoryBytes_.fetch_sub(docSize, std::memory_order_relaxed);
+        // ---- Proceed: the pre-v2.11.0 expiry, per document. ---------------
+        uint64_t docSize = 0;
+        {
+            std::shared_lock<std::shared_mutex> globalLock(globalMutex_);
+            auto collIt = collections_.find(cand.collection);
+            if (collIt == collections_.end()) continue;   // collection dropped
+            auto* coll = collIt->second.get();
+
+            std::unique_lock<std::shared_mutex> collLock(coll->mutex);
+            auto docIt = coll->documents.find(cand.id);
+            if (docIt == coll->documents.end()) {
+                // Already gone (deleted, or evicted). Drop the stale index
+                // entry so it is not re-examined every sweep forever.
+                removeFromExpirationIndex(*coll, cand.id, cand.expiresAt);
+                continue;
+            }
+            // Re-check the expiry under the lock: between phase 1 and here an
+            // update may have extended or cleared the TTL, and expiring on a
+            // stale reading would delete a document that is no longer expired.
+            if (docIt->second.expiresAt == 0 || docIt->second.expiresAt > now) {
+                continue;
+            }
+            docSize = estimateDocumentSize(docIt->second);
+            removeFromExpirationIndex(*coll, cand.id, docIt->second.expiresAt);
+            coll->documents.erase(docIt);
 
-                emitPersist(collName, id, std::nullopt, EventType::EXPIRE);
-                emitEvent(EventType::EXPIRE, collName, id);
+            // v2.0 dual-write under lock — expire is semantically a delete
+            // for the substrate. Mirror it as DELETE so LMDB doesn't keep
+            // an orphan doc whose MemoryStore version has expired. Same
+            // for the vector sub-db; del_vector is a no-op if no vector.
+            mirrorWriteToDocStore(cand.collection, cand.id, std::nullopt, EventType::DELETE);
+            mirrorVectorToDocStore(cand.collection, cand.id, nullptr, EventType::DELETE);
+        }
 
-                collLock.lock();
-            }
+        ++expired;
+        {
+            std::lock_guard<std::mutex> statsLock(statsMutex_);
+            stats_.totalDocuments--;
+            stats_.expiredCount++;
         }
+        estimatedMemoryBytes_.fetch_sub(docSize, std::memory_order_relaxed);
+
+        emitPersist(cand.collection, cand.id, std::nullopt, EventType::EXPIRE);
+        emitEvent(EventType::EXPIRE, cand.collection, cand.id);
     }
 
     return expired;

+ 131 - 1
service/src/memory_store.hpp

@@ -216,6 +216,59 @@ public:
             : mirrorDriftCount_->load(std::memory_order_relaxed);
     }
 
+    /**
+     * v2.11.0 close-out — DRIFT ACCRUED AFTER THE BOOT PATH FINISHED, which is
+     * what a destructive relation policy must gate on. Read the whole of this
+     * before using mirrorDriftCount() for a new gate.
+     *
+     * mirrorDriftCount_ is never reset, and the post-replay re-mirror pass
+     * (remirrorDocuments()) deliberately bumps it - and deliberately does NOT
+     * flip health - for every row it could not write. That is correct for
+     * READS: a stale LMDB row means MemoryStore is ahead, and a nonzero drift
+     * count is exactly what routes reads to MemoryStore.
+     *
+     * It was NOT correct as the gate for cascade/set_null deletes and
+     * CreateRelation. Those refuse while drift is nonzero, drift is never
+     * reset, and a row the boot pass cannot repair is re-attempted and
+     * re-failed on EVERY boot - so one unrepairable row permanently disabled
+     * every destructive relation policy service-wide, on a condition
+     * runPendingRemirror() explicitly treats as advisory ("recovery is NOT
+     * failed by this"), and the error text told the operator to restart, which
+     * cannot help.
+     *
+     * DatabaseService::initialize() calls markMirrorDriftBaseline() as its last
+     * act, so everything the boot path accrued - the re-mirror pass, the
+     * backfill, a malformed legacy collection key - lands in the baseline, and
+     * this function reports only drift a LIVE write caused. Live drift still
+     * latches for the process lifetime and still refuses, which is the part
+     * that was right: a mirror that broke under traffic is exactly when a
+     * cascade must not push a stale LMDB body back into MemoryStore.
+     *
+     * ⚠ HONEST LIMIT, accepted deliberately (operator's ruling): the specific
+     * rows the boot pass left stale are still stale, and a cascade that touches
+     * one of them can still overwrite a fresher MemoryStore body with the stale
+     * LMDB copy. Excluding them from the gate trades that narrow, per-row risk
+     * for not disabling every destructive policy on every collection forever.
+     * A per-row remedy would need the failed id set retained, which is the
+     * unbounded-retention problem finding 3 removed. Repair the row (rewrite
+     * it through the ordinary write path) to clear the staleness itself.
+     */
+    void markMirrorDriftBaseline() noexcept {
+        mirrorDriftBaseline_.store(mirrorDriftCount(), std::memory_order_relaxed);
+    }
+    [[nodiscard]] uint64_t mirrorDriftBaseline() const noexcept {
+        return mirrorDriftBaseline_.load(std::memory_order_relaxed);
+    }
+    [[nodiscard]] uint64_t mirrorDriftSinceBaseline() const noexcept {
+        const uint64_t total = mirrorDriftCount();
+        const uint64_t base = mirrorDriftBaseline_.load(std::memory_order_relaxed);
+        // Saturating: the baseline can only ever be <= total (it is a snapshot
+        // of the same monotonic counter), but subtracting unsigned without the
+        // guard would wrap into "billions of drift" if that ever stopped being
+        // true, which fails in the loudest possible wrong direction.
+        return total > base ? total - base : 0;
+    }
+
     /**
      * v2.11.0 final review (finding 5) — does a write to `collection` need
      * T11's pre-write undo snapshot?
@@ -628,10 +681,71 @@ public:
 
     // ===== TTL Management =====
 
+    /**
+     * v2.11.0 close-out — WHAT A TTL EXPIRY MUST DO ABOUT THE EXPIRING
+     * DOCUMENT'S CHILDREN. Returned by the relation hook below, one call per
+     * candidate document, with NO MemoryStore lock held.
+     *
+     * Operator's ruling: a TTL-deleted parent handles its children exactly the
+     * way a manually deleted parent does - like MySQL. Before this, the sweeper
+     * erased the document and mirrored a DELETE with no relation enforcement at
+     * all, so a `restrict`-protected parent silently vanished and orphaned
+     * every child.
+     */
+    enum class TtlExpiryAction {
+        // No relation has anything to say about this document (the common case,
+        // and what a build with no relations declared always returns): expire
+        // it exactly as pre-v2.11.0 did - erase, mirror a DELETE, emit EXPIRE.
+        Proceed,
+        // The expiry must NOT happen: a `restrict` relation blocks it (a manual
+        // delete would have failed with FAILED_PRECONDITION), or the cascade
+        // machinery refused because the LMDB mirror is unhealthy/drifted.
+        // ⚠ The document is left in place WITH ITS EXPIRY RE-ARMED, so the next
+        // sweep retries - which means it OUTLIVES ITS TTL for as long as the
+        // block lasts. That is a retention-policy surprise, so each skip logs a
+        // WARN naming the relation and the blocking child count, and
+        // Stats::ttlExpiryBlockedByRelation counts them.
+        Skip,
+        // The hook already performed the whole delete through the ordinary
+        // cascade machinery (relations/relation_cascade.cpp: WAL for the parent
+        // and every child mutation, fsync, one atomic LMDB commit, then the
+        // MemoryStore apply, with replication + Subscribe events driven per
+        // mutation). The sweeper must NOT touch the document again - doing so
+        // would double-log the delete to the WAL and re-mirror it.
+        Handled
+    };
+    using TtlExpiryRelationHook =
+        std::function<TtlExpiryAction(const std::string& qualifiedCollection,
+                                      const std::string& id)>;
+
+    /**
+     * Install the hook above. DatabaseService does this once, at boot, after
+     * the RelationManager is loaded - MemoryStore has no access to
+     * RelationManager, LmdbDocumentStore, PersistenceManager or
+     * CollectionConfigManager, all four of which a cascade needs.
+     *
+     * Unset (the default, and every unit fixture that does not need it) is
+     * exactly the pre-v2.11.0 behaviour.
+     */
+    void setTtlExpiryRelationHook(TtlExpiryRelationHook hook);
+
     /**
      * Expire documents that have passed their TTL.
      * Called automatically by background thread.
-     * @return Number of documents expired
+     *
+     * ⚠ LOCKING (v2.11.0 close-out): this runs in TWO PHASES and the split is
+     * load-bearing. Phase 1 collects candidates under the global read lock plus
+     * each collection's write lock, exactly as before. Phase 2 releases both and
+     * then processes one document at a time, taking that collection's lock
+     * afresh per document. The relation hook is only ever called in phase 2,
+     * with NO MemoryStore lock held, because a cascade re-enters MemoryStore
+     * (applyCascadeToMemory takes each affected collection's lock in turn, and
+     * getOrCreateCollection takes globalMutex_ EXCLUSIVELY): calling it from
+     * inside phase 1 would self-deadlock on globalMutex_ - a std::shared_mutex
+     * is not upgradable - and would deadlock outright against a request thread
+     * whenever the child collection and the parent collection differed.
+     *
+     * @return Number of documents actually expired (Skip does not count)
      */
     uint64_t expireDocuments();
 
@@ -642,6 +756,13 @@ public:
         uint64_t totalCollections = 0;
         uint64_t estimatedMemoryBytes = 0;
         uint64_t expiredCount = 0;
+        // v2.11.0 close-out — TTL expiries the sweeper REFUSED because a
+        // relation blocked them (a `restrict` relation with live children, or
+        // the cascade path refusing while the LMDB mirror is unhealthy/drifted).
+        // Each one means a document is still present PAST its TTL and will be
+        // retried on the next sweep. Nonzero is an operator signal, not an
+        // error: the alternative was silently orphaning the children.
+        uint64_t ttlExpiryBlockedByRelation = 0;
         uint64_t insertCount = 0;
         uint64_t updateCount = 0;
         uint64_t deleteCount = 0;
@@ -1141,8 +1262,17 @@ private:
     // unset, mirrorWriteToDocStore / mirrorVectorToDocStore are no-ops
     // (preserves test/bootstrap paths that don't wire LMDB).
     DocumentStoreResolver docStoreResolver_;
+    // v2.11.0 close-out — see setTtlExpiryRelationHook(). Written once at boot
+    // (before the expiration thread does anything that consults it) and read
+    // from the cleanup thread only.
+    TtlExpiryRelationHook ttlExpiryRelationHook_;
+
     std::atomic<bool>* mirrorHealthy_ = nullptr;
     std::atomic<uint64_t>* mirrorDriftCount_ = nullptr;
+    // v2.11.0 close-out — drift already accrued when the boot path finished.
+    // See markMirrorDriftBaseline(). Owned here rather than in DatabaseService
+    // because relation_cascade.cpp only ever gets a MemoryStore&.
+    std::atomic<uint64_t> mirrorDriftBaseline_{0};
 
     // Per-document last-write WAL sequence map. Populated by the persist
     // callback after each WAL append and by recovery's WAL replay. Used

+ 57 - 1
service/src/migrations/migration_runner.cpp

@@ -11,9 +11,11 @@
 
 namespace smartbotic::database {
 
-MigrationRunner::MigrationRunner(MemoryStore& store, ViewManager& view_manager, Config config)
+MigrationRunner::MigrationRunner(MemoryStore& store, ViewManager& view_manager,
+                                 RelationManager& relation_manager, Config config)
     : store_(store)
     , view_manager_(view_manager)
+    , relation_manager_(relation_manager)
     , config_(std::move(config))
 {
     // Ensure migrations collection exists
@@ -341,6 +343,60 @@ bool MigrationRunner::applyOperation(const nlohmann::json& operation) {
         return true;
     }
 
+    // v2.11.0 close-out — create_relation. Modelled on create_view above: same
+    // file shape, same idempotency contract ("already exists" is success on
+    // replay), and the same decision to validate through the manager rather
+    // than writing the declaration record directly.
+    if (type == "create_relation") {
+        RelationInfo r;
+        r.name = operation.value("name", "");
+        r.child = operation.value("child", "");
+        r.childField = operation.value("child_field", "");
+        r.parent = operation.value("parent", "");
+        if (r.name.empty() || r.child.empty() || r.childField.empty() || r.parent.empty()) {
+            spdlog::error("create_relation: name, child, child_field and parent are all "
+                          "required (got name='{}' child='{}' child_field='{}' parent='{}')",
+                          r.name, r.child, r.childField, r.parent);
+            return false;
+        }
+
+        // VALIDATE, do not coerce — the same call the CreateRelation RPC makes,
+        // for the reason recorded in the header: an unrecognised value used to
+        // become Restrict silently, which stopped being a safe default the
+        // moment cascade/set_null became genuinely destructive. Absent means
+        // the documented default.
+        const std::string onDelete = operation.value("on_delete", "");
+        if (onDelete.empty()) {
+            r.onDelete = OnDelete::Restrict;
+        } else if (auto parsed = parseOnDelete(onDelete); parsed.has_value()) {
+            r.onDelete = *parsed;
+        } else {
+            spdlog::error("create_relation '{}': on_delete must be one of {} (got '{}')",
+                          r.name, kOnDeleteValues, onDelete);
+            return false;
+        }
+        r.validateOnWrite = operation.value("validate_on_write", false);
+
+        std::string err;
+        if (!relation_manager_.createRelation(r, err)) {
+            // Idempotent on replay, exactly as create_view is. Everything else
+            // (cross-project, malformed name, unknown collection) is a real
+            // failure and must fail the migration - a relation the author
+            // believes is armed and is not is the whole class of bug this
+            // feature exists to prevent.
+            if (err.find("already exists") != std::string::npos) {
+                spdlog::debug("create_relation '{}': already exists (idempotent)", r.name);
+                return true;
+            }
+            spdlog::error("create_relation '{}': {}", r.name, err);
+            return false;
+        }
+        spdlog::info("create_relation '{}': {}.{} -> {} (on_delete={})",
+                     r.name, r.child, r.childField, r.parent,
+                     onDelete.empty() ? "restrict" : onDelete);
+        return true;
+    }
+
     if (type == "insert") {
         std::string collection = operation.value("collection", "");
         std::string id = operation.value("id", "");

+ 34 - 1
service/src/migrations/migration_runner.hpp

@@ -1,6 +1,7 @@
 #pragma once
 
 #include "../memory_store.hpp"
+#include "../relations/relation_manager.hpp"
 #include "../views/view_manager.hpp"
 
 #include <filesystem>
@@ -50,6 +51,32 @@ namespace smartbotic::database {
  *   GTE/LT/LTE/IN/CONTAINS/EXISTS/REGEX/SEARCH — via string or integer
  *   op field), and `default_sort` (field + descending). Idempotent:
  *   "already exists" is treated as success on migration replay.
+ * - create_relation: Declare a referential-integrity relation (v2.11.0).
+ *   Required: `name`, `child`, `child_field`, `parent`. Optional:
+ *   `on_delete` (restrict|cascade|set_null|no_action, default restrict) and
+ *   `validate_on_write` (bool, default false). Idempotent: "already
+ *   exists" is treated as success on migration replay.
+ *
+ *   ⚠ Goes through RelationManager::createRelation, deliberately and not as
+ *   a shortcut, so the same-project rule and the on_delete validation apply
+ *   exactly as they do to the CreateRelation RPC. An unrecognised on_delete
+ *   FAILS the migration rather than being coerced to restrict - the same
+ *   decision the RPC made in the v2.11.0 final review, and it matters more
+ *   here: a typo in a file that ships in a deb would otherwise arm restrict
+ *   on every install while the author believed cascade was armed.
+ *
+ *   ⚠ Names are project-qualified the same way every other collection name
+ *   in a migration file is: a bare `users` resolves to `default:users`. A
+ *   consumer declaring relations in a non-default project must write the
+ *   qualified `<project>:<name>` form, and all four of name/child/parent
+ *   must name the SAME project (no LMDB transaction spans two project envs).
+ *
+ *   Arming and the bootstrap scan over existing rows are NOT done here:
+ *   DatabaseService::runMigrations() calls applyRelationDeclarations() after
+ *   the run, which arms every declaration and builds any reverse index whose
+ *   sub-db is absent. Doing it per-op would duplicate that logic; doing it
+ *   nowhere would leave a migration-declared relation unenforced until the
+ *   next restart.
  */
 class MigrationRunner {
 public:
@@ -59,7 +86,12 @@ public:
         bool failOnError = true;
     };
 
-    MigrationRunner(MemoryStore& store, ViewManager& view_manager, Config config);
+    // v2.11.0 close-out — `relation_manager` is what the create_relation op
+    // declares through. A reference, not a pointer: every construction site
+    // has one, and an optional RelationManager would make "the op silently did
+    // nothing" a reachable state.
+    MigrationRunner(MemoryStore& store, ViewManager& view_manager,
+                    RelationManager& relation_manager, Config config);
 
     /**
      * Run all pending migrations.
@@ -110,6 +142,7 @@ private:
 
     MemoryStore& store_;
     ViewManager& view_manager_;
+    RelationManager& relation_manager_;
     Config config_;
 
     /**

+ 9 - 0
service/src/persistence/persistence_manager.cpp

@@ -467,6 +467,15 @@ void PersistenceManager::runPendingRemirror(MemoryStore& store, RecoveryOutcome&
                       "this - a repair pass that cannot repair one row must not "
                       "stop the service from starting.", remirror.failed);
     }
+
+    // v2.11.0 close-out — release the id list. RecoveryOutcome is held by
+    // DatabaseService for the process lifetime; the pass runs exactly once and
+    // nothing reads the list after it, so keeping it resident pinned ~100
+    // bytes per replayed row for nothing (millions of rows under
+    // --recovery-mode=wal_only). shrink_to_fit() as well as clear(), because
+    // clear() alone keeps the capacity - which is the entire allocation.
+    outcome.pendingRemirror.clear();
+    outcome.pendingRemirror.shrink_to_fit();
 }
 
 uint64_t PersistenceManager::logInsert(const std::string& collection, const Document& doc,

+ 10 - 0
service/src/persistence/persistence_manager.hpp

@@ -88,6 +88,16 @@ struct RecoveryOutcome {
     // declared. See MemoryStore::remirrorDocuments()'s doc comment.
     //
     // Empty after a recovery that replayed nothing, and after ForceEmpty.
+    //
+    // ⚠ RELEASED BY runPendingRemirror() ONCE THE PASS HAS RUN (v2.11.0
+    // close-out). RecoveryOutcome lives as DatabaseService::recovery_outcome_
+    // for the whole process lifetime, so leaving this populated pinned the
+    // entire replayed id set - ~100 bytes per entry, and the entry count is
+    // an install's whole history under --recovery-mode=wal_only - resident
+    // for nothing, since the pass runs exactly once and nothing reads the
+    // list afterwards. The COUNTS (updatesRemirroredAfterReplay /
+    // updatesRemirrorFailed) are what operator surfaces report and they are
+    // deliberately kept.
     std::vector<std::pair<std::string, std::string>> pendingRemirror;
 
     bool isNonTrivial() const {

+ 106 - 7
service/src/relations/relation_cascade.cpp

@@ -352,19 +352,31 @@ bool executeCascade(RelationManager& relations,
     // exactly the contract CascadeBlocked already carries. Same refusal
     // shape, and the same reasoning, as the CreateIndex(unique) and
     // CreateRelation gates.
-    if (!memStore.mirrorHealthy() || memStore.mirrorDriftCount() != 0) {
+    // ⚠ v2.11.0 close-out — the drift half of this gate is
+    // mirrorDriftSinceBaseline(), NOT mirrorDriftCount(). Drift is never
+    // reset and the boot path's post-replay re-mirror pass bumps it for every
+    // row it cannot repair - a condition runPendingRemirror() deliberately
+    // treats as advisory and which recurs on every boot, so the raw counter
+    // made ONE unrepairable row refuse every cascade/set_null delete
+    // service-wide, permanently, while the message above told the operator to
+    // restart. See MemoryStore::markMirrorDriftBaseline() for the full
+    // reasoning and for the residual per-row risk this accepts.
+    if (!memStore.mirrorHealthy() || memStore.mirrorDriftSinceBaseline() != 0) {
         throw CascadeBlocked(
             "refusing the cascade/set_null delete of '" + qualifiedParentCollection +
             "/" + parentId + "': the LMDB mirror is " +
-            (memStore.mirrorHealthy() ? "drifted (" + std::to_string(memStore.mirrorDriftCount()) +
-                                         " row(s) known stale)"
-                                      : "unhealthy") +
+            (memStore.mirrorHealthy()
+                 ? "drifted since startup (" +
+                       std::to_string(memStore.mirrorDriftSinceBaseline()) +
+                       " row(s) a live write left stale)"
+                 : "unhealthy") +
             ", so the child documents this cascade would rewrite are read from a substrate "
             "that may be behind MemoryStore - applying them would overwrite fresher data "
             "with older data. Restrict/no_action deletes are unaffected. Fix the mirror "
-            "(see the ERROR lines naming the failed collection and id) and restart to "
-            "clear drift, or use on_delete=no_action if the dangling reference is "
-            "acceptable.");
+            "(see the ERROR lines naming the failed collection and id): drift accrued "
+            "under live traffic latches for this process, so a restart clears it ONLY if "
+            "the underlying cause is gone. Or use on_delete=no_action if the dangling "
+            "reference is acceptable.");
     }
 
     const CascadePlan plan = planCascade(relations, store, configManager,
@@ -375,4 +387,91 @@ bool executeCascade(RelationManager& relations,
     return parentExisted;
 }
 
+MemoryStore::TtlExpiryAction ttlExpiryDecision(
+    RelationManager& relations,
+    smartbotic::db::storage::LmdbDocumentStore& store,
+    PersistenceManager& persistence,
+    MemoryStore& memStore,
+    CollectionConfigManager& configManager,
+    const std::string& qualifiedParentCollection,
+    const std::string& parentId,
+    const CascadeNotifyFn& notify) {
+    using Action = MemoryStore::TtlExpiryAction;
+
+    // ---- Scope: the no-relations case must cost exactly one map lookup. ----
+    // Same early-out shape the write path uses. Nothing below this line runs
+    // for a collection that is nobody's parent, which is every collection on
+    // every install that declares no relations.
+    if (qualifiedParentCollection.empty() || qualifiedParentCollection[0] == '_') {
+        return Action::Proceed;
+    }
+    const auto parentRelations = relations.relationsWithParent(qualifiedParentCollection);
+    if (parentRelations.empty()) return Action::Proceed;
+
+    try {
+        // relationsEnforced gates the sweeper exactly as it gates the Delete
+        // handler: off means every on_delete policy is skipped, so a TTL expiry
+        // behaves as it did pre-v2.11.0 rather than half-enforcing. Named local
+        // - RelationEnforcer binds the flag by reference.
+        const CollectionCfg cfg = configManager.configFor(qualifiedParentCollection);
+        if (!cfg.relationsEnforced) return Action::Proceed;
+
+        // ---- restrict: a manual delete FAILS, so the expiry must not happen.
+        RelationEnforcer enforcer(relations, store, cfg.relationsEnforced);
+        std::string err;
+        if (!enforcer.canDelete(qualifiedParentCollection, parentId, err)) {
+            // ⚠ Stated where an operator will see it: this document now
+            // OUTLIVES ITS TTL, indefinitely, for as long as a child keeps
+            // referencing it. That is the accepted price of not orphaning the
+            // children, and the alternative is what this change removes.
+            spdlog::warn("TTL expiry of '{}/{}' is BLOCKED by a restrict relation, so the "
+                         "document remains past its TTL and will be retried on the next "
+                         "sweep (MemoryStore stat ttlExpiryBlockedByRelation counts these): "
+                         "{}", qualifiedParentCollection, parentId, err);
+            return Action::Skip;
+        }
+
+        // ---- cascade / set_null: only reach for the cascade machinery when a
+        // destructive policy actually has work to do. A restrict-only or
+        // no_action-only parent falls through to the ordinary expiry, which
+        // keeps the EXPIRE event and the expiredCount stat exactly as they were.
+        bool hasDestructive = false;
+        for (const auto& r : parentRelations) {
+            if (r.onDelete == OnDelete::Cascade || r.onDelete == OnDelete::SetNull) {
+                hasDestructive = true;
+                break;
+            }
+        }
+        if (!hasDestructive) return Action::Proceed;
+
+        try {
+            (void)executeCascade(relations, store, persistence, memStore, configManager,
+                                 qualifiedParentCollection, parentId, notify);
+            return Action::Handled;
+        } catch (const CascadeBlocked& e) {
+            // Either a restrict-protected grandchild, or the mirror
+            // unhealthy/drifted refusal. Nothing was written in either case
+            // (both are thrown before writeCascadeWal()), so skipping is a
+            // clean no-op and the next sweep retries. Expiring the parent
+            // anyway is exactly the orphaning this function exists to remove.
+            spdlog::warn("TTL expiry of '{}/{}' skipped: {} - the document remains past "
+                         "its TTL and will be retried on the next sweep",
+                         qualifiedParentCollection, parentId, e.what());
+            return Action::Skip;
+        }
+    } catch (const std::exception& e) {
+        // A malformed name, MDB_READERS_FULL, a mid-cascade fault. Never expire
+        // a parent whose children could not be handled.
+        // ⚠ A cascade that threw AFTER writeCascadeWal() fsynced is already
+        // durable and will apply on the next restart - the same caveat the
+        // Delete handler documents. Skipping is still right: the sweeper must
+        // not also delete the parent behind that cascade's back.
+        spdlog::error("TTL expiry of '{}/{}' failed while handling its children ({}) - the "
+                      "document is left in place and will be retried; if the failure came "
+                      "after the cascade's WAL fsync, that cascade will still apply on the "
+                      "next restart", qualifiedParentCollection, parentId, e.what());
+        return Action::Skip;
+    }
+}
+
 } // namespace smartbotic::database

+ 60 - 0
service/src/relations/relation_cascade.hpp

@@ -218,6 +218,11 @@
 #include <vector>
 
 #include "document.hpp"
+// v2.11.0 close-out — for MemoryStore::TtlExpiryAction, the return type of
+// ttlExpiryDecision() below. MemoryStore was previously only forward-declared
+// here; the enum has to be a complete type at this point, and it belongs on
+// MemoryStore because MemoryStore's sweeper is what acts on it.
+#include "memory_store.hpp"
 #include "relation_manager.hpp"
 
 namespace smartbotic::db::storage {
@@ -353,6 +358,61 @@ void applyCascadeToMemory(MemoryStore& store,
 // written anywhere. Returns commitCascadeLmdb()'s result — whether the
 // parent document itself was present and removed — for the handler's
 // DeleteResponse.deleted field.
+// v2.11.0 close-out — WHAT A TTL EXPIRY MUST DO ABOUT THE EXPIRING PARENT'S
+// CHILDREN, and the cascade itself when one is required.
+//
+// Operator's ruling: a TTL-deleted parent handles its children exactly the way
+// a manually deleted parent does, like MySQL. Before this,
+// MemoryStore::expireDocuments() erased the document and mirrored a DELETE with
+// no relation enforcement at all, so a `restrict`-protected parent silently
+// vanished and orphaned every child - strictly the worst of the available
+// behaviours.
+//
+// This function is the sweeper's whole decision, and it lives HERE rather than
+// in DatabaseService so it is directly unit-testable against the same fixtures
+// the cascade tests use (DatabaseService::ttlExpiryRelationDecision() is a thin
+// binder over it, and MemoryStore calls it through
+// MemoryStore::TtlExpiryRelationHook). By policy:
+//   restrict  -> Skip. A manual delete FAILS, so the expiry must not happen.
+//                The document stays, its expiry stays armed, the next sweep
+//                retries - which means it OUTLIVES ITS TTL for as long as a
+//                child references it. Deliberate, and signalled: a WARN per
+//                skip plus MemoryStore::Stats::ttlExpiryBlockedByRelation.
+//   cascade   -> Handled, via executeCascade() below. NOT a second cascade
+//                implementation: WAL-before-LMDB is the whole reason this path
+//                exists (MemoryStore is rebuilt from snapshot + WAL, never from
+//                LMDB, so a sweeper that wrote only LMDB would have its child
+//                deletions resurrected by the next boot's replay).
+//   set_null  -> Handled, same call. Scalar reference nulled; an array-valued
+//                reference has the id pulled and the document kept, exactly as
+//                the request path does.
+//   no_action -> Proceed. Reference left dangling, as today.
+// and Proceed also when no relation names this collection as a parent (the
+// common case, early-returned before any LMDB or config work), or when the
+// collection's relationsEnforced is off (the documented escape hatch, gating
+// every on_delete policy and not just restrict, exactly as the Delete handler
+// does).
+//
+// Skips - never orphans - if executeCascade() refuses because the LMDB mirror is
+// unhealthy or has drifted since READY, and if anything throws. `notify` is
+// forwarded to executeCascade() so a TTL-driven cascade drives replication and
+// Subscribe events per mutation like any other; without it a follower would
+// never see the child deletions and would diverge permanently.
+//
+// ⚠ MUST be called with NO MemoryStore lock held - it re-enters MemoryStore
+// (applyCascadeToMemory takes each affected collection's lock, and
+// getOrCreateCollection takes globalMutex_ exclusively). That is exactly what
+// expireDocuments()' two-phase structure guarantees.
+MemoryStore::TtlExpiryAction ttlExpiryDecision(
+    RelationManager& relations,
+    smartbotic::db::storage::LmdbDocumentStore& store,
+    PersistenceManager& persistence,
+    MemoryStore& memStore,
+    CollectionConfigManager& configManager,
+    const std::string& qualifiedParentCollection,
+    const std::string& parentId,
+    const CascadeNotifyFn& notify = nullptr);
+
 bool executeCascade(RelationManager& relations,
                     smartbotic::db::storage::LmdbDocumentStore& store,
                     PersistenceManager& persistence,

+ 6 - 0
tests/CMakeLists.txt

@@ -832,6 +832,10 @@ add_test(NAME test_view_manager_paging COMMAND test_view_manager_paging)
 add_executable(test_relation_manager
     test_relation_manager.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/relations/relation_manager.cpp
+    # v2.11.0 close-out — the create_relation migration op is tested here rather
+    # than in a binary of its own: MigrationRunner needs exactly the MemoryStore
+    # + ViewManager + RelationManager trio this target already links.
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/migrations/migration_runner.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/views/view_manager.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/views/projection.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/memory_store.cpp
@@ -858,4 +862,6 @@ else()
     target_include_directories(test_relation_manager PRIVATE ${SPDLOG_INCLUDE_DIRS})
 endif()
 target_link_libraries(test_relation_manager PRIVATE Threads::Threads)
+# migration_runner.cpp uses RAND_bytes for the generate_secret op.
+target_link_libraries(test_relation_manager PRIVATE OpenSSL::Crypto)
 add_test(NAME test_relation_manager COMMAND test_relation_manager)

+ 39 - 0
tests/load_test/test_relations.sh

@@ -200,5 +200,44 @@ echo
 echo "=== phase: verify the self-heal rebuilt the index and enforcement works again ==="
 "$DRIVER" "127.0.0.1:$PORT" "$PROJECT_A" "$PROJECT_B" verify_selfheal || fail "verify_selfheal phase"
 
+echo
+echo "=== phase: DropProject must clean up its relation declarations (close-out) ==="
+# Declarations live in the GLOBAL _relations collection, so dropping a project
+# used to leave them pointing at a namespace that no longer exists. There is no
+# client surface for CreateProject/DropProject relation cleanup, and no unit
+# fixture builds a DatabaseService, so this is the only level it can be tested at.
+PROJECT_C=relproj_c
+rpc CreateProject "{\"name\":\"$PROJECT_C\"}" >/dev/null
+C_PARENT_B64="$(b64 '{"name":"wf-c1"}')"
+C_CHILD_B64="$(b64 '{"workflowId":"wf-c1"}')"
+rpc Upsert "{\"collection\":\"$PROJECT_C:workflows\",\"id\":\"wf-c1\",\"data\":\"$C_PARENT_B64\"}" >/dev/null
+rpc Upsert "{\"collection\":\"$PROJECT_C:executions\",\"id\":\"ex-c1\",\"data\":\"$C_CHILD_B64\"}" >/dev/null
+OUT="$(rpc CreateRelation "{\"name\":\"$PROJECT_C:wf_exec_c\",\"child\":\"$PROJECT_C:executions\",\"child_field\":\"workflowId\",\"parent\":\"$PROJECT_C:workflows\",\"on_delete\":\"restrict\"}" || true)"
+grep -q '"success": true' <<<"$OUT" \
+    || { echo "$OUT"; fail "could not declare a relation in the throwaway project"; }
+OUT="$(rpc ListRelations "{\"project\":\"$PROJECT_C\"}" || true)"
+grep -q "wf_exec_c" <<<"$OUT" \
+    || { echo "$OUT"; fail "the declaration is not there to begin with - fixture is wrong"; }
+echo "  declared $PROJECT_C:wf_exec_c"
+
+# ⚠ DropProjectResponse's field is `dropped`, not `success` (CreateRelation's
+# is `success`) - checking for the wrong one made this phase fail on a drop that
+# had actually worked.
+OUT="$(rpc DropProject "{\"name\":\"$PROJECT_C\"}" || true)"
+grep -q '"dropped": true' <<<"$OUT" \
+    || { echo "$OUT"; fail "DropProject failed"; }
+
+OUT="$(rpc ListRelations "{\"project\":\"$PROJECT_C\"}" || true)"
+grep -q "wf_exec_c" <<<"$OUT" \
+    && { echo "$OUT"; fail "DropProject left a relation declaration pointing at a dropped namespace"; }
+OUT="$(rpc GetRelationInfo "{\"name\":\"$PROJECT_C:wf_exec_c\"}" || true)"
+grep -q '"found": true' <<<"$OUT" \
+    && { echo "$OUT"; fail "the declaration is still individually resolvable after the drop"; }
+# And the drop must not have taken anyone else's declarations with it.
+OUT="$(rpc ListRelations "{\"project\":\"$PROJECT_A\"}" || true)"
+grep -q "wf_exec" <<<"$OUT" \
+    || { echo "$OUT"; fail "DropProject removed relations belonging to ANOTHER project"; }
+echo "  dropped with the project, and project A's relations are untouched"
+
 echo
 echo "ALL RELATIONS E2E CHECKS PASSED"

+ 533 - 0
tests/test_relation_enforcement.cpp

@@ -58,6 +58,7 @@ using smartbotic::database::describeDeleteImpacts;
 using smartbotic::database::executeCascade;
 using smartbotic::database::findRelationBlocks;
 using smartbotic::database::planCascade;
+using smartbotic::database::ttlExpiryDecision;
 using smartbotic::database::writeCascadeWal;
 using smartbotic::db::storage::LmdbDocumentStore;
 using smartbotic::db::storage::LmdbEnv;
@@ -2079,6 +2080,529 @@ void test_can_reject_writes_gates_the_undo_snapshot() {
     ms.stop();
 }
 
+
+// =========================================================================
+// v2.11.0 close-out — TTL EXPIRY MUST HANDLE CHILDREN LIKE A MANUAL DELETE.
+//
+// Operator's ruling. Before this, MemoryStore::expireDocuments() erased the
+// document and mirrored a DELETE with no relation enforcement whatsoever, so a
+// restrict-protected parent carrying a TTL silently vanished and orphaned every
+// child - strictly the worst of the three available behaviours.
+//
+// Every test below drives the REAL sweeper (MemoryStore::expireDocuments()) with
+// the REAL decision function (relations/relation_cascade.cpp's
+// ttlExpiryDecision(), which DatabaseService binds as the production hook), so
+// what is under test is the shipped path and not a re-statement of it.
+// =========================================================================
+
+// Residency probe. MemoryStore::get() deliberately hides an EXPIRED document
+// (Document::isExpired()), so "is it still there" cannot be asked with get()
+// when the whole point is that the document is past its TTL and still present.
+// getAllDocuments() does not filter.
+bool residentInMemory(MemoryStore& mstore, const std::string& collection,
+                      const std::string& id) {
+    for (const auto& d : mstore.getAllDocuments(collection)) {
+        if (d.id == id) return true;
+    }
+    return false;
+}
+
+// Install the production decision function as the sweeper's hook. Exactly what
+// DatabaseService::ttlExpiryRelationDecision() does, minus the per-project store
+// resolution (this fixture has one store) and the replication notifier.
+void installTtlHook(MemoryStore& mstore, RelationManager& rm, LmdbDocumentStore& store,
+                    PersistenceManager& pm, CollectionConfigManager& cfgManager,
+                    const smartbotic::database::CascadeNotifyFn& notify = nullptr) {
+    mstore.setTtlExpiryRelationHook(
+        [&rm, &store, &pm, &mstore, &cfgManager, notify](const std::string& coll,
+                                                          const std::string& id) {
+            return ttlExpiryDecision(rm, store, pm, mstore, cfgManager, coll, id, notify);
+        });
+}
+
+// A parent carrying a TTL, plus one child, plus a declared relation. Returns
+// nothing; the caller owns every object so the fixtures stay explicit (this
+// file's established style).
+void seedTtlParentAndChild(MemoryStore& mstore, LmdbDocumentStore& store,
+                           RelationManager& rm, OnDelete policy,
+                           const std::string& childField = "workflowId",
+                           const nlohmann::json& childData = {{"workflowId", "wf-1"}}) {
+    RelationInfo rel;
+    rel.name = "default:exec_wf";
+    rel.child = "default:executions";
+    rel.childField = childField;
+    rel.parent = "default:workflows";
+    rel.onDelete = policy;
+    std::string mgrErr;
+    check(rm.createRelation(rel, mgrErr), "declared the relation under test");
+    store.set_relations("executions",
+                        {RelationRef{"exec_wf", childField, "workflows", false, true}});
+
+    // ⚠ expiresAt = 1 (one millisecond after the epoch) rather than "now minus
+    // something": expireDocuments() compares against currentTimeMs(), so 1 is
+    // unambiguously expired on any clock and the test cannot race the sweep.
+    Document parent;
+    parent.id = "wf-1";
+    parent.collection = "workflows";
+    parent.set_data({{"name", "wf-1"}});
+    parent.expiresAt = 1;
+    store.put("workflows", "wf-1", parent);
+    mstore.loadDocument("default:workflows", parent);
+
+    Document child;
+    child.id = "ex-1";
+    child.collection = "executions";
+    child.set_data(childData);
+    store.put("executions", "ex-1", child);
+    mstore.loadDocument("default:executions", child);
+}
+
+// restrict: a MANUAL delete fails, so the expiry must not happen. The document
+// is left in place, its expiry stays armed, and the block is signalled.
+//
+// Would fail before the fix on its first assertion: the pre-v2.11.0 sweeper
+// returned 1 and erased the parent, leaving ex-1 pointing at nothing.
+void test_ttl_restrict_blocks_the_expiry() {
+    TmpEnv t("ttl-restrict");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-restrict-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    seedTtlParentAndChild(mstore, store, rm, OnDelete::Restrict);
+    check(store.relation_index_child_count("exec_wf", "wf-1") == 1, "one child indexed");
+    installTtlHook(mstore, rm, store, p.pm, cfgManager);
+
+    const uint64_t expired = mstore.expireDocuments();
+    check(expired == 0, "the restrict-protected parent was NOT expired");
+    check(residentInMemory(mstore, "default:workflows", "wf-1"),
+          "the parent is still in MemoryStore - it OUTLIVES its TTL, deliberately");
+    check(store.get("workflows", "wf-1").has_value(), "the parent is still in LMDB");
+    check(mstore.get("default:executions", "ex-1").has_value(), "the child was not orphaned");
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 1,
+          "the skip is counted, so an operator can see a document is stuck past its TTL");
+    check(mstore.getStats().expiredCount == 0, "and it is not counted as expired");
+
+    // The expiry must stay ARMED so the next sweep retries - a skip that also
+    // dropped the expiration index entry would leave the document permanently
+    // unexpirable even after the last child went away.
+    const uint64_t again = mstore.expireDocuments();
+    check(again == 0, "still blocked on the second sweep");
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 2,
+          "the second sweep re-examined it, so the expiry is still armed");
+
+    // And once the child is gone, the same sweep expires it - the block is the
+    // relation's, not a permanent quarantine.
+    store.del("executions", "ex-1");
+    check(mstore.remove("default:executions", "ex-1"), "child removed");
+    const uint64_t third = mstore.expireDocuments();
+    check(third == 1, "with no children left, the retry finally expires the parent");
+    check(!mstore.get("default:workflows", "wf-1").has_value(), "the parent is gone now");
+
+    p.pm.stop();
+    mstore.stop();
+}
+
+// cascade with a SCALAR reference: children are deleted, exactly as
+// executeCascade does for a manual delete, and the WAL carries every mutation.
+//
+// Would fail before the fix: the old sweeper wrote no WAL entry for ex-1 at all
+// (it only mirrored a DELETE for the parent), so walHasDelete for the child is
+// the assertion that a hand-rolled, LMDB-only sweeper cascade cannot satisfy.
+void test_ttl_cascade_deletes_children_like_a_manual_delete() {
+    TmpEnv t("ttl-cascade");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-cascade-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    seedTtlParentAndChild(mstore, store, rm, OnDelete::Cascade);
+
+    // The notify wiring: a TTL-driven cascade must drive replication and
+    // Subscribe events per mutation, or a follower never sees the child
+    // deletions and diverges permanently.
+    struct Notification { std::string collection, id; bool hadDoc; smartbotic::database::EventType et; };
+    std::vector<Notification> notifications;
+    auto notify = [&](const std::string& coll, const std::string& id,
+                      const std::optional<Document>& doc, smartbotic::database::EventType et) {
+        notifications.push_back({coll, id, doc.has_value(), et});
+    };
+    installTtlHook(mstore, rm, store, p.pm, cfgManager, notify);
+
+    const uint64_t expired = mstore.expireDocuments();
+    check(expired == 1, "the parent was expired");
+    check(!mstore.get("default:executions", "ex-1").has_value(),
+          "the child was CASCADED, not orphaned");
+    check(!store.get("executions", "ex-1").has_value(), "and it is gone from LMDB too");
+    check(!store.get("workflows", "wf-1").has_value(), "the parent is gone from LMDB");
+    check(store.relation_index_child_count("exec_wf", "wf-1") == 0,
+          "the reverse-index posting went with it");
+    check(notifications.size() == 2,
+          "replication/events fired once per child plus once for the parent");
+
+    // The property a hand-rolled sweeper cascade breaks: MemoryStore is rebuilt
+    // from snapshot + WAL, NEVER from LMDB, so a cascade whose child deletions
+    // exist only in LMDB has them RESURRECTED on the next boot.
+    p.pm.stop();
+    auto entries = replayRawWal(p.path);
+    check(walHasDelete(entries, "default:executions", "ex-1"),
+          "the WAL carries the CHILD's delete - without this the next boot resurrects it");
+    check(walHasDelete(entries, "default:workflows", "wf-1"),
+          "the WAL carries the parent's delete");
+
+    mstore.stop();
+}
+
+// set_null with a SCALAR reference: the field is nulled and the child KEPT.
+void test_ttl_set_null_nulls_the_scalar_and_keeps_the_child() {
+    TmpEnv t("ttl-setnull");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-setnull-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    seedTtlParentAndChild(mstore, store, rm, OnDelete::SetNull);
+    installTtlHook(mstore, rm, store, p.pm, cfgManager);
+
+    const uint64_t expired = mstore.expireDocuments();
+    check(expired == 1, "the parent was expired");
+    auto child = store.get("executions", "ex-1");
+    check(child.has_value(), "the child SURVIVED - set_null keeps the document");
+    if (child) {
+        check(child->data().contains("workflowId") && child->data()["workflowId"].is_null(),
+              "and its reference is null, not deleted and not left dangling");
+    }
+    auto memChild = mstore.get("default:executions", "ex-1");
+    check(memChild.has_value(), "the child survived in MemoryStore too");
+    if (memChild) {
+        check(memChild->data()["workflowId"].is_null(), "with the same nulled reference");
+    }
+    check(store.relation_index_child_count("exec_wf", "wf-1") == 0,
+          "the posting is gone, since the reference no longer names wf-1");
+
+    p.pm.stop();
+    auto entries = replayRawWal(p.path);
+    check(walHasUpdate(entries, "default:executions", "ex-1"),
+          "the child's UPDATE is in the WAL - the mutation survives a restart");
+    mstore.stop();
+}
+
+// An ARRAY-valued reference: the id is PULLED and the document KEPT, under
+// cascade as well as set_null. The place a MySQL mental model actively
+// misleads, and the same rule the request path follows - deleting a node
+// because one of its three credentials expired would be worse than useless.
+void test_ttl_array_reference_pulls_the_id_and_keeps_the_document() {
+    TmpEnv t("ttl-array");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-array-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    // cascade, deliberately: the array rule must hold under the policy a
+    // MySQL-trained reader expects to delete the row.
+    seedTtlParentAndChild(mstore, store, rm, OnDelete::Cascade, "workflowIds",
+                          {{"workflowIds", {"wf-1", "wf-2"}}});
+    check(store.relation_index_child_count("exec_wf", "wf-1") == 1,
+          "the array element is indexed as a posting");
+    installTtlHook(mstore, rm, store, p.pm, cfgManager);
+
+    const uint64_t expired = mstore.expireDocuments();
+    check(expired == 1, "the parent was expired");
+    auto child = store.get("executions", "ex-1");
+    check(child.has_value(), "the child SURVIVED - an array match never deletes the document");
+    if (child) {
+        const auto& ids = child->data()["workflowIds"];
+        check(ids.is_array() && ids.size() == 1, "exactly one id was pulled");
+        check(ids.is_array() && !ids.empty() && ids[0] == "wf-2",
+              "and it was the expiring parent's id, not the surviving one");
+    }
+    check(store.relation_index_child_count("exec_wf", "wf-1") == 0, "posting for wf-1 gone");
+    check(store.relation_index_child_count("exec_wf", "wf-2") == 1, "posting for wf-2 intact");
+
+    p.pm.stop();
+    mstore.stop();
+}
+
+// no_action: the expiry proceeds and the reference is left dangling, which is
+// what it did before this change and must keep doing.
+void test_ttl_no_action_expires_and_leaves_the_reference_dangling() {
+    TmpEnv t("ttl-noaction");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-noaction-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    seedTtlParentAndChild(mstore, store, rm, OnDelete::NoAction);
+    installTtlHook(mstore, rm, store, p.pm, cfgManager);
+
+    const uint64_t expired = mstore.expireDocuments();
+    check(expired == 1, "no_action does not block the expiry");
+    check(!mstore.get("default:workflows", "wf-1").has_value(), "the parent expired");
+    auto child = mstore.get("default:executions", "ex-1");
+    check(child.has_value(), "the child is untouched");
+    if (child) {
+        check(child->data()["workflowId"] == "wf-1",
+              "and its reference still names the gone parent - dangling, as declared");
+    }
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 0, "nothing was blocked");
+
+    p.pm.stop();
+    mstore.stop();
+}
+
+// SCOPE: the common case - no relation names this collection as a parent - must
+// behave exactly as it did before, including the LMDB DELETE mirror and the
+// EXPIRE event. This is the regression guard for the two-phase rewrite of
+// expireDocuments(), which is what actually changed for every install.
+void test_ttl_with_no_relations_expires_exactly_as_before() {
+    TmpEnv t("ttl-plain");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-plain-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    std::atomic<bool> healthy{true};
+    std::atomic<uint64_t> drift{0};
+    mstore.setDocumentStoreMirror(
+        [&store](std::string_view) -> smartbotic::db::storage::DocumentStore* { return &store; },
+        &healthy, &drift);
+
+    int expireEvents = 0;
+    mstore.setEventCallback([&](const smartbotic::database::DatabaseEvent& ev) {
+        if (ev.type == smartbotic::database::EventType::EXPIRE) ++expireEvents;
+    });
+
+    Document doc;
+    doc.id = "s-1";
+    doc.collection = "sessions";
+    doc.set_data({{"token", "abc"}});
+    doc.expiresAt = 1;
+    store.put("sessions", "s-1", doc);
+    mstore.loadDocument("default:sessions", doc);
+
+    // Hook installed, with the relation cache EMPTY - the early-return path.
+    installTtlHook(mstore, rm, store, p.pm, cfgManager);
+
+    const uint64_t expired = mstore.expireDocuments();
+    check(expired == 1, "an unrelated document still expires");
+    check(!mstore.get("default:sessions", "s-1").has_value(), "gone from MemoryStore");
+    check(!store.get("sessions", "s-1").has_value(),
+          "and the DELETE was mirrored to LMDB, as before");
+    check(expireEvents == 1, "exactly one EXPIRE event, as before");
+    check(mstore.getStats().expiredCount == 1, "counted as expired");
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 0, "nothing blocked");
+
+    // A document whose TTL is in the future is not touched by any of this.
+    Document later;
+    later.id = "s-2";
+    later.collection = "sessions";
+    later.set_data({{"token", "def"}});
+    later.expiresAt = 1;
+    later.expiresAt = static_cast<uint64_t>(1) << 62;   // far future
+    mstore.loadDocument("default:sessions", later);
+    check(mstore.expireDocuments() == 0, "an unexpired document is left alone");
+    check(mstore.get("default:sessions", "s-2").has_value(), "and is still there");
+
+    p.pm.stop();
+    mstore.stop();
+}
+
+// THE ORDERING PROPERTY, observed rather than asserted from the outside: the
+// cascade's WAL entries are durable BEFORE the LMDB transaction commits.
+//
+// Induced honestly, with the v2.4.4 identity sentinel: the child sub-db's
+// sentinel is overwritten with the wrong name using raw LMDB, so
+// commitCascadeLmdb()'s open_for_write() throws and its WriteTxn aborts
+// UNWRITTEN. Everything before it has already happened. If the sweeper wrote
+// LMDB first (or hand-rolled its own cascade), the WAL would be empty here and
+// the child's deletion would be lost on the next boot.
+void test_ttl_cascade_wal_is_durable_before_the_lmdb_commit() {
+    TmpEnv t("ttl-wal-first");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("ttl-wal-first-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    seedTtlParentAndChild(mstore, store, rm, OnDelete::Cascade);
+    installTtlHook(mstore, rm, store, p.pm, cfgManager);
+
+    // Overwrite the child sub-db's identity sentinel with a name that is not
+    // its own. open_for_write verifies it on every cached-handle reuse.
+    {
+        MDB_txn* txn = nullptr;
+        check(mdb_txn_begin(t.env.raw(), nullptr, 0, &txn) == 0, "tamper txn opened");
+        MDB_dbi dbi = 0;
+        check(mdb_dbi_open(txn, "executions", 0, &dbi) == 0, "child sub-db opened for tampering");
+        const std::string_view sentinel{"\0__subdb_identity__", 19};
+        MDB_val k{sentinel.size(), const_cast<char*>(sentinel.data())};
+        std::string wrong = "not_executions";
+        MDB_val v{wrong.size(), wrong.data()};
+        check(mdb_put(txn, dbi, &k, &v, 0) == 0, "sentinel overwritten");
+        check(mdb_txn_commit(txn) == 0, "tamper txn committed");
+    }
+
+    const uint64_t expired = mstore.expireDocuments();
+    check(expired == 0, "the sweeper did NOT expire the parent when the cascade faulted");
+    check(residentInMemory(mstore, "default:workflows", "wf-1"),
+          "the parent is still in MemoryStore - never expired behind a failed cascade");
+    check(store.get("workflows", "wf-1").has_value(),
+          "and still in LMDB - commitCascadeLmdb's WriteTxn aborted unwritten");
+    check(store.get("executions", "ex-1").has_value(), "the child is still in LMDB too");
+    check(mstore.getStats().ttlExpiryBlockedByRelation == 1, "the skip is counted");
+
+    // ...and yet the WAL already describes the whole cascade. That is the
+    // ordering: WAL fsynced, THEN the LMDB commit attempted.
+    p.pm.stop();
+    auto entries = replayRawWal(p.path);
+    check(walHasDelete(entries, "default:executions", "ex-1"),
+          "the child's DELETE is already durable in the WAL, before any LMDB commit");
+    check(walHasDelete(entries, "default:workflows", "wf-1"),
+          "so is the parent's - the next boot's replay applies the cascade regardless");
+
+    mstore.stop();
+}
+
+// =========================================================================
+// v2.11.0 close-out — BOOT-PASS DRIFT MUST NOT DISABLE DESTRUCTIVE POLICIES.
+//
+// mirror_drift_count_ is never reset, and the post-replay re-mirror pass bumps
+// it for every row it cannot repair - a condition runPendingRemirror() treats as
+// advisory and which recurs on every boot, since the restart replays the same
+// failing row. Gated on the raw counter, ONE unrepairable row refused every
+// cascade/set_null delete and every CreateRelation for the process's life, with
+// an error message that told the operator to restart.
+// =========================================================================
+void test_boot_pass_drift_does_not_disable_the_cascade() {
+    TmpEnv t("rel-drift-baseline");
+    LmdbDocumentStore store(t.env);
+    TmpPersistence p("rel-drift-baseline-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+
+    std::atomic<bool> healthy{true};
+    std::atomic<uint64_t> drift{0};
+    mstore.setDocumentStoreMirror(
+        [&store](std::string_view) -> smartbotic::db::storage::DocumentStore* { return &store; },
+        &healthy, &drift);
+
+    // The boot path: three rows the re-mirror pass could not write. Health is
+    // deliberately NOT flipped (the v2.8.1 lesson), so drift is the only signal.
+    drift.store(3);
+    check(mstore.mirrorDriftCount() == 3, "the raw counter still reports every stale row");
+    mstore.markMirrorDriftBaseline();       // what initialize() does last
+    check(mstore.mirrorDriftBaseline() == 3, "the boot path's drift is in the baseline");
+    check(mstore.mirrorDriftSinceBaseline() == 0,
+          "and nothing has drifted since READY");
+
+    seedTtlParentAndChild(mstore, store, rm, OnDelete::Cascade);
+
+    bool parentExisted = false;
+    try {
+        parentExisted = executeCascade(rm, store, p.pm, mstore, cfgManager,
+                                       "default:workflows", "wf-1", nullptr);
+    } catch (const std::exception& e) {
+        check(false, "the cascade was refused because of drift the BOOT PASS accrued");
+        std::cerr << "  (" << e.what() << ")\n";
+    }
+    check(parentExisted, "the cascade ran despite 3 rows the boot pass left stale");
+    check(!store.get("executions", "ex-1").has_value(), "and it actually cascaded");
+
+    // A LIVE write's drift still latches and still refuses - that half was
+    // right, and re-basing the gate must not have removed it.
+    drift.fetch_add(1);
+    check(mstore.mirrorDriftSinceBaseline() == 1, "post-READY drift is visible");
+    Document parent2;
+    parent2.id = "wf-2";
+    parent2.collection = "workflows";
+    parent2.set_data({{"name", "wf-2"}});
+    store.put("workflows", "wf-2", parent2);
+    mstore.loadDocument("default:workflows", parent2);
+    Document child2;
+    child2.id = "ex-2";
+    child2.collection = "executions";
+    child2.set_data({{"workflowId", "wf-2"}});
+    store.put("executions", "ex-2", child2);
+    mstore.loadDocument("default:executions", child2);
+
+    bool refused = false;
+    try {
+        executeCascade(rm, store, p.pm, mstore, cfgManager, "default:workflows", "wf-2", nullptr);
+    } catch (const CascadeBlocked&) {
+        refused = true;
+    }
+    check(refused, "drift accrued AFTER READY still refuses the cascade");
+    check(store.get("executions", "ex-2").has_value(), "and nothing was touched");
+
+    p.pm.stop();
+    mstore.stop();
+}
+
+// v2.11.0 close-out — runPendingRemirror() must RELEASE the id list.
+// RecoveryOutcome lives as DatabaseService::recovery_outcome_ for the process
+// lifetime; the pass runs exactly once and nothing reads the list afterwards, so
+// keeping it resident pinned ~100 bytes per replayed row (an install's whole
+// history under --recovery-mode=wal_only) for nothing.
+void test_pending_remirror_list_is_released_after_the_pass() {
+    TmpPersistence p("remirror-release");
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+
+    smartbotic::database::RecoveryOutcome outcome;
+    for (int i = 0; i < 500; ++i) {
+        outcome.pendingRemirror.emplace_back("default:widgets", "w-" + std::to_string(i));
+    }
+    check(outcome.pendingRemirror.size() == 500, "the list starts populated");
+
+    p.pm.runPendingRemirror(mstore, outcome);
+
+    check(outcome.pendingRemirror.empty(),
+          "the id list is cleared once the pass has run");
+    check(outcome.pendingRemirror.capacity() == 0,
+          "and its CAPACITY is released - clear() alone keeps the whole allocation");
+
+    mstore.stop();
+}
+
 int main() {
     std::cout << "=== test_relation_enforcement ===\n";
     test_index_follows_the_child_field();
@@ -2112,6 +2636,15 @@ int main() {
     test_remirror_windows_are_bounded_by_bytes_and_by_count();
     test_cascade_refuses_while_the_mirror_is_unhealthy_or_drifted();
     test_can_reject_writes_gates_the_undo_snapshot();
+    test_ttl_restrict_blocks_the_expiry();
+    test_ttl_cascade_deletes_children_like_a_manual_delete();
+    test_ttl_set_null_nulls_the_scalar_and_keeps_the_child();
+    test_ttl_array_reference_pulls_the_id_and_keeps_the_document();
+    test_ttl_no_action_expires_and_leaves_the_reference_dangling();
+    test_ttl_with_no_relations_expires_exactly_as_before();
+    test_ttl_cascade_wal_is_durable_before_the_lmdb_commit();
+    test_boot_pass_drift_does_not_disable_the_cascade();
+    test_pending_remirror_list_is_released_after_the_pass();
 
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;

+ 201 - 0
tests/test_relation_manager.cpp

@@ -13,9 +13,15 @@
 
 #include <nlohmann/json.hpp>
 
+#include <filesystem>
+#include <fstream>
+#include <unistd.h>
+
 #include "document.hpp"
 #include "memory_store.hpp"
+#include "migrations/migration_runner.hpp"
 #include "relations/relation_manager.hpp"
+#include "views/view_manager.hpp"
 
 using namespace smartbotic::database;
 
@@ -330,6 +336,199 @@ void test_on_delete_is_validated_not_coerced() {
           "and it falls back to restrict, the non-destructive direction");
 }
 
+
+// =========================================================================
+// v2.11.0 close-out — the `create_relation` MIGRATION OP.
+//
+// Declaring schema in migration files is how consumers ship views
+// (shadowman-cpp: /opt/shadowman/share/shadowman/migrations/json, callerai:
+// /etc/callerai/migrations), and until this op existed they could not declare a
+// relation at all. Modelled on create_view: same file shape, same idempotency,
+// and it goes through RelationManager::createRelation so the same-project rule
+// and the on_delete validation apply rather than being bypassed.
+// =========================================================================
+
+std::filesystem::path makeMigrationDir(const std::string& tag) {
+    auto dir = std::filesystem::temp_directory_path() /
+               ("mig-relation-" + tag + "-" + std::to_string(::getpid()));
+    std::filesystem::remove_all(dir);
+    std::filesystem::create_directories(dir);
+    return dir;
+}
+
+void writeMigration(const std::filesystem::path& dir, const std::string& file,
+                    const std::string& body) {
+    std::ofstream out(dir / file);
+    out << body;
+}
+
+void test_create_relation_migration_op_declares_and_is_idempotent() {
+    auto dir = makeMigrationDir("basic");
+    writeMigration(dir, "001_relations.json", R"({
+      "version": "001",
+      "name": "declare_exec_wf",
+      "operations": [
+        {"type": "create_collection", "collection": "workflows"},
+        {"type": "create_collection", "collection": "executions"},
+        {"type": "create_relation",
+         "name": "exec_wf",
+         "child": "executions",
+         "child_field": "workflowId",
+         "parent": "workflows",
+         "on_delete": "cascade",
+         "validate_on_write": true}
+      ]
+    })");
+
+    Fixture f;
+    ViewManager vm(f.store);
+    RelationManager rm(f.store);
+    MigrationRunner::Config cfg;
+    cfg.directory = dir;
+
+    {
+        MigrationRunner runner(f.store, vm, rm, cfg);
+        check(runner.runMigrations(), "the migration ran");
+    }
+
+    // Bare names in a migration file qualify to `default:`, exactly as every
+    // other collection name in a migration file does.
+    auto got = rm.getRelation("default:exec_wf");
+    check(got.has_value(), "the create_relation op DECLARED the relation");
+    if (got) {
+        check(got->child == "default:executions", "child is project-qualified");
+        check(got->childField == "workflowId", "child_field carried through");
+        check(got->parent == "default:workflows", "parent is project-qualified");
+        check(got->onDelete == OnDelete::Cascade, "on_delete was parsed, not defaulted");
+        check(got->validateOnWrite, "validate_on_write carried through");
+    }
+
+    // Idempotency has TWO layers and both matter. The runner skips an
+    // already-applied migration file, so re-running is a no-op at that level;
+    // the op itself must ALSO tolerate "already exists", which is what a
+    // consumer re-shipping the same declaration under a new version number
+    // hits.
+    {
+        MigrationRunner runner(f.store, vm, rm, cfg);
+        check(runner.runMigrations(), "re-running the same migrations still succeeds");
+    }
+    writeMigration(dir, "002_again.json", R"({
+      "version": "002",
+      "name": "declare_exec_wf_again",
+      "operations": [
+        {"type": "create_relation",
+         "name": "exec_wf", "child": "executions",
+         "child_field": "workflowId", "parent": "workflows",
+         "on_delete": "cascade"}
+      ]
+    })");
+    {
+        MigrationRunner runner(f.store, vm, rm, cfg);
+        check(runner.runMigrations(),
+              "a SECOND migration re-declaring the same relation succeeds - "
+              "'already exists' is not a failure on replay");
+    }
+    check(rm.listRelations("default").size() == 1,
+          "and it did not duplicate the declaration");
+
+    std::filesystem::remove_all(dir);
+}
+
+void test_create_relation_migration_op_validates() {
+    // A typo'd on_delete must FAIL the migration, not silently arm restrict.
+    // The reason it matters more here than at the RPC: a typo in a file that
+    // ships in a deb would otherwise be wrong on every install, forever.
+    {
+        auto dir = makeMigrationDir("typo");
+        writeMigration(dir, "001_typo.json", R"({
+          "version": "001", "name": "typo",
+          "operations": [
+            {"type": "create_relation", "name": "bad_rel", "child": "executions",
+             "child_field": "workflowId", "parent": "workflows",
+             "on_delete": "Cascade"}
+          ]
+        })");
+        Fixture f;
+        ViewManager vm(f.store);
+        RelationManager rm(f.store);
+        MigrationRunner::Config cfg;
+        cfg.directory = dir;
+        MigrationRunner runner(f.store, vm, rm, cfg);
+        check(!runner.runMigrations(), "an unrecognised on_delete FAILS the migration");
+        check(!rm.getRelation("default:bad_rel").has_value(),
+              "and nothing was declared - not coerced to restrict");
+        std::filesystem::remove_all(dir);
+    }
+
+    // A cross-project declaration must be refused by RelationManager, which is
+    // the whole point of routing through createRelation rather than writing the
+    // `_relations` record directly.
+    {
+        auto dir = makeMigrationDir("xproj");
+        writeMigration(dir, "001_xproj.json", R"({
+          "version": "001", "name": "xproj",
+          "operations": [
+            {"type": "create_relation", "name": "a:rel", "child": "a:executions",
+             "child_field": "workflowId", "parent": "b:workflows"}
+          ]
+        })");
+        Fixture f;
+        ViewManager vm(f.store);
+        RelationManager rm(f.store);
+        MigrationRunner::Config cfg;
+        cfg.directory = dir;
+        MigrationRunner runner(f.store, vm, rm, cfg);
+        check(!runner.runMigrations(),
+              "a cross-project relation is refused through the migration op too");
+        check(!rm.getRelation("a:rel").has_value(), "and nothing was declared");
+        std::filesystem::remove_all(dir);
+    }
+
+    // Missing required fields fail rather than declaring a half-relation.
+    {
+        auto dir = makeMigrationDir("missing");
+        writeMigration(dir, "001_missing.json", R"({
+          "version": "001", "name": "missing",
+          "operations": [
+            {"type": "create_relation", "name": "half_rel", "child": "executions"}
+          ]
+        })");
+        Fixture f;
+        ViewManager vm(f.store);
+        RelationManager rm(f.store);
+        MigrationRunner::Config cfg;
+        cfg.directory = dir;
+        MigrationRunner runner(f.store, vm, rm, cfg);
+        check(!runner.runMigrations(), "a create_relation missing child_field/parent fails");
+        check(!rm.getRelation("default:half_rel").has_value(), "and declares nothing");
+        std::filesystem::remove_all(dir);
+    }
+
+    // Absent on_delete means the documented default, and is NOT an error.
+    {
+        auto dir = makeMigrationDir("default-od");
+        writeMigration(dir, "001_default.json", R"({
+          "version": "001", "name": "defaulted",
+          "operations": [
+            {"type": "create_relation", "name": "def_rel", "child": "executions",
+             "child_field": "workflowId", "parent": "workflows"}
+          ]
+        })");
+        Fixture f;
+        ViewManager vm(f.store);
+        RelationManager rm(f.store);
+        MigrationRunner::Config cfg;
+        cfg.directory = dir;
+        MigrationRunner runner(f.store, vm, rm, cfg);
+        check(runner.runMigrations(), "an absent on_delete is accepted");
+        auto got = rm.getRelation("default:def_rel");
+        check(got.has_value(), "and the relation is declared");
+        check(got && got->onDelete == OnDelete::Restrict, "with restrict, the documented default");
+        check(got && !got->validateOnWrite, "and validate_on_write defaulting to false");
+        std::filesystem::remove_all(dir);
+    }
+}
+
 int main() {
     std::cout << "=== test_relation_manager ===\n";
     test_relations_are_project_scoped_and_survive_reload();
@@ -339,6 +538,8 @@ int main() {
     test_bare_and_qualified_names_are_the_same_relation();
     test_load_rekeys_a_legacy_bare_named_record();
     test_on_delete_is_validated_not_coerced();
+    test_create_relation_migration_op_declares_and_is_idempotent();
+    test_create_relation_migration_op_validates();
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;
 }