소스 검색

feat(relations): T12 - cascade and set_null, WAL-first

cascade and set_null are now destructive; restrict/no_action unchanged.

- relations/relation_cascade.{hpp,cpp}: planCascade (read-only resolve),
  writeCascadeWal (WAL + fsync, step 3), commitCascadeLmdb (one atomic
  WriteTxn, step 4), applyCascadeToMemory (step 5), executeCascade
  (orchestrates 3->4->5). Array-valued references collapse cascade and
  set_null to the same behaviour: pull the id, keep the document. Only a
  scalar reference under cascade deletes the child.
- document_store_lmdb: beginWrite()/commitAndCache() let external code
  (the cascade) open one caller-owned WriteTxn spanning parent + every
  child without leaking the private LmdbEnv&.
- memory_store: unloadDocument(), the delete-side counterpart to
  loadDocument()/loadDocumentWithHistory() - memory-only, no WAL/mirror,
  since the cascade caller already did those explicitly and the ordinary
  remove()/update() would double-log.
- persistence_manager: flushWal() forces a synchronous fsync before the
  LMDB commit, per the WAL-first design (MemoryStore rebuilds from WAL at
  boot, not LMDB).
- wal.cpp: WriteAheadLog::sync() now does a REAL fsync via a reopened fd,
  not just ofstream::flush() - a pre-existing gap that would have made
  flushWal() a name without the property it promises.
- database_grpc_impl: Delete() routes through executeCascade() when the
  parent has any Cascade/SetNull relation and restrict has already
  permitted the delete; gated on relationsEnforced like restrict is.
- Three new tests in test_relation_enforcement.cpp: scalar cascade with
  WAL verified via a raw WAL replay, array-pull-keeps-document, and a
  crash-between-WAL-and-commit simulation proving recovery converges from
  WAL alone.

Verified by hand against a live server incl. a real kill -9 + restart:
no cascade mutation was resurrected.
fszontagh 1 개월 전
부모
커밋
7d5d26a855

+ 5 - 2
cli/main.cpp

@@ -745,8 +745,11 @@ bool execCommand(smartbotic::database::Client& client,
                 printError("usage: relation-create <name> <child> <child_field> <parent> "
                            "[on_delete] [validate_on_write]\n"
                            "  on_delete: restrict (default) | cascade | set_null | no_action\n"
-                           "  note: cascade/set_null are accepted and persisted but currently "
-                           "behave as permit");
+                           "  note (v2.11.0+): cascade deletes the referencing document "
+                           "(scalar reference) or pulls the id from the array and keeps the "
+                           "document (array reference - cascade and set_null are the same "
+                           "for arrays); set_null nulls the scalar field. no_action permits "
+                           "the delete and leaves the reference dangling.");
                 return false;
             }
             const std::string onDelete = params.size() > 4 ? params[4] : "restrict";

+ 1 - 0
service/CMakeLists.txt

@@ -60,6 +60,7 @@ set(DATABASE_SERVICE_SOURCES
     src/relations/relation_manager.cpp
     src/relations/relation_index.cpp
     src/relations/relation_enforcement.cpp
+    src/relations/relation_cascade.cpp
     src/config/collection_config_manager.cpp
     src/config/config_loader.cpp
     src/storage/lmdb_env.cpp

+ 45 - 6
service/src/database_grpc_impl.cpp

@@ -6,6 +6,7 @@
 #include "storage/document_store_lmdb.hpp"
 #include "storage/project_store.hpp"
 #include "relations/relation_enforcement.hpp"
+#include "relations/relation_cascade.hpp"
 #include "json_parse.hpp"
 #include "persistence/wal.hpp"
 #include "views/projection.hpp"
@@ -818,12 +819,13 @@ grpc::Status DatabaseGrpcImpl::Delete(
             "cannot write to view '" + request->collection() + "': views are read-only");
     }
 
-    // v2.11.0 T6a — referential integrity. After the access gate (authorisation
-    // before integrity - a caller must not learn about child counts on a
-    // collection they cannot read), before store_.remove() so a blocked delete
-    // never mutates anything. Only meaningful on the LMDB substrate, where the
-    // reverse index lives; if this project has no LMDB store for some reason,
-    // there is nothing to enforce against and the delete proceeds as before.
+    // v2.11.0 T6a/T12 — referential integrity. After the access gate
+    // (authorisation before integrity - a caller must not learn about
+    // child counts on a collection they cannot read), before store_.remove()
+    // so a blocked delete never mutates anything. Only meaningful on the
+    // LMDB substrate, where the reverse index lives; if this project has no
+    // LMDB store for some reason, there is nothing to enforce against and
+    // the delete proceeds as before.
     if (!request->collection().empty() && request->collection()[0] != '_') {
         try {
             const auto rc = smartbotic::database::resolveCollection(request->collection());
@@ -838,6 +840,43 @@ grpc::Status DatabaseGrpcImpl::Delete(
                     if (!enforcer.canDelete(request->collection(), request->id(), err)) {
                         return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, err);
                     }
+
+                    // v2.11.0 T12 — cascade/set_null. restrict has already had
+                    // its say (canDelete, above); this is reached only when the
+                    // delete is otherwise permitted. relationsEnforced gates
+                    // this exactly like it gates restrict: disabled means every
+                    // on_delete policy is skipped, not just restrict, so a
+                    // collection with enforcement off keeps today's permissive
+                    // (dangling-reference) behaviour rather than half-enforcing.
+                    if (cfg.relationsEnforced) {
+                        bool hasDestructive = false;
+                        for (const auto& r : relation_manager_.relationsWithParent(request->collection())) {
+                            if (r.onDelete == smartbotic::database::OnDelete::Cascade ||
+                                r.onDelete == smartbotic::database::OnDelete::SetNull) {
+                                hasDestructive = true;
+                                break;
+                            }
+                        }
+                        if (hasDestructive) {
+                            bool parentExisted = false;
+                            try {
+                                parentExisted = smartbotic::database::executeCascade(
+                                    relation_manager_, *lmdb, persistence_, store_,
+                                    request->collection(), request->id());
+                            } catch (const std::exception& e) {
+                                spdlog::error("relations: cascade delete of '{}/{}' failed: {}",
+                                             request->collection(), request->id(), e.what());
+                                return grpc::Status(grpc::StatusCode::INTERNAL,
+                                    std::string("cascade delete failed: ") + e.what());
+                            }
+                            // executeCascade already deleted the parent (and its
+                            // vector, if any) atomically with every child
+                            // mutation — do not fall through to the plain
+                            // store_.remove() path below.
+                            response->set_deleted(parentExisted);
+                            return grpc::Status::OK;
+                        }
+                    }
                 }
             }
         } catch (const std::invalid_argument& e) {

+ 47 - 0
service/src/memory_store.cpp

@@ -825,6 +825,53 @@ bool MemoryStore::remove(const std::string& collection, const std::string& id) {
     return true;
 }
 
+bool MemoryStore::unloadDocument(const std::string& collection, const std::string& id) {
+    CollectionData* coll = getOrCreateCollection(collection);
+
+    // v1.7.0 T6 quiesce: mark before taking the write lock. Same reasoning
+    // as remove() — this still mutates coll->documents under the same lock.
+    InFlightWriteGuard inFlight(*coll, id);
+
+    std::unique_lock<std::shared_mutex> lock(coll->mutex);
+
+    auto it = coll->documents.find(id);
+    if (it == coll->documents.end()) {
+        return false;
+    }
+
+    uint64_t docSize = estimateDocumentSize(it->second);
+
+    if (it->second.expiresAt > 0) {
+        removeFromExpirationIndex(*coll, id, it->second.expiresAt);
+    }
+
+    saveToHistory(*coll, it->second);
+    coll->documents.erase(it);
+    removeVector(*coll, id);
+    coll->updatedAt = currentTimeMs();
+
+    // Deliberately NO mirrorWriteToDocStore()/mirrorVectorToDocStore() call
+    // here — see the header comment. The caller already committed this
+    // exact mutation to LMDB, atomically with the rest of the cascade,
+    // before calling this.
+
+    lock.unlock();
+
+    {
+        std::lock_guard<std::mutex> statsLock(statsMutex_);
+        stats_.totalDocuments--;
+        stats_.deleteCount++;
+    }
+
+    estimatedMemoryBytes_.fetch_sub(docSize, std::memory_order_relaxed);
+
+    // Deliberately NO emitPersist()/emitEvent() call — see the header
+    // comment. The caller already WAL-logged this exact mutation before
+    // the LMDB commit that preceded this step.
+
+    return true;
+}
+
 bool MemoryStore::exists(const std::string& collection, const std::string& id) const {
     const CollectionData* coll = getCollection(collection);
     if (!coll) {

+ 24 - 0
service/src/memory_store.hpp

@@ -265,6 +265,30 @@ public:
      */
     bool remove(const std::string& collection, const std::string& id);
 
+    /**
+     * v2.11.0 T12 — memory-only counterpart to loadDocument()/
+     * loadDocumentWithHistory() for a delete. Cascade (relations/
+     * relation_cascade.cpp) WAL-logs and commits the LMDB mutation itself,
+     * in that order, BEFORE this runs — see the Delete handler's cascade
+     * path in database_grpc_impl.cpp. Calling the ordinary remove() here
+     * would re-log the delete to WAL (a redundant second entry) and
+     * re-attempt the LMDB mirror (a redundant WriteTxn against a row
+     * already gone) for a mutation that is already durable — harmless but
+     * wasteful, and it blurs the "WAL and the LMDB commit each happen
+     * exactly once, in that order" property the crash test depends on.
+     *
+     * Mirrors what remove() does structurally (history, expiration index,
+     * vector, stats) MINUS the two things the caller already handled: the
+     * LMDB mirror and the WAL log. No persist/event callback fires either,
+     * for the same reason — this step is memory bookkeeping only.
+     *
+     * Returns false if the id was not present, which is expected and not
+     * an error: MemoryStore is a bounded cache and may simply not hold a
+     * child the cascade is deleting from LMDB (see the v2.4.4 notes on
+     * eviction). Nothing to undo there.
+     */
+    bool unloadDocument(const std::string& collection, const std::string& id);
+
     /**
      * Check if a document exists.
      */

+ 5 - 0
service/src/persistence/persistence_manager.cpp

@@ -397,6 +397,11 @@ uint64_t PersistenceManager::logDelete(const std::string& collection, const std:
     return seq;
 }
 
+void PersistenceManager::flushWal() {
+    if (!wal_) return;
+    wal_->sync();
+}
+
 uint64_t PersistenceManager::logUpsert(const std::string& collection, const Document& doc,
                                         const std::string& originNodeId) {
     if (!running_.load()) return 0;

+ 20 - 0
service/src/persistence/persistence_manager.hpp

@@ -175,6 +175,26 @@ public:
     uint64_t logVecDelete(const std::string& collection, const std::string& docId,
                           const std::string& originNodeId = "");
 
+    /**
+     * v2.11.0 T12 — force the WAL to fsync NOW, synchronously.
+     *
+     * The background walSyncLoop() fsyncs every walSyncIntervalMs (100ms by
+     * default), which is fine for the ordinary per-document write path
+     * (that path mirrors to LMDB first, under MemoryStore's collection
+     * lock, and only logs to WAL — durably or not — afterward; see
+     * memory_store.cpp's remove()/update()). A cascade delete is the one
+     * place that ordering is deliberately reversed: WAL entries for the
+     * parent and every child mutation must be durable (written AND
+     * fsynced) BEFORE the LMDB transaction that mutates them commits,
+     * because MemoryStore is rebuilt at boot from snapshot + WAL replay,
+     * NOT from LMDB. Relying on the 100ms timer would leave a window where
+     * a crash lands entries in the WAL FILE but not yet fsynced to disk —
+     * indistinguishable, after a real crash, from "never written" — so the
+     * cascade path calls this explicitly instead of waiting on the timer.
+     * See relations/relation_cascade.cpp's writeCascadeWal().
+     */
+    void flushWal();
+
     /**
      * Force an immediate snapshot.
      * Blocks until snapshot is complete.

+ 21 - 4
service/src/persistence/wal.cpp

@@ -8,6 +8,9 @@
 #include <iomanip>
 #include <sstream>
 
+#include <fcntl.h>
+#include <unistd.h>
+
 namespace smartbotic::database {
 
 // ===== CRC32 Implementation =====
@@ -443,10 +446,24 @@ uint64_t WriteAheadLog::append(WalEntry entry) {
 }
 
 void WriteAheadLog::sync() {
-    if (currentFile_.is_open()) {
-        currentFile_.flush();
-        // Note: For true durability, we'd use fsync() here
-        // std::filesystem doesn't provide this, would need OS-specific code
+    if (!currentFile_.is_open()) return;
+    currentFile_.flush();
+    // v2.11.0 T12 — real fsync, not just flush(). flush() only pushes the
+    // C++ stream buffer into the OS page cache; on a crash before the OS
+    // itself writes that page back to disk, a "flushed" WAL entry is
+    // indistinguishable from one never written at all. The cascade
+    // WAL-first design (PersistenceManager::flushWal(), called before the
+    // LMDB commit — see relations/relation_cascade.cpp) depends on
+    // durability being REAL here, not merely buffered: a crash between an
+    // LMDB-only cascade write and an un-fsynced "durable" WAL entry is
+    // exactly the resurrection bug this task exists to close. std::ofstream
+    // has no portable fsync, so this reopens the same path by fd — cheap
+    // (one open/fsync/close), and correct on the POSIX targets this
+    // project ships for.
+    const int fd = ::open(currentFilePath().c_str(), O_WRONLY);
+    if (fd >= 0) {
+        ::fsync(fd);
+        ::close(fd);
     }
 }
 

+ 290 - 0
service/src/relations/relation_cascade.cpp

@@ -0,0 +1,290 @@
+#include "relation_cascade.hpp"
+
+#include <chrono>
+#include <limits>
+#include <unordered_map>
+
+#include <spdlog/spdlog.h>
+
+#include "../memory_store.hpp"
+#include "../persistence/persistence_manager.hpp"
+#include "../project_addressing.hpp"
+#include "../storage/document_store_lmdb.hpp"
+#include "../storage/filter_eval.hpp"
+#include "../storage/lmdb_txn.hpp"
+
+namespace smartbotic::database {
+
+namespace {
+
+uint64_t nowMs() {
+    return static_cast<uint64_t>(
+        std::chrono::duration_cast<std::chrono::milliseconds>(
+            std::chrono::system_clock::now().time_since_epoch())
+            .count());
+}
+
+std::vector<std::string> splitDotPath(const std::string& path) {
+    std::vector<std::string> parts;
+    size_t start = 0;
+    while (start <= path.size()) {
+        const size_t dot = path.find('.', start);
+        parts.push_back(path.substr(start, dot == std::string::npos ? std::string::npos : dot - start));
+        if (dot == std::string::npos) break;
+        start = dot + 1;
+    }
+    return parts;
+}
+
+// Set `root`'s value at `path` to null. Mirrors filter_eval::getJsonPath's
+// read semantics: dotted paths descend OBJECTS only, arrays are never
+// indexed into. No-ops (rather than throws) if the path doesn't resolve —
+// the caller only reaches here after confirming the field held a live
+// reference, but a defensive no-op is cheaper than a second round of
+// validation and strictly safer than crashing mid-cascade over a race.
+void setDotPathNull(nlohmann::json& root, const std::string& path) {
+    const auto parts = splitDotPath(path);
+    if (parts.empty()) return;
+    nlohmann::json* cur = &root;
+    for (size_t i = 0; i + 1 < parts.size(); ++i) {
+        if (!cur->is_object()) return;
+        auto it = cur->find(parts[i]);
+        if (it == cur->end()) return;
+        cur = &(*it);
+    }
+    if (!cur->is_object()) return;
+    (*cur)[parts.back()] = nullptr;
+}
+
+// Remove every string array element equal to `id` at `path`, in place.
+// No-op if the path doesn't resolve to an array.
+void pullDotPathArrayId(nlohmann::json& root, const std::string& path, const std::string& id) {
+    const auto parts = splitDotPath(path);
+    if (parts.empty()) return;
+    nlohmann::json* cur = &root;
+    for (size_t i = 0; i + 1 < parts.size(); ++i) {
+        if (!cur->is_object()) return;
+        auto it = cur->find(parts[i]);
+        if (it == cur->end()) return;
+        cur = &(*it);
+    }
+    if (!cur->is_object()) return;
+    auto it = cur->find(parts.back());
+    if (it == cur->end() || !it->is_array()) return;
+    nlohmann::json pruned = nlohmann::json::array();
+    for (const auto& el : *it) {
+        if (!(el.is_string() && el.get<std::string>() == id)) pruned.push_back(el);
+    }
+    *it = std::move(pruned);
+}
+
+// Document-level wrappers: materialise the full tree, mutate the copy,
+// re-encode. Document::data()/set_data() are the sanctioned way to touch
+// the binary-backed representation from outside doc_binary.{hpp,cpp} — see
+// document.hpp's class comment.
+void setFieldNullInPlace(Document& doc, const std::string& path) {
+    nlohmann::json data = doc.data();
+    setDotPathNull(data, path);
+    doc.set_data(data);
+}
+
+void pullArrayIdInPlace(Document& doc, const std::string& path, const std::string& id) {
+    nlohmann::json data = doc.data();
+    pullDotPathArrayId(data, path, id);
+    doc.set_data(data);
+}
+
+// Running state for one child document while planCascade() folds every
+// relation that touches it. Keyed by (bare child collection, child id) so
+// two different relations from the same parent landing on the same child
+// accumulate onto ONE working copy instead of one clobbering the other's
+// edit — see the header comment on the merge rule.
+struct WorkingChild {
+    std::string collectionQualified;
+    std::string collectionBare;
+    std::string childId;
+    bool deleted = false;
+    Document doc;
+};
+
+} // namespace
+
+CascadePlan planCascade(RelationManager& relations,
+                        smartbotic::db::storage::LmdbDocumentStore& store,
+                        const std::string& qualifiedParentCollection,
+                        const std::string& parentId) {
+    CascadePlan plan;
+
+    try {
+        const auto parentRc = resolveCollection(qualifiedParentCollection);
+        plan.parentHadVector = store.get_vector(parentRc.collection, parentId).has_value();
+    } catch (const std::exception&) {
+        // Malformed parent name would already have failed the delete
+        // upstream (canDelete/the handler's own resolveCollection); treat
+        // as "no vector" rather than let a cosmetic lookup abort planning.
+    }
+
+    std::vector<WorkingChild> working;
+    std::unordered_map<std::string, size_t> indexOf;   // "<bare-coll>\x1f<id>" -> working[]
+
+    for (const auto& r : relations.relationsWithParent(qualifiedParentCollection)) {
+        if (r.onDelete != OnDelete::Cascade && r.onDelete != OnDelete::SetNull) continue;
+
+        std::string bareRelation;
+        smartbotic::database::ResolvedCollection childRc;
+        try {
+            bareRelation = resolveCollection(r.name).collection;
+            childRc = resolveCollection(r.child);
+        } catch (const std::exception& e) {
+            spdlog::error("relations: cascade planning skipped malformed relation '{}': {}",
+                         r.name, e.what());
+            continue;
+        }
+
+        // Unbounded — relation_index_children truncates at `limit`, and a
+        // cascade that silently dropped children past some cap would be
+        // worse than no cascade at all (a dangling reference an operator
+        // can find with `relations check`, versus data quietly left
+        // referencing a deleted parent forever). SIZE_MAX asks for
+        // everything; see the .hpp file header.
+        const auto childIds = store.relation_index_children(
+            bareRelation, parentId, std::numeric_limits<size_t>::max());
+
+        for (const auto& childId : childIds) {
+            const std::string key = childRc.collection + "\x1f" + childId;
+            size_t idx;
+            auto found = indexOf.find(key);
+            if (found == indexOf.end()) {
+                auto childDoc = store.get(childRc.collection, childId);
+                if (!childDoc) continue;   // raced away between index read and here
+                WorkingChild w;
+                w.collectionQualified = r.child;
+                w.collectionBare = childRc.collection;
+                w.childId = childId;
+                w.doc = std::move(*childDoc);
+                working.push_back(std::move(w));
+                idx = working.size() - 1;
+                indexOf.emplace(key, idx);
+            } else {
+                idx = found->second;
+            }
+
+            WorkingChild& w = working[idx];
+            if (w.deleted) continue;   // already slated for deletion — nothing further to do
+
+            auto val = smartbotic::db::storage::filter_eval::resolveFilterValue(w.doc, r.childField);
+            if (!val) continue;   // field no longer present — race, skip this relation's effect
+
+            if (val->is_array()) {
+                // ⚠ THE ARRAY RULE — cascade and set_null collapse. See the
+                // .hpp file header: pull the id, keep the document, no
+                // matter which policy is declared.
+                pullArrayIdInPlace(w.doc, r.childField, parentId);
+            } else if (val->is_string() && val->get<std::string>() == parentId) {
+                if (r.onDelete == OnDelete::Cascade) {
+                    w.deleted = true;
+                    continue;   // don't touch w.doc further — it's being deleted
+                }
+                setFieldNullInPlace(w.doc, r.childField);
+            } else {
+                continue;   // no longer actually references parentId — race, skip
+            }
+            w.doc.version += 1;
+            w.doc.updatedAt = nowMs();
+        }
+    }
+
+    plan.mutations.reserve(working.size());
+    for (auto& w : working) {
+        CascadeMutation mut;
+        mut.childCollectionQualified = w.collectionQualified;
+        mut.childCollectionBare = w.collectionBare;
+        mut.childId = w.childId;
+        if (w.deleted) {
+            mut.kind = CascadeMutation::Kind::DeleteChild;
+            mut.hadVector = store.get_vector(w.collectionBare, w.childId).has_value();
+        } else {
+            mut.kind = CascadeMutation::Kind::UpdateChild;
+            mut.updatedDoc = std::move(w.doc);
+        }
+        plan.mutations.push_back(std::move(mut));
+    }
+
+    return plan;
+}
+
+void writeCascadeWal(PersistenceManager& persistence,
+                     const std::string& qualifiedParentCollection,
+                     const std::string& parentId,
+                     const CascadePlan& plan) {
+    for (const auto& m : plan.mutations) {
+        if (m.kind == CascadeMutation::Kind::DeleteChild) {
+            persistence.logDelete(m.childCollectionQualified, m.childId);
+            if (m.hadVector) persistence.logVecDelete(m.childCollectionQualified, m.childId);
+        } else {
+            persistence.logUpdate(m.childCollectionQualified, *m.updatedDoc);
+        }
+    }
+    persistence.logDelete(qualifiedParentCollection, parentId);
+    if (plan.parentHadVector) persistence.logVecDelete(qualifiedParentCollection, parentId);
+
+    // MANDATORY — see the .hpp file header's WAL-before-LMDB explanation.
+    // Every entry above must be durable before commitCascadeLmdb() runs.
+    persistence.flushWal();
+}
+
+bool commitCascadeLmdb(smartbotic::db::storage::LmdbDocumentStore& store,
+                       const std::string& qualifiedParentCollection,
+                       const std::string& parentId,
+                       const CascadePlan& plan) {
+    const auto parentRc = resolveCollection(qualifiedParentCollection);
+
+    smartbotic::db::storage::WriteTxn wtxn = store.beginWrite();
+    std::vector<std::pair<std::string, unsigned int>> to_cache;
+
+    for (const auto& m : plan.mutations) {
+        if (m.kind == CascadeMutation::Kind::DeleteChild) {
+            store.del(wtxn, m.childCollectionBare, m.childId, to_cache);
+            if (m.hadVector) store.del_vector(wtxn, m.childCollectionBare, m.childId, to_cache);
+        } else {
+            store.put(wtxn, m.childCollectionBare, m.childId, *m.updatedDoc, to_cache);
+        }
+    }
+
+    const bool parentExisted = store.del(wtxn, parentRc.collection, parentId, to_cache);
+    if (plan.parentHadVector) store.del_vector(wtxn, parentRc.collection, parentId, to_cache);
+
+    // ONE commit for the whole cascade — all-or-nothing. commitAndCache()
+    // caches every handle in to_cache only AFTER this commit succeeds.
+    store.commitAndCache(wtxn, to_cache);
+    return parentExisted;
+}
+
+void applyCascadeToMemory(MemoryStore& memStore,
+                          const std::string& qualifiedParentCollection,
+                          const std::string& parentId,
+                          const CascadePlan& plan) {
+    for (const auto& m : plan.mutations) {
+        if (m.kind == CascadeMutation::Kind::DeleteChild) {
+            memStore.unloadDocument(m.childCollectionQualified, m.childId);
+        } else {
+            memStore.loadDocumentWithHistory(m.childCollectionQualified, *m.updatedDoc);
+        }
+    }
+    memStore.unloadDocument(qualifiedParentCollection, parentId);
+}
+
+bool executeCascade(RelationManager& relations,
+                    smartbotic::db::storage::LmdbDocumentStore& store,
+                    PersistenceManager& persistence,
+                    MemoryStore& memStore,
+                    const std::string& qualifiedParentCollection,
+                    const std::string& parentId) {
+    const CascadePlan plan = planCascade(relations, store, qualifiedParentCollection, parentId);
+    writeCascadeWal(persistence, qualifiedParentCollection, parentId, plan);
+    const bool parentExisted = commitCascadeLmdb(store, qualifiedParentCollection, parentId, plan);
+    applyCascadeToMemory(memStore, qualifiedParentCollection, parentId, plan);
+    return parentExisted;
+}
+
+} // namespace smartbotic::database

+ 157 - 0
service/src/relations/relation_cascade.hpp

@@ -0,0 +1,157 @@
+// v2.11.0 T12 — cascade and set_null, made destructive and WAL-first.
+//
+// Today (through T9/Phase A) a relation declaring `cascade` or `set_null`
+// behaves exactly like `no_action`: the delete goes through and the
+// reference is left dangling. This module makes them actually act, while
+// leaving `restrict` (relation_enforcement.hpp) untouched — that check
+// still runs first, upstream of everything here, and still aborts the
+// whole delete with FAILED_PRECONDITION when it blocks.
+//
+// ⚠ THE PLACE A RELATIONAL MENTAL MODEL ACTIVELY MISLEADS: for an
+// ARRAY-VALUED reference, `cascade` and `set_null` COLLAPSE TO THE SAME
+// BEHAVIOUR — pull the parent id out of the array and keep the document.
+// A MySQL-trained instinct says `cascade` means "delete the child row";
+// that is only true for a SCALAR reference. Deleting a node because one of
+// its three credentials went away would be worse than useless. Only a
+// scalar reference under `cascade` deletes the child document. Every array
+// match — under either policy — and every scalar match under `set_null`,
+// mutates the child in place instead.
+//
+// ⚠ WAL-BEFORE-LMDB, AND WHY: MemoryStore is rebuilt at boot from snapshot
+// + WAL replay, NOT from LMDB (LMDB is a read-serving mirror, kept in sync
+// under MemoryStore's per-collection lock on the ordinary write path — see
+// memory_store.cpp's mirrorWriteToDocStore). A cascade that committed its
+// LMDB transaction before the WAL entries for the parent and every child
+// mutation were written and fsynced would, on a crash in that window, come
+// back on restart with the children RESURRECTED — reconstructed from WAL
+// exactly as they were before the cascade, now pointing at a parent that
+// LMDB (correctly) no longer has. Ordering it the other way — WAL first,
+// fsynced, THEN one atomic LMDB commit — means a crash in that window
+// instead leaves LMDB one step behind a WAL that already describes the
+// full cascade; the next boot's WAL replay reconstructs the correct
+// (post-cascade) MemoryStore regardless of whether LMDB got that far, and
+// replaying an already-applied delete/update against MemoryStore is a
+// no-op (MemoryStore::remove()/unloadDocument() on an absent id just
+// returns false). See writeCascadeWal()/commitCascadeLmdb() below for the
+// exact sequence, and tests/test_relation_enforcement.cpp for the crash
+// simulation that pins this.
+//
+// Sequence, in full (mirrors the design doc):
+//   1. Resolve children through the reverse index — planCascade().
+//   2. restrict still blocks upstream (relation_enforcement.hpp) — this
+//      module is never reached if that already refused the delete.
+//   3. WAL-log the parent delete + every child mutation, fsync —
+//      writeCascadeWal().
+//   4. One atomic LMDB WriteTxn: mutate children, update the reverse
+//      index (automatic — see maintainRelations in document_store_lmdb.cpp),
+//      delete the parent, commit — commitCascadeLmdb().
+//   5. Apply the same mutations to MemoryStore, taking each affected
+//      collection's lock in turn — applyCascadeToMemory().
+// executeCascade() runs 3 → 4 → 5 in order; it is what DatabaseGrpcImpl::
+// Delete calls once RelationEnforcer::canDelete() has already permitted
+// the delete.
+//
+// Scope note: this module does NOT queue replication or publish Subscribe
+// events for the child mutations or the parent delete — see the Task 12
+// report for why (the ordinary per-doc write path gets both for free from
+// MemoryStore's persistCallback_, which this module deliberately bypasses
+// to avoid double-logging to WAL; reaching replication/events without that
+// callback needs its own follow-up, out of this task's scope).
+
+#pragma once
+
+#include <optional>
+#include <string>
+#include <vector>
+
+#include "document.hpp"
+#include "relation_manager.hpp"
+
+namespace smartbotic::db::storage {
+class LmdbDocumentStore;
+}
+
+namespace smartbotic::database {
+
+class PersistenceManager;
+class MemoryStore;
+
+// One resolved child-side mutation a cascade/set_null relation requires.
+struct CascadeMutation {
+    enum class Kind { DeleteChild, UpdateChild };
+    Kind kind = Kind::DeleteChild;
+    std::string childCollectionQualified;   // for MemoryStore/WAL (e.g. "default:executions")
+    std::string childCollectionBare;        // for LmdbDocumentStore (e.g. "executions")
+    std::string childId;
+    bool hadVector = false;                 // DeleteChild only — also drop the vector
+    std::optional<Document> updatedDoc;     // UpdateChild only — the full post-mutation doc
+};
+
+struct CascadePlan {
+    std::vector<CascadeMutation> mutations;
+    bool parentHadVector = false;
+};
+
+// Step 1 — read-only. Resolves every Cascade/SetNull relation whose parent
+// is qualifiedParentCollection, walks the reverse index for parentId under
+// each (unbounded — a cascade must see every child, never a truncated
+// sample), reads the current child documents, and decides Delete vs
+// Update per the array rule in the file header. Mutates nothing.
+//
+// If the SAME child is targeted by two different relations from this
+// parent (e.g. two fields on one child collection both referencing it),
+// and the two would disagree (one wants Delete, the other Update), Delete
+// wins — a child slated for deletion is not resurrected by a later Update
+// in the same plan. This is a narrow, deliberately simple merge rule; see
+// the Task 12 report for the reasoning.
+CascadePlan planCascade(RelationManager& relations,
+                        smartbotic::db::storage::LmdbDocumentStore& store,
+                        const std::string& qualifiedParentCollection,
+                        const std::string& parentId);
+
+// Step 3 — WAL-log the parent delete and every mutation in `plan`, then
+// force an fsync (PersistenceManager::flushWal()). MUST run, and MUST
+// complete, before commitCascadeLmdb() — see the file header.
+void writeCascadeWal(PersistenceManager& persistence,
+                     const std::string& qualifiedParentCollection,
+                     const std::string& parentId,
+                     const CascadePlan& plan);
+
+// Step 4 — one atomic LMDB WriteTxn: apply every mutation in `plan`
+// (deleting or updating each child, which also maintains the reverse index
+// and any secondary indexes automatically — see maintainRelations/
+// maintainIndexes), then delete the parent (and its vector, if any).
+// All-or-nothing: any exception leaves neither a child nor the parent
+// mutated — the WriteTxn's destructor aborts unwritten. Returns whether the
+// PARENT document itself was present and removed (independent of how many
+// children were mutated — a cascade run against an already-gone parent id
+// still cleans up any dangling children the reverse index still names, the
+// same repair `relations check` (Task 7) exists to find).
+bool commitCascadeLmdb(smartbotic::db::storage::LmdbDocumentStore& store,
+                       const std::string& qualifiedParentCollection,
+                       const std::string& parentId,
+                       const CascadePlan& plan);
+
+// Step 5 — apply the same mutations to MemoryStore, taking each affected
+// collection's lock IN TURN (one call per document, not one lock spanning
+// the whole cascade). Memory-only, no WAL/mirror/callback side effects —
+// see MemoryStore::unloadDocument()'s header comment for why.
+void applyCascadeToMemory(MemoryStore& store,
+                          const std::string& qualifiedParentCollection,
+                          const std::string& parentId,
+                          const CascadePlan& plan);
+
+// Orchestrates steps 3 → 4 → 5, in order, with nothing in between. What
+// DatabaseGrpcImpl::Delete calls once RelationEnforcer::canDelete() (step
+// 2 — restrict) has already permitted the delete; this function does not
+// re-check restrict and must not be reached if that refused. Returns
+// commitCascadeLmdb()'s result — whether the parent document itself was
+// present and removed — for the handler's DeleteResponse.deleted field.
+bool executeCascade(RelationManager& relations,
+                    smartbotic::db::storage::LmdbDocumentStore& store,
+                    PersistenceManager& persistence,
+                    MemoryStore& memStore,
+                    const std::string& qualifiedParentCollection,
+                    const std::string& parentId);
+
+} // namespace smartbotic::database

+ 13 - 0
service/src/storage/document_store_lmdb.cpp

@@ -398,6 +398,19 @@ LmdbDocumentStore::cachedDbi(std::string_view collection) {
     return it->second;
 }
 
+WriteTxn LmdbDocumentStore::beginWrite() {
+    return WriteTxn(env_);
+}
+
+void LmdbDocumentStore::commitAndCache(
+    WriteTxn& wtxn,
+    const std::vector<std::pair<std::string, unsigned int>>& to_cache) {
+    // Commit BEFORE any handle is cached - the v2.8.0 lesson, same as every
+    // no-txn wrapper in this file.
+    wtxn.commit();
+    for (const auto& [sub, d] : to_cache) cacheCommittedDbi(sub, d);
+}
+
 std::optional<unsigned int>
 LmdbDocumentStore::try_open_for_read(ReadTxn& rtxn,
                                       std::string_view collection) {

+ 28 - 0
service/src/storage/document_store_lmdb.hpp

@@ -380,6 +380,34 @@ public:
                            const float* data,
                            size_t count)> callback) override;
 
+    // v2.11.0 T12 — open a caller-owned WriteTxn spanning several
+    // txn-accepting put()/del()/put_vector()/del_vector() calls across
+    // MULTIPLE collections (parent + every cascade-affected child + the
+    // relation/index sub-dbs those touch) — exactly what an atomic cascade
+    // needs and what the txn-accepting overloads above were built for (see
+    // their header comment). Never opens a second MDB_env: this reuses
+    // env_, the same instance every internal caller already shares.
+    //
+    // The caller owns wtxn's lifetime: call every put()/del() it needs
+    // against ONE shared `to_cache`, then commitAndCache() exactly once.
+    // Letting wtxn go out of scope uncommitted (or calling .abort())
+    // aborts everything done through it — nothing partial can survive.
+    class WriteTxn beginWrite();
+
+    // v2.11.0 T12 — commit `wtxn`, then cache every handle accumulated in
+    // `to_cache`, in that order. Companion to beginWrite(): the same
+    // commit-then-cache discipline every no-txn wrapper in this file
+    // already applies internally (put(), del(), put_vector(), del_vector()),
+    // exposed for a caller that opened its own WriteTxn via beginWrite() and
+    // shared one to_cache across several calls. Never cache before commit -
+    // a handle whose opening transaction later aborts is closed by LMDB, and
+    // caching it anyway is the v2.8.0 bug (EINVAL for the life of the
+    // process). This is the ONLY sanctioned way for outside code to reach
+    // the private per-handle cache — deliberately, so "commit before cache"
+    // cannot be gotten wrong by a caller that forgets the ordering.
+    void commitAndCache(class WriteTxn& wtxn,
+                        const std::vector<std::pair<std::string, unsigned int>>& to_cache);
+
 private:
     // Sub-db handle cache. MDB_dbi is typedef'd to unsigned int — held as
     // that bare type to avoid pulling <lmdb.h> into this header.

+ 16 - 0
tests/CMakeLists.txt

@@ -610,12 +610,15 @@ add_executable(test_relation_enforcement
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/relations/relation_index.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/relations/relation_manager.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/relations/relation_enforcement.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/relations/relation_cascade.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/memory_store.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/views/view_manager.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/views/projection.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/config/collection_config_manager.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/persistence/history_store.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/persistence/wal.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/persistence/snapshot.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/persistence/persistence_manager.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/json_parse.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/doc_binary.cpp
 )
@@ -644,6 +647,19 @@ endif()
 find_package(Threads REQUIRED)
 target_link_libraries(test_relation_enforcement PRIVATE Threads::Threads)
 
+# v2.11.0 T12 — persistence_manager.cpp now linked in (WAL-first cascade
+# tests need a real PersistenceManager); it pulls in snapshot.cpp, which
+# uses LZ4 compression, same as test_migrate_v1_to_v2 above.
+if(PKG_CONFIG_FOUND)
+    pkg_check_modules(RELENF_LZ4 QUIET liblz4)
+endif()
+if(RELENF_LZ4_FOUND)
+    target_link_libraries(test_relation_enforcement PRIVATE ${RELENF_LZ4_LIBRARIES})
+    target_include_directories(test_relation_enforcement PRIVATE ${RELENF_LZ4_INCLUDE_DIRS})
+else()
+    target_link_libraries(test_relation_enforcement PRIVATE lz4)
+endif()
+
 add_test(NAME test_relation_enforcement COMMAND test_relation_enforcement)
 
 # v2.9.0 — secondary index key encoding. Pins the one property that matters:

+ 369 - 5
tests/test_relation_enforcement.cpp

@@ -11,6 +11,7 @@
 // RelationManager + RelationEnforcer + LmdbDocumentStore directly, no
 // grpc::ServerContext. The RPC-level check lands in Task 6/9.
 
+#include <algorithm>
 #include <atomic>
 #include <cstdio>
 #include <filesystem>
@@ -23,6 +24,8 @@
 
 #include "document.hpp"
 #include "memory_store.hpp"
+#include "persistence/persistence_manager.hpp"
+#include "relations/relation_cascade.hpp"
 #include "relations/relation_enforcement.hpp"
 #include "relations/relation_manager.hpp"
 #include "storage/document_store_lmdb.hpp"
@@ -30,16 +33,27 @@
 
 namespace fs = std::filesystem;
 
+using smartbotic::database::CascadeMutation;
+using smartbotic::database::CascadePlan;
 using smartbotic::database::Document;
 using smartbotic::database::MemoryStore;
 using smartbotic::database::OnDelete;
+using smartbotic::database::PersistenceManager;
 using smartbotic::database::RelationBlock;
 using smartbotic::database::RelationEnforcer;
 using smartbotic::database::RelationInfo;
 using smartbotic::database::RelationManager;
 using smartbotic::database::RelationImpact;
+using smartbotic::database::WalEntry;
+using smartbotic::database::WalOpType;
+using smartbotic::database::WriteAheadLog;
+using smartbotic::database::applyCascadeToMemory;
+using smartbotic::database::commitCascadeLmdb;
 using smartbotic::database::describeDeleteImpacts;
+using smartbotic::database::executeCascade;
 using smartbotic::database::findRelationBlocks;
+using smartbotic::database::planCascade;
+using smartbotic::database::writeCascadeWal;
 using smartbotic::db::storage::LmdbDocumentStore;
 using smartbotic::db::storage::LmdbEnv;
 using smartbotic::db::storage::LmdbEnvOpts;
@@ -82,6 +96,66 @@ struct TmpEnv {
     TmpEnv& operator=(const TmpEnv&) = delete;
 };
 
+// v2.11.0 T12 — a PersistenceManager over its own tmpdir, for the WAL-first
+// cascade tests. Separate from TmpEnv (which wraps an LmdbEnv) because the
+// two are deliberately independent stores in this test file, exactly as
+// they are in the real service (executeCascade takes both, wired together
+// only by this test/the Delete handler, never by MemoryStore's own
+// persistCallback_ — see relation_cascade.hpp's file header).
+struct TmpPersistence {
+    std::string path;
+    PersistenceManager::Config cfg;
+    PersistenceManager pm;
+    explicit TmpPersistence(const char* tag)
+        : path(make_tmpdir(tag)), cfg(makeConfig(path)), pm(cfg) {}
+    ~TmpPersistence() {
+        std::error_code ec;
+        fs::remove_all(path, ec);
+    }
+    TmpPersistence(const TmpPersistence&) = delete;
+    TmpPersistence& operator=(const TmpPersistence&) = delete;
+
+private:
+    static PersistenceManager::Config makeConfig(const std::string& dir) {
+        PersistenceManager::Config c;
+        c.dataDir = dir;
+        return c;
+    }
+};
+
+// Replay every WAL entry under `dataDir`/wal directly, bypassing
+// PersistenceManager/MemoryStore entirely — the most direct way to answer
+// "does the WAL file actually contain this entry", independent of whatever
+// LMDB or MemoryStore ended up with.
+std::vector<WalEntry> replayRawWal(const std::string& dataDir) {
+    WriteAheadLog::Config wcfg;
+    wcfg.walDir = std::filesystem::path(dataDir) / "wal";
+    WriteAheadLog wal(wcfg);
+    std::vector<WalEntry> out;
+    wal.replay(0, [&](const WalEntry& e) { out.push_back(e); });
+    return out;
+}
+
+bool walHasDelete(const std::vector<WalEntry>& entries,
+                  const std::string& collection, const std::string& id) {
+    for (const auto& e : entries) {
+        if (e.opType == WalOpType::DELETE && e.collection == collection && e.documentId == id) {
+            return true;
+        }
+    }
+    return false;
+}
+
+bool walHasUpdate(const std::vector<WalEntry>& entries,
+                  const std::string& collection, const std::string& id) {
+    for (const auto& e : entries) {
+        if (e.opType == WalOpType::UPDATE && e.collection == collection && e.documentId == id) {
+            return true;
+        }
+    }
+    return false;
+}
+
 // -------------------------------------------------------------------------
 
 void test_index_follows_the_child_field() {
@@ -225,10 +299,17 @@ void test_restrict_blocks_and_names_the_blockers() {
     mstore.stop();
 }
 
-// Cascade/set_null are NOT yet destructive (Task 12) - a relation declaring
-// either one must behave exactly like no_action here, not block, and not
-// half-implement destruction.
-void test_cascade_and_set_null_permit_for_now() {
+// v2.11.0 T12 — cascade/set_null ARE now destructive (see the tests further
+// below: test_cascade_deletes_scalar_children_wal_first,
+// test_array_reference_pulls_id_and_keeps_document), but that destruction
+// happens entirely OUTSIDE findRelationBlocks()/canDelete() - those two only
+// ever implement `restrict`, deliberately (see relation_enforcement.hpp's
+// file header: "Only restrict blocks"). A relation declaring cascade or
+// set_null must therefore still never BLOCK a delete via this path, exactly
+// like no_action - the actual cascade/set_null execution is a separate step
+// (executeCascade, called by the Delete handler only after this check has
+// already permitted the delete).
+void test_cascade_and_set_null_never_block_via_restrict_check() {
     MemoryStore mstore(MemoryStore::Config{});
     mstore.start();
     RelationManager rm(mstore);
@@ -429,6 +510,286 @@ void test_describe_delete_no_relations_declared() {
     mstore.stop();
 }
 
+// -------------------------------------------------------------------------
+// v2.11.0 T12 — cascade/set_null made destructive, WAL-first.
+// -------------------------------------------------------------------------
+
+// The brief's first required test: cascade delete of a parent with three
+// scalar-referencing children. Children gone, index entries gone, and -
+// the part most likely to be skipped - a WAL entry exists for every child
+// mutation (not just the parent). Would fail if regressed: if
+// commitCascadeLmdb() forgot to delete a child, that child's LMDB row or
+// index posting would survive; if writeCascadeWal() forgot a child, the raw
+// WAL replay would not contain its DELETE entry - this reads the WAL file
+// directly (replayRawWal), not through any code path that could paper over
+// a missing log call.
+void test_cascade_deletes_scalar_children_wal_first() {
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+
+    TmpEnv t("rel-cascade-scalar");
+    LmdbDocumentStore store(t.env);
+    store.set_relations("executions", {{"exec_wf", "workflowId"}});
+
+    TmpPersistence p("rel-cascade-scalar-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    auto putBoth = [&](const std::string& id, const nlohmann::json& data) {
+        Document d; d.id = id; d.collection = "executions"; d.set_data(data);
+        store.put("executions", id, d);
+        mstore.loadDocument("default:executions", d);
+    };
+    putBoth("e1", {{"workflowId", "wf-1"}});
+    putBoth("e2", {{"workflowId", "wf-1"}});
+    putBoth("e3", {{"workflowId", "wf-1"}});
+
+    {
+        Document parent; parent.id = "wf-1"; parent.collection = "workflows";
+        parent.set_data({{"name", "example"}});
+        store.put("workflows", "wf-1", parent);
+        mstore.loadDocument("default:workflows", parent);
+    }
+
+    RelationInfo rel;
+    rel.name = "default:exec_wf";
+    rel.child = "default:executions";
+    rel.childField = "workflowId";
+    rel.parent = "default:workflows";
+    rel.onDelete = OnDelete::Cascade;
+    std::string mgrErr;
+    check(rm.createRelation(rel, mgrErr), "declared the cascade relation");
+
+    check(store.relation_index_child_count("exec_wf", "wf-1") == 3, "three children indexed");
+
+    const bool parentExisted = executeCascade(rm, store, p.pm, mstore, "default:workflows", "wf-1");
+    check(parentExisted, "the parent document was present and removed");
+
+    // LMDB: children gone, index gone.
+    check(!store.get("executions", "e1").has_value(), "e1 gone from LMDB");
+    check(!store.get("executions", "e2").has_value(), "e2 gone from LMDB");
+    check(!store.get("executions", "e3").has_value(), "e3 gone from LMDB");
+    check(!store.get("workflows", "wf-1").has_value(), "parent gone from LMDB");
+    check(store.relation_index_child_count("exec_wf", "wf-1") == 0,
+          "reverse index has no postings left for wf-1");
+
+    // MemoryStore: same, applied via applyCascadeToMemory / unloadDocument.
+    check(!mstore.get("default:executions", "e1").has_value(), "e1 gone from MemoryStore");
+    check(!mstore.get("default:executions", "e2").has_value(), "e2 gone from MemoryStore");
+    check(!mstore.get("default:executions", "e3").has_value(), "e3 gone from MemoryStore");
+    check(!mstore.get("default:workflows", "wf-1").has_value(), "parent gone from MemoryStore");
+
+    // WAL: every child mutation AND the parent got its own DELETE entry -
+    // read straight from the WAL file, not inferred from LMDB/MemoryStore
+    // state (which could be right for the wrong reason if WAL logging were
+    // silently skipped).
+    p.pm.stop();
+    auto entries = replayRawWal(p.path);
+    check(walHasDelete(entries, "default:executions", "e1"), "WAL has DELETE for e1");
+    check(walHasDelete(entries, "default:executions", "e2"), "WAL has DELETE for e2");
+    check(walHasDelete(entries, "default:executions", "e3"), "WAL has DELETE for e3");
+    check(walHasDelete(entries, "default:workflows", "wf-1"), "WAL has DELETE for the parent");
+
+    mstore.stop();
+}
+
+// The brief's second required test, and the one the file header calls out
+// as the place a MySQL mental model actively misleads: an array-valued
+// reference under `cascade` pulls the id and KEEPS the document, exactly
+// like set_null would. Would fail if regressed: a naive port of "cascade
+// deletes the child" would delete the node here, which this test catches
+// directly (node must still exist afterward), not just check the array
+// shrank.
+void test_array_reference_pulls_id_and_keeps_document() {
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+
+    TmpEnv t("rel-cascade-array");
+    LmdbDocumentStore store(t.env);
+    store.set_relations("nodes", {{"node_creds", "config.credentialIds"}});
+
+    TmpPersistence p("rel-cascade-array-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    {
+        Document n; n.id = "n1"; n.collection = "nodes";
+        n.set_data({{"config", {{"credentialIds", {"c1", "c2", "c3"}}}}});
+        store.put("nodes", "n1", n);
+        mstore.loadDocument("default:nodes", n);
+    }
+    {
+        Document parent; parent.id = "c1"; parent.collection = "credentials";
+        parent.set_data({{"name", "prod-key"}});
+        store.put("credentials", "c1", parent);
+        mstore.loadDocument("default:credentials", parent);
+    }
+
+    // Cascade, deliberately - this is exactly the policy a MySQL-trained
+    // instinct expects to delete the node. It must not.
+    RelationInfo rel;
+    rel.name = "default:node_creds";
+    rel.child = "default:nodes";
+    rel.childField = "config.credentialIds";
+    rel.parent = "default:credentials";
+    rel.onDelete = OnDelete::Cascade;
+    std::string mgrErr;
+    check(rm.createRelation(rel, mgrErr), "declared the cascade relation on an array field");
+
+    const bool parentExisted = executeCascade(rm, store, p.pm, mstore, "default:credentials", "c1");
+    check(parentExisted, "the credential document was present and removed");
+
+    // The node survives, in BOTH stores, with c1 pulled and the other two ids intact.
+    auto lmdbNode = store.get("nodes", "n1");
+    check(lmdbNode.has_value(), "node n1 still exists in LMDB - not deleted");
+    if (lmdbNode) {
+        auto ids = lmdbNode->data()["config"]["credentialIds"];
+        check(ids.is_array() && ids.size() == 2, "two ids remain");
+        check(std::find(ids.begin(), ids.end(), nlohmann::json("c1")) == ids.end(),
+              "c1 was pulled");
+        check(std::find(ids.begin(), ids.end(), nlohmann::json("c2")) != ids.end(),
+              "c2 survives");
+        check(std::find(ids.begin(), ids.end(), nlohmann::json("c3")) != ids.end(),
+              "c3 survives");
+    }
+
+    auto memNode = mstore.get("default:nodes", "n1");
+    check(memNode.has_value(), "node n1 still exists in MemoryStore - not deleted");
+    if (memNode) {
+        auto ids = memNode->data()["config"]["credentialIds"];
+        check(ids.is_array() && ids.size() == 2, "MemoryStore copy agrees: two ids remain");
+    }
+
+    check(store.relation_index_child_count("node_creds", "c1") == 0,
+          "the pulled posting is gone from the reverse index");
+    check(store.relation_index_child_count("node_creds", "c2") == 1,
+          "c2's posting is untouched");
+
+    check(!store.get("credentials", "c1").has_value(), "the credential itself IS deleted");
+
+    p.pm.stop();
+    auto entries = replayRawWal(p.path);
+    check(walHasUpdate(entries, "default:nodes", "n1"),
+          "WAL has an UPDATE for the node (not a DELETE) - the array rule");
+    check(walHasDelete(entries, "default:credentials", "c1"), "WAL has DELETE for the parent");
+
+    mstore.stop();
+}
+
+// The brief's third required test: simulate a crash between the WAL write
+// (step 3) and the LMDB commit (step 4) by calling exactly those primitives
+// in that order and stopping there - never calling commitCascadeLmdb() or
+// applyCascadeToMemory(). This is what a real crash right after
+// PersistenceManager::flushWal() returns looks like from the next boot's
+// point of view: the WAL file has everything, LMDB has nothing.
+//
+// "Restart" = a FRESH MemoryStore, recovered via a second PersistenceManager
+// instance over the same dataDir (recover() replays from sequence 0 since
+// no snapshot exists - the fresh-install path in persistence_manager.cpp).
+// Recovery must converge on the correct post-cascade state even though the
+// LMDB transaction that would have made it "real" never ran - proving
+// MemoryStore's correctness after a crash depends on the WAL, not on how
+// far the LMDB commit got. Then recovering a SECOND time onto the SAME
+// (already-recovered) store checks that replaying an already-applied
+// delete is a no-op - it must not throw, and the state must not change.
+//
+// ⚠ Scope of what this test can and cannot prove: it exercises the WAL
+// primitives in isolation, so it proves the WAL entries are self-sufficient
+// for correct recovery. It does NOT, by itself, prove that executeCascade()
+// calls writeCascadeWal() before commitCascadeLmdb() - that ordering is
+// pinned by executeCascade() being three plain sequential statements next
+// to this comment in relation_cascade.cpp, not by a runtime fault
+// injection. See the Task 12 report for why a real fault-injection test
+// was judged not worth the complexity here.
+void test_crash_between_wal_and_commit_recovers() {
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+
+    TmpEnv t("rel-cascade-crash");
+    LmdbDocumentStore store(t.env);
+    store.set_relations("executions", {{"exec_wf", "workflowId"}});
+
+    TmpPersistence p("rel-cascade-crash-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    auto putBoth = [&](const std::string& id, const nlohmann::json& data) {
+        Document d; d.id = id; d.collection = "executions"; d.set_data(data);
+        store.put("executions", id, d);
+        mstore.loadDocument("default:executions", d);
+    };
+    putBoth("e1", {{"workflowId", "wf-1"}});
+    putBoth("e2", {{"workflowId", "wf-1"}});
+
+    {
+        Document parent; parent.id = "wf-1"; parent.collection = "workflows";
+        parent.set_data({{"name", "example"}});
+        store.put("workflows", "wf-1", parent);
+        mstore.loadDocument("default:workflows", parent);
+    }
+
+    RelationInfo rel;
+    rel.name = "default:exec_wf";
+    rel.child = "default:executions";
+    rel.childField = "workflowId";
+    rel.parent = "default:workflows";
+    rel.onDelete = OnDelete::Cascade;
+    std::string mgrErr;
+    check(rm.createRelation(rel, mgrErr), "declared the cascade relation");
+
+    // Steps 1+3 only - simulating a crash immediately after flushWal()
+    // returns, before commitCascadeLmdb()/applyCascadeToMemory() ever run.
+    CascadePlan plan = planCascade(rm, store, "default:workflows", "wf-1");
+    check(plan.mutations.size() == 2, "planned both children");
+    writeCascadeWal(p.pm, "default:workflows", "wf-1", plan);
+    p.pm.stop();   // closes the WAL file, like a process exiting
+
+    // Proof the "crash" really happened: LMDB was never touched by this
+    // cascade attempt - the children and parent are still exactly as they
+    // were.
+    check(store.get("executions", "e1").has_value(),
+          "LMDB was never committed - e1 is still there (this is the point)");
+    check(store.get("executions", "e2").has_value(),
+          "LMDB was never committed - e2 is still there");
+    check(store.get("workflows", "wf-1").has_value(),
+          "LMDB was never committed - the parent is still there");
+
+    // "Restart": fresh MemoryStore, fresh PersistenceManager over the SAME
+    // dataDir, recover().
+    MemoryStore freshStore(MemoryStore::Config{});
+    freshStore.start();
+    PersistenceManager::Config cfg2;
+    cfg2.dataDir = p.path;
+    PersistenceManager pm2(cfg2);
+    auto outcome = pm2.recover(freshStore);
+    check(outcome.kind != smartbotic::database::RecoveryOutcome::Kind::Failed,
+          "recovery did not fail");
+    check(outcome.walEntriesReplayed >= 3,
+          "replayed at least the parent delete + two child deletes");
+
+    check(!freshStore.get("default:executions", "e1").has_value(),
+          "recovery converges: e1 is gone from the recovered MemoryStore");
+    check(!freshStore.get("default:executions", "e2").has_value(),
+          "recovery converges: e2 is gone from the recovered MemoryStore");
+    check(!freshStore.get("default:workflows", "wf-1").has_value(),
+          "recovery converges: the parent is gone from the recovered MemoryStore");
+
+    // Replaying an already-applied delete is a no-op: recover() a SECOND
+    // time, onto the SAME already-recovered store. Must not throw and must
+    // not change the outcome.
+    auto outcome2 = pm2.recover(freshStore);
+    check(outcome2.kind != smartbotic::database::RecoveryOutcome::Kind::Failed,
+          "second recovery pass did not fail either");
+    check(!freshStore.get("default:executions", "e1").has_value(),
+          "still gone after replaying the same DELETE a second time - idempotent");
+
+    freshStore.stop();
+    mstore.stop();
+}
+
 }  // namespace
 
 int main() {
@@ -438,11 +799,14 @@ int main() {
     test_unrelated_update_leaves_the_posting_alone();
     test_posting_visible_after_reopen();
     test_restrict_blocks_and_names_the_blockers();
-    test_cascade_and_set_null_permit_for_now();
+    test_cascade_and_set_null_never_block_via_restrict_check();
     test_relations_enforced_false_skips_the_check();
     test_describe_delete_reports_restrict_and_no_action();
     test_describe_delete_zero_children_does_not_block();
     test_describe_delete_no_relations_declared();
+    test_cascade_deletes_scalar_children_wal_first();
+    test_array_reference_pulls_id_and_keeps_document();
+    test_crash_between_wal_and_commit_recovers();
 
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;