Prechádzať zdrojové kódy

fix(relations): T12 review round 3 - I1 LMDB convergence gap, plus 2 Minors

Addresses the last open Critical/Important finding and two requested Minors:

- I1: after a crash between writeCascadeWal()'s fsync and
  commitCascadeLmdb()'s commit, a replayed UPDATE/UPSERT (a set_null or
  array-pull mutation) never reached LMDB on the next boot, because WAL
  replay's INSERT/UPDATE path (loadDocument()/loadDocumentWithHistory())
  never mirrors - unlike DELETE replay, which converges via
  MemoryStore::remove()'s own mirror call. Traced the callers of
  logInsert/logUpdate/logDelete (exactly two: the ordinary persist
  callback, which always mirrors before logging, and relation_cascade.cpp,
  which is the only caller that logs before mirroring) to confirm this bug
  is currently only reachable through cascade, even though the underlying
  non-mirroring mechanism is pre-existing, ordinary code. Fixed with the
  narrower of the two options the review offered: PersistenceManager::
  recover() now tracks every distinct (collection, id) applied as
  UPDATE/UPSERT during WAL replay and re-mirrors each to LMDB exactly once
  via a new MemoryStore::remirrorDocument(), after replay completes -
  bounded by distinct ids touched since the last snapshot, not by WAL
  entry count, and generalized to any replayed UPDATE (not cascade-
  special-cased, since WAL entries carry no such marker and the
  generalization is a free, idempotent no-op for the ordinary case).
  Confirmed by reverting the fix and observing 3 assertions fail, then
  restoring it.
- Corrected an inaccurate claim from round 2: relation_cascade.hpp and
  database_grpc_impl.cpp said a post-WAL-commit failure means "neither
  store reflects it yet" - true only if commitCascadeLmdb() threw; if
  applyCascadeToMemory() threw instead, LMDB already committed and is
  already visible via LMDB-first reads. Both now state both cases.
- Documented (not fixed, per instruction) a schema shape where a
  restrict-blocking grandchild that is itself one of the plan's own
  children makes the parent permanently undeletable, even though applying
  the whole plan atomically would leave nothing dangling.

New test: test_replayed_cascade_update_remirrors_to_lmdb_after_crash_window,
driving the real cascade path through the actual crash window and
asserting LMDB convergence, not just MemoryStore's.
test_relation_enforcement: 116 -> 131 assertions.
fszontagh 1 mesiac pred
rodič
commit
67c1cc7104

+ 50 - 23
service/src/database_grpc_impl.cpp

@@ -883,33 +883,60 @@ grpc::Status DatabaseGrpcImpl::Delete(
                                 // itself uses.
                                 return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, e.what());
                             } catch (const std::exception& e) {
-                                // v2.11.0 T12 review (I4) — WAL-first means a
-                                // failure here can occur AFTER the WAL entries
-                                // for this cascade were already durably
-                                // fsynced (writeCascadeWal() ran inside
-                                // executeCascade() before commitCascadeLmdb()/
-                                // applyCascadeToMemory(), either of which could
-                                // be what actually threw). This is NOT
-                                // necessarily a no-op failure: the mutation may
-                                // already be committed to the WAL and WILL be
-                                // applied on the next restart's replay even
-                                // though this call reports failure now, and
-                                // neither store reflects it yet in the
-                                // meantime. There is no compensating "un-write"
-                                // of the WAL entry — replaying it again is
-                                // safe, erasing it is not. Check server logs
-                                // and, if in doubt, the current document state
-                                // before retrying.
+                                // v2.11.0 T12 review (I4, round 2 correction)
+                                // — WAL-first means a failure here can occur
+                                // AFTER the WAL entries for this cascade were
+                                // already durably fsynced (writeCascadeWal()
+                                // ran inside executeCascade() before
+                                // commitCascadeLmdb()/applyCascadeToMemory(),
+                                // either of which could be what actually
+                                // threw). This is NOT necessarily a no-op
+                                // failure. What the caller can and cannot
+                                // assume differs by WHICH of those two threw,
+                                // and neither is distinguishable from out
+                                // here:
+                                //   - commitCascadeLmdb() threw: LMDB was
+                                //     never committed (the WriteTxn aborts
+                                //     unwritten), so LMDB still reflects the
+                                //     pre-cascade state; MemoryStore is
+                                //     unchanged too (step 5 never ran). Reads
+                                //     see the pre-cascade state everywhere,
+                                //     for now — but the WAL entries are
+                                //     already durable, so the NEXT RESTART's
+                                //     replay applies the whole cascade
+                                //     regardless of this failure.
+                                //   - applyCascadeToMemory() threw partway:
+                                //     LMDB already committed (step 4
+                                //     succeeded) — reads being LMDB-first,
+                                //     the cascade IS ALREADY VISIBLE for
+                                //     collections read through LMDB, even
+                                //     though this call is reporting failure.
+                                //     MemoryStore may be only PARTIALLY
+                                //     applied (some mutations done, some
+                                //     not), and `notify` may have already
+                                //     fired for the mutations that did
+                                //     apply.
+                                // There is no compensating "un-write" of the
+                                // WAL entry in either case — replaying it
+                                // again is safe (idempotent), erasing it is
+                                // not. Check server logs and current
+                                // document state (LMDB, not just MemoryStore)
+                                // before retrying or assuming nothing
+                                // happened.
                                 spdlog::error(
                                     "relations: cascade delete of '{}/{}' failed - the WAL entries for "
-                                    "this cascade may already be durable and will apply on next "
-                                    "restart even though this call is reporting failure: {}",
+                                    "this cascade may already be durable (will apply on next restart "
+                                    "regardless) and, if the failure was in the post-LMDB-commit step, "
+                                    "the mutation may ALREADY be visible via LMDB-first reads even "
+                                    "though this call is reporting failure: {}",
                                     request->collection(), request->id(), e.what());
                                 return grpc::Status(grpc::StatusCode::INTERNAL,
-                                    "cascade delete failed after some or all of its WAL entries may "
-                                    "already be durable - it may still apply on the next restart even "
-                                    "though this call failed; check server logs and current state "
-                                    "before retrying: " + std::string(e.what()));
+                                    "cascade delete failed - its WAL entries may already be durable "
+                                    "(will apply on the next restart regardless of this failure), and "
+                                    "if the failure occurred after the LMDB commit, the mutation may "
+                                    "ALREADY be visible via LMDB-first reads even though this call "
+                                    "failed; check server logs and current document state (not just "
+                                    "this response) before retrying: " + std::string(e.what()));
                             }
                             // executeCascade already deleted the parent (and its
                             // vector, if any) atomically with every child

+ 21 - 0
service/src/memory_store.cpp

@@ -872,6 +872,27 @@ bool MemoryStore::unloadDocument(const std::string& collection, const std::strin
     return true;
 }
 
+bool MemoryStore::remirrorDocument(const std::string& collection, const std::string& id) {
+    // v2.11.0 T12 review (I1) — see the header comment for why this exists.
+    const CollectionData* coll = getCollection(collection);
+    if (!coll) return false;
+
+    std::optional<Document> doc;
+    {
+        std::shared_lock<std::shared_mutex> lock(coll->mutex);
+        auto it = coll->documents.find(id);
+        if (it == coll->documents.end()) return false;
+        doc = it->second;
+    }
+
+    // No CollectionData mutation here — just reading the current state and
+    // pushing it to LMDB. mirrorWriteToDocStore() is itself a no-op if the
+    // resolver was never wired (legacy/test bootstrap), matching every
+    // other caller's documented behaviour.
+    mirrorWriteToDocStore(collection, id, doc, EventType::UPDATE);
+    return true;
+}
+
 bool MemoryStore::exists(const std::string& collection, const std::string& id) const {
     const CollectionData* coll = getCollection(collection);
     if (!coll) {

+ 42 - 0
service/src/memory_store.hpp

@@ -289,6 +289,48 @@ public:
      */
     bool unloadDocument(const std::string& collection, const std::string& id);
 
+    /**
+     * v2.11.0 T12 review (I1) — explicitly re-mirror an id's CURRENT
+     * in-memory state into the LMDB DocumentStore, without touching
+     * MemoryStore itself or the WAL.
+     *
+     * WHY THIS EXISTS: loadDocument()/loadDocumentWithHistory() — what WAL
+     * replay uses to reconstruct MemoryStore for INSERT/UPDATE/UPSERT
+     * entries at boot — do NOT mirror to LMDB (unlike every ordinary write
+     * path, where mirrorDocOrUndo()/mirrorWriteToDocStore() always run
+     * BEFORE the WAL entry is even logged — see update()/insertWithVector()
+     * above). That asymmetry is harmless for the ordinary write path,
+     * because a WAL entry there can only exist if its LMDB write already
+     * committed successfully first (mirror failure prevents the WAL log
+     * call from ever being reached). It stops being harmless the moment
+     * something logs a WAL entry BEFORE its LMDB write — which
+     * relations/relation_cascade.cpp's cascade delete deliberately does
+     * (WAL-first, for its own crash-safety reasons: see that file's
+     * header). A crash between that WAL write and the LMDB commit leaves a
+     * WAL UPDATE entry (a set_null or array-pull mutation) whose LMDB
+     * write never actually happened; DELETE entries self-heal because
+     * remove() DOES mirror, but UPDATE/UPSERT entries replayed via
+     * loadDocumentWithHistory() do not, and LMDB then permanently serves
+     * the pre-cascade row until something else happens to rewrite it.
+     *
+     * PersistenceManager::recover() calls this once per DISTINCT id that
+     * WAL replay applied as UPDATE/UPSERT (deduplicated, and skipped if a
+     * later DELETE for that id was also replayed — remove() already
+     * mirrored that) — a targeted post-replay pass, not a mirror call on
+     * every replayed entry, which would multiply LMDB writes by however
+     * many times a hot document was rewritten since the last snapshot for
+     * no benefit (the ordinary case's LMDB write already happened at
+     * original write time and needs no repeating).
+     *
+     * Safe to call whether or not this specific id actually needed it:
+     * re-mirroring an already-correct row is an idempotent overwrite with
+     * the same content. Returns false (no-op) if the id is not currently
+     * in MemoryStore (nothing to mirror) or the mirror isn't wired
+     * (legacy/test bootstrap path, same as mirrorWriteToDocStore's own
+     * no-op condition).
+     */
+    bool remirrorDocument(const std::string& collection, const std::string& id);
+
     /**
      * Check if a document exists.
      */

+ 36 - 1
service/src/persistence/persistence_manager.cpp

@@ -4,6 +4,8 @@
 
 #include <optional>
 #include <stdexcept>
+#include <unordered_map>
+#include <utility>
 
 namespace smartbotic::database {
 
@@ -203,11 +205,44 @@ RecoveryOutcome PersistenceManager::recover(MemoryStore& store) {
             outcome.failureReason = "Failed to open WAL for replay";
             return false;
         }
-        uint64_t replayed = wal_->replay(fromSequence, [&store](const WalEntry& entry) {
+        // v2.11.0 T12 review (I1) — track every (collection, id) applied as
+        // UPDATE/UPSERT during this replay, deduplicated, so it can be
+        // explicitly re-mirrored to LMDB once replay finishes. Erased again
+        // on a later DELETE for the same id: remove() (DELETE's own replay
+        // path) already mirrors, so there is nothing left to re-mirror, and
+        // re-mirroring a doc that replay just deleted from MemoryStore would
+        // be wrong (remirrorDocument() would correctly no-op on a missing
+        // id, but skipping it here avoids the pointless lookup). See
+        // MemoryStore::remirrorDocument()'s doc comment for why this step
+        // exists at all: WAL replay's INSERT/UPDATE path never mirrors on
+        // its own, unlike every ordinary write, and
+        // relations/relation_cascade.cpp's WAL-first cascade can leave a
+        // replayed UPDATE whose LMDB write never actually happened.
+        std::unordered_map<std::string, std::pair<std::string, std::string>> touchedByUpdate;
+        uint64_t replayed = wal_->replay(fromSequence, [&](const WalEntry& entry) {
             applyWalEntry(store, entry);
+            const std::string key = entry.collection + "\x1f" + entry.documentId;
+            if (entry.opType == WalOpType::UPDATE || entry.opType == WalOpType::UPSERT) {
+                touchedByUpdate[key] = {entry.collection, entry.documentId};
+            } else if (entry.opType == WalOpType::DELETE) {
+                touchedByUpdate.erase(key);
+            }
         });
         outcome.walEntriesReplayed = replayed;
         spdlog::info("Replayed {} WAL entries from sequence {}", replayed, fromSequence);
+
+        uint64_t remirrored = 0;
+        for (const auto& [key, collAndId] : touchedByUpdate) {
+            (void)key;
+            if (store.remirrorDocument(collAndId.first, collAndId.second)) ++remirrored;
+        }
+        outcome.updatesRemirroredAfterReplay = remirrored;
+        if (remirrored > 0) {
+            spdlog::info("Re-mirrored {} document(s) to LMDB after WAL replay "
+                        "(closes the window where a WAL-first writer, e.g. a "
+                        "cascade delete, logged an UPDATE whose LMDB write "
+                        "never ran before a crash)", remirrored);
+        }
         // Anchor the WAL's sequence_ counter to the snapshot's walSequence.
         // Without this, the steady-state outcome of `truncateBefore(walSeq)`
         // after every snapshot — which deletes every WAL file because all

+ 14 - 0
service/src/persistence/persistence_manager.hpp

@@ -48,6 +48,20 @@ struct RecoveryOutcome {
     size_t snapshotsAttempted = 0;
     size_t snapshotsAvailable = 0;
 
+    // v2.11.0 T12 review (I1) — how many distinct (collection, id) pairs
+    // WAL replay applied as UPDATE/UPSERT and then explicitly re-mirrored
+    // to LMDB afterward (MemoryStore::remirrorDocument(), called once per
+    // id from PersistenceManager::recover() — see that method and
+    // remirrorDocument()'s doc comment for why this step exists: WAL
+    // replay's INSERT/UPDATE path does not mirror on its own, unlike every
+    // ordinary write, and relations/relation_cascade.cpp's WAL-first
+    // cascade can leave a replayed UPDATE whose LMDB write never actually
+    // ran). Exposed for tests and operator visibility; not itself an
+    // error signal — a nonzero count on a box that has never run a
+    // cascade just means ordinary UPDATEs got a redundant (harmless,
+    // idempotent) re-mirror.
+    uint64_t updatesRemirroredAfterReplay = 0;
+
     bool isNonTrivial() const {
         return kind != Kind::TrivialSuccess && kind != Kind::FreshInstall;
     }

+ 44 - 12
service/src/relations/relation_cascade.hpp

@@ -87,6 +87,28 @@
 // NOT gated by the grandchild collection's own relationsEnforced — being
 // more cautious than strictly required here is deliberate.
 //
+// ⚠ MINOR, DOCUMENTED NOT FIXED: a schema shape this makes permanently
+// undeletable. If the restrict-blocking grandchild is ITSELF one of the
+// children this same plan would delete — e.g. workflows --cascade-->
+// logs, and logs --restrict--> executions, where executions ALSO cascades
+// from workflows and references a log row — findRelationBlocks() still
+// sees that log row's posting (the check runs against the CURRENT reverse
+// index, not against "what this same plan is also about to remove"), so
+// the whole cascade refuses even though, if the plan were applied as one
+// atomic unit, both the log row and its blocker would vanish together and
+// nothing would actually be left dangling. This is conservative and safe
+// (a false refusal, never a false permit), but an operator who designs a
+// schema shaped like this will find the parent refuses to delete no
+// matter what, with no cascade/set_null policy able to route around it —
+// only dropping or loosening the grandchild relation, or deleting the
+// grandchild row out of band first, resolves it. Fixing this properly
+// means checking each blocked-by posting against the REST OF THE PLAN
+// (not just the live index) before refusing, which needs the plan to be
+// built as a fixed point over multiple passes rather than the current
+// single pass over relationsWithParent() — judged not worth doing for
+// what is likely a rare schema shape, but worth an operator knowing why
+// their delete refuses.
+//
 // Sequence, in full (mirrors the design doc):
 //   1. Resolve children through the reverse index — planCascade(). Also
 //      where the restrict/grandchild check above runs, and where
@@ -106,20 +128,30 @@
 // Delete calls once RelationEnforcer::canDelete() has already permitted
 // the delete.
 //
-// ⚠ FAILURE AFTER THE WAL COMMIT (review finding I4): steps 3-5 are not
-// itself transactional as a WHOLE — only step 4 (the LMDB txn) is
-// all-or-nothing. If commitCascadeLmdb() or applyCascadeToMemory() throws
-// AFTER writeCascadeWal() has already fsynced, the mutation is already
-// DURABLE (the next boot's WAL replay will apply it) even though the
-// caller receives an error and neither store reflects it yet. This is a
+// ⚠ FAILURE AFTER THE WAL COMMIT (review finding I4, corrected in round 2):
+// steps 3-5 are not themselves transactional as a WHOLE — only step 4 (the
+// LMDB txn) is all-or-nothing. If commitCascadeLmdb() or
+// applyCascadeToMemory() throws AFTER writeCascadeWal() has already
+// fsynced, the mutation is already DURABLE (the next boot's WAL replay
+// will apply it) even though the caller receives an error. What "neither
+// store reflects it yet" would get wrong: that is only true if
+// commitCascadeLmdb() is what threw (LMDB's WriteTxn aborts unwritten, so
+// LMDB is genuinely still pre-cascade, and step 5 never ran either). If
+// applyCascadeToMemory() is what threw instead, LMDB has ALREADY committed
+// (step 4 succeeded) — and reads are LMDB-first, so the cascade IS ALREADY
+// VISIBLE through ordinary reads at that point, even though the caller is
+// being told the call failed. MemoryStore in that case may be only
+// PARTIALLY applied (some mutations done, some not — applyCascadeToMemory()
+// applies one document at a time, not atomically), and `notify` may have
+// already fired for whichever mutations did apply before the throw.
+// DatabaseGrpcImpl::Delete's error text distinguishes these two cases by
+// stating both possibilities rather than asserting the wrong one. This is a
 // direct, unavoidable consequence of WAL-first: making the WAL entry
 // contingent on the LMDB commit succeeding would reopen exactly the
-// resurrection window this module exists to close. DatabaseGrpcImpl::Delete
-// says so explicitly in the error text it returns rather than leaving the
-// caller to assume nothing happened; there is no compensating "un-write"
-// of the WAL entry, because replaying it again is safe (idempotent) but
-// erasing it is not (a second failure between erasure and re-fsync would
-// then silently lose the cascade for real).
+// resurrection window this module exists to close. There is no
+// compensating "un-write" of the WAL entry, because replaying it again is
+// safe (idempotent) but erasing it is not (a second failure between
+// erasure and re-fsync would then silently lose the cascade for real).
 //
 // Replication + Subscribe events (review finding C2): the ordinary
 // per-document write path gets both for free because

+ 125 - 0
tests/test_relation_enforcement.cpp

@@ -959,6 +959,130 @@ void test_cascade_refuses_when_grandchild_is_restrict_protected() {
     mstore.stop();
 }
 
+// v2.11.0 T12 review round 3 (I1) - a crash between writeCascadeWal()'s
+// fsync and commitCascadeLmdb() leaves an UpdateChild mutation's WAL UPDATE
+// entry durable while its LMDB write never ran. Unlike DeleteChild (which
+// self-heals on replay because DELETE replay goes through
+// MemoryStore::remove(), which mirrors), UPDATE/UPSERT replay goes through
+// loadDocumentWithHistory(), which never mirrored - so before this fix,
+// LMDB would permanently keep serving the pre-cascade array/field. This
+// test drives the REAL cascade path (planCascade + writeCascadeWal against
+// an array reference, so the mutation really is an UpdateChild - not the
+// isolated primitive) through exactly that crash window, "restarts" with
+// the LMDB mirror wired the same way DatabaseService wires it in production
+// (setupComponents() wires the mirror BEFORE persistence_->recover() runs -
+// confirmed by reading database_service.cpp's initialize()), and asserts
+// LMDB converges, not just MemoryStore.
+void test_replayed_cascade_update_remirrors_to_lmdb_after_crash_window() {
+    MemoryStore mstore(MemoryStore::Config{});
+    mstore.start();
+    RelationManager rm(mstore);
+    rm.loadFromStore();
+    CollectionConfigManager cfgManager(mstore);
+    cfgManager.loadFromStore();
+
+    TmpEnv t("rel-remirror");
+    LmdbDocumentStore store(t.env);
+    store.set_relations("nodes", {{"node_creds", "config.credentialIds"}});
+
+    TmpPersistence p("rel-remirror-wal");
+    check(p.pm.start(), "persistence manager started");
+
+    // Seed through the REAL WAL (logInsert), same discipline as the C1 fix -
+    // these assertions must depend on genuine WAL replay, not just on
+    // store.put()/whatever MemoryStore starts with.
+    Document node; node.id = "n1"; node.collection = "nodes";
+    node.set_data({{"config", {{"credentialIds", {"c1", "c2", "c3"}}}}});
+    store.put("nodes", "n1", node);
+    p.pm.logInsert("default:nodes", node);
+
+    Document parent; parent.id = "c1"; parent.collection = "credentials";
+    parent.set_data({{"name", "prod-key"}});
+    store.put("credentials", "c1", parent);
+    p.pm.logInsert("default:credentials", parent);
+
+    RelationInfo rel;
+    rel.name = "default:node_creds";
+    rel.child = "default:nodes";
+    rel.childField = "config.credentialIds";
+    rel.parent = "default:credentials";
+    rel.onDelete = OnDelete::Cascade;   // array reference -> UpdateChild (pull), not DeleteChild
+    std::string mgrErr;
+    check(rm.createRelation(rel, mgrErr), "declared the cascade relation on an array field");
+
+    // Steps 1+3 only - simulating a crash immediately after flushWal()
+    // returns, before commitCascadeLmdb()/applyCascadeToMemory() ever run.
+    CascadePlan plan = planCascade(rm, store, cfgManager, "default:credentials", "c1");
+    check(plan.mutations.size() == 1, "planned the one array-pull mutation");
+    check(plan.mutations[0].kind == CascadeMutation::Kind::UpdateChild,
+          "confirmed UpdateChild - this is exactly the case DeleteChild does NOT cover");
+    writeCascadeWal(p.pm, "default:credentials", "c1", plan);
+    p.pm.stop();   // closes the WAL file, like a process exiting
+
+    // Proof the "crash" really happened: LMDB was never touched by this
+    // cascade attempt - the node still has all three ids.
+    auto beforeRecovery = store.get("nodes", "n1");
+    check(beforeRecovery.has_value(), "n1 still in LMDB");
+    if (beforeRecovery) {
+        auto ids = beforeRecovery->data()["config"]["credentialIds"];
+        check(ids.is_array() && ids.size() == 3,
+              "LMDB was never committed - still the PRE-cascade array (this is the point)");
+    }
+
+    // "Restart": fresh MemoryStore with the LMDB mirror wired to the SAME
+    // LmdbDocumentStore (matching production ordering - the mirror is wired
+    // in DatabaseService::setupComponents(), which runs BEFORE
+    // persistence_->recover() in initialize()), fresh PersistenceManager
+    // over the same dataDir, recover().
+    MemoryStore freshStore(MemoryStore::Config{});
+    freshStore.start();
+    std::atomic<bool> mirrorHealthy{true};
+    std::atomic<uint64_t> mirrorDrift{0};
+    freshStore.setDocumentStoreMirror(
+        [&store](std::string_view) -> smartbotic::db::storage::DocumentStore* { return &store; },
+        &mirrorHealthy, &mirrorDrift);
+
+    PersistenceManager::Config cfg2;
+    cfg2.dataDir = p.path;
+    PersistenceManager pm2(cfg2);
+    auto outcome = pm2.recover(freshStore);
+    check(outcome.kind != smartbotic::database::RecoveryOutcome::Kind::Failed,
+          "recovery did not fail");
+    check(outcome.updatesRemirroredAfterReplay >= 1,
+          "recover() re-mirrored at least the node's replayed UPDATE");
+
+    // MemoryStore converges (this part already worked before this fix).
+    auto memNode = freshStore.get("default:nodes", "n1");
+    check(memNode.has_value(), "n1 present in the recovered MemoryStore");
+    if (memNode) {
+        auto ids = memNode->data()["config"]["credentialIds"];
+        check(ids.is_array() && ids.size() == 2, "MemoryStore has the post-cascade array");
+    }
+    check(!freshStore.get("default:credentials", "c1").has_value(),
+          "parent gone from the recovered MemoryStore (DELETE replay already worked pre-fix)");
+
+    // LOAD-BEARING: LMDB now ALSO reflects the post-cascade array - only
+    // true because recover()'s remirror pass ran. Without the I1 fix this
+    // would still show 3 ids, exactly like `beforeRecovery` above.
+    auto lmdbAfter = store.get("nodes", "n1");
+    check(lmdbAfter.has_value(), "n1 still in LMDB after recovery");
+    if (lmdbAfter) {
+        auto ids = lmdbAfter->data()["config"]["credentialIds"];
+        check(ids.is_array() && ids.size() == 2,
+              "LMDB converged to the post-cascade array - the I1 fix");
+        check(std::find(ids.begin(), ids.end(), nlohmann::json("c1")) == ids.end(),
+              "c1 is gone from LMDB too, not just MemoryStore");
+    }
+    // The parent delete self-healed on its own even before this fix
+    // (DELETE replay mirrors via MemoryStore::remove()) - confirmed here so
+    // the test pins the WHOLE combined scenario, not just the new half.
+    check(!store.get("credentials", "c1").has_value(),
+          "parent also gone from LMDB after recovery");
+
+    freshStore.stop();
+    mstore.stop();
+}
+
 }  // namespace
 
 int main() {
@@ -977,6 +1101,7 @@ int main() {
     test_array_reference_pulls_id_and_keeps_document();
     test_crash_between_wal_and_commit_recovers();
     test_cascade_refuses_when_grandchild_is_restrict_protected();
+    test_replayed_cascade_update_remirrors_to_lmdb_after_crash_window();
 
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;