소스 검색

fix(relations): T13 review round 2 - gate on relationsEnforced, close two holes

Four fixes from code review, plus one additional test the coordinator asked
for:

1. validate_on_write now reads relationsEnforced and is skipped when it is
   false - the documented escape hatch (bulk import, write pressure) did not
   previously stop it, so the only way to get past a rejection was dropping
   and re-declaring the relation. RelationRef gains a relationsEnforced
   field, snapshotted at both arming sites (armRelationsForChild,
   applyRelationDeclarations) from CollectionConfigManager, since the
   storage layer stays free of that dependency the same way it stays free
   of RelationManager. ConfigureCollection re-arms the collection's
   relations when it flips the flag, so the snapshot doesn't go stale.
   Fixed the now-false "no arming step" comment on
   CollectionCfg::relationsEnforced.

2. Reworded the race-test's block comment in test_relation_enforcement.cpp:
   it claimed to prove the race is "actually closed, not merely narrowed,"
   but its sequence (committed get, committed del, then put) would reject
   identically under a narrowed window too. Placement inside the
   transaction is established by reading the code, not by this test - the
   comment now says so.

3. RelationManager::loadFromStore did not re-check the same-project rule
   createRelation enforces, so a hand-written or legacy cross-project
   _relations record would arm with its parent resolved against the wrong
   project's env. Now skipped and logged, like applyRelationDeclarations
   already does for unparseable records.

4. Guarded backfillIntoDocStore's per-row applyDualWriteMirror call, which
   was uncaught and would have failed initialize() over one rejected row
   (UniqueViolation had the same hole since T11; fixed for both).

Added test_validate_on_write_rolls_back_an_already_applied_sibling_relation:
proves a real LMDB abort undoes an already-applied index mutation from an
earlier relation in the same write, not just "the rejecting relation's
sub-db never existed" (which the round-1 assertions could not distinguish).
Also added the relationsEnforced escape-hatch test and a RelationManager
cross-project-skip test.

Verified each new/changed assertion by reverting and rerunning: gate
disabled -> 3 dependent failures; full guard disabled -> 18 dependent
failures (171 = 189 - 18); cross-project check reverted -> 1 dependent
failure. All restored and green after.
fszontagh 1 개월 전
부모
커밋
9575041044

+ 13 - 5
service/src/config/collection_config_manager.hpp

@@ -62,11 +62,19 @@ struct CollectionCfg {
     //
     // Consumed by Task 4's delete enforcement (restrict/no_action on delete):
     // read via configFor(collection).relationsEnforced before refusing a
-    // delete that would orphan a child row. Nothing else applies this at
-    // boot yet - there is no separate "arming" step, because enforcement is
-    // just a per-call read of this flag, not state that needs to be pushed
-    // into another component the way indexedFields is pushed into
-    // LmdbDocumentStore::set_indexed_fields().
+    // delete that would orphan a child row - that consumer IS just a
+    // per-call read of this flag, no arming needed.
+    //
+    // v2.11.0 T13 round 2 — validate_on_write is different: it runs inside
+    // LmdbDocumentStore, which cannot read CollectionConfigManager itself
+    // (keyed by qualified name; the per-project storage layer deliberately
+    // doesn't carry that - see RelationRef's comment). So this flag IS
+    // pushed into another component for that consumer, the same way
+    // indexedFields is pushed via LmdbDocumentStore::set_indexed_fields():
+    // ConfigureCollection re-arms the collection's relations
+    // (armRelationsForChild) whenever it changes this flag, so
+    // RelationRef::relationsEnforced stays a fresh snapshot rather than
+    // going stale.
     bool relationsEnforced = true;
 
     // v2.11.0 T8 — persisted only. Nothing enforces uniqueness yet (Task 11

+ 18 - 1
service/src/database_grpc_impl.cpp

@@ -2862,6 +2862,13 @@ void DatabaseGrpcImpl::armRelationsForChild(const std::string& childQualified) {
         auto* lmdb = dynamic_cast<smartbotic::db::storage::LmdbDocumentStore*>(ds);
         if (lmdb == nullptr) return;  // no LMDB substrate for this project; nothing to arm
 
+        // v2.11.0 T13 round 2 — snapshot relationsEnforced for this child at
+        // arm time, since the storage layer cannot read CollectionConfigManager
+        // itself (see RelationRef's comment). ConfigureCollection re-arms
+        // after flipping the flag, so this only ever goes stale between that
+        // RPC's write and its own re-arm call - never observably.
+        const bool enforced = config_manager_.configFor(childQualified).relationsEnforced;
+
         std::vector<smartbotic::db::storage::RelationRef> refs;
         for (const auto& rel : relation_manager_.relationsWithChild(childQualified)) {
             const auto rn = smartbotic::database::resolveCollection(rel.name);
@@ -2871,7 +2878,7 @@ void DatabaseGrpcImpl::armRelationsForChild(const std::string& childQualified) {
             // parent is guaranteed to live in this same project's env.
             const auto pc = smartbotic::database::resolveCollection(rel.parent);
             refs.push_back(smartbotic::db::storage::RelationRef{
-                rn.collection, rel.childField, pc.collection, rel.validateOnWrite});
+                rn.collection, rel.childField, pc.collection, rel.validateOnWrite, enforced});
         }
         lmdb->set_relations(rc.collection, refs);
         spdlog::info("v2.11 relations: re-armed {} relation(s) for child '{}'",
@@ -3709,6 +3716,16 @@ grpc::Status DatabaseGrpcImpl::ConfigureCollection(
         response->set_success(ok);
         if (!ok) {
             response->set_error(err);
+        } else if (request->config().has_relations_enforced()) {
+            // v2.11.0 T13 round 2 — RelationRef::relationsEnforced is a
+            // snapshot taken at arm time (the storage layer cannot read
+            // CollectionConfigManager itself - see RelationRef's comment), so
+            // flipping the flag here must re-arm this collection's relations
+            // or validate_on_write keeps consulting the OLD value until
+            // something else happens to re-declare a relation. A no-op
+            // (empty refs) when this collection is not a child in any
+            // relation - armRelationsForChild handles that already.
+            armRelationsForChild(request->collection());
         }
         return grpc::Status::OK;
     } catch (const std::exception& e) {

+ 26 - 4
service/src/database_service.cpp

@@ -408,8 +408,13 @@ void DatabaseService::applyRelationDeclarations() {
             // as armRelationsForChild's mirror of this construction.
             const auto pc = resolveCollection(r.parent);
             const std::string key = rc.project + ":" + rc.collection;
+            // v2.11.0 T13 round 2 — snapshot relationsEnforced at boot-time
+            // arming, same reasoning as armRelationsForChild's mirror of
+            // this construction. `key` is already the qualified child name
+            // configFor expects.
+            const bool enforced = config_manager_->configFor(key).relationsEnforced;
             byChild[key].push_back(smartbotic::db::storage::RelationRef{
-                rn.collection, r.childField, pc.collection, r.validateOnWrite});
+                rn.collection, r.childField, pc.collection, r.validateOnWrite, enforced});
             keyToProjectCollection[key] = {rc.project, rc.collection};
         } catch (const std::exception& e) {
             // Advisory per relation: one unparseable declaration must not
@@ -552,9 +557,26 @@ bool DatabaseService::backfillIntoDocStore() {
         for (const auto& doc : docs) {
             std::optional<Document> opt_doc(doc);
             const uint64_t drift_before = mirror_drift_count_.load(std::memory_order_relaxed);
-            smartbotic::db::storage::applyDualWriteMirror(
-                ds, mirror_healthy_, mirror_drift_count_,
-                pc.collection, doc.id, opt_doc, EventType::INSERT);
+            // v2.11.0 T13 round 2 — UniqueViolation/MissingParentReference are
+            // rethrown UNCAUGHT by applyDualWriteMirror (deliberately - see
+            // both exceptions' header comments), unlike a genuine mirror
+            // fault which is swallowed and only shows up as a drift bump
+            // below. Uncaught here would escape this loop, backfillIntoDocStore
+            // and initialize() itself, refusing startup over one rejected row.
+            // Practically unreachable today - backfill only runs migrating a
+            // v1.x dataset, which predates both `_relations` and unique-index
+            // declarations - but catching removes that reasoning burden for
+            // the next reader, same as any other row here that fails.
+            try {
+                smartbotic::db::storage::applyDualWriteMirror(
+                    ds, mirror_healthy_, mirror_drift_count_,
+                    pc.collection, doc.id, opt_doc, EventType::INSERT);
+            } catch (const std::exception& e) {
+                spdlog::error("v2.3 backfill: mirror rejected {}/{}: {} - row stays "
+                              "in MemoryStore only, LMDB does not have it",
+                              qualified, doc.id, e.what());
+                ++total_failures;
+            }
             if (mirror_drift_count_.load(std::memory_order_relaxed) > drift_before) {
                 ++total_failures;
             }

+ 33 - 0
service/src/relations/relation_manager.cpp

@@ -86,6 +86,39 @@ void RelationManager::loadFromStore() {
             if (d.id.empty()) continue;
             RelationInfo r = fromJson(d.data());
             if (r.name.empty()) continue;
+
+            // v2.11.0 T13 round 2 — createRelation refuses a cross-project
+            // declaration at write time (see below), but this load loop is
+            // the OTHER entry point into the cache and did not re-check it.
+            // A legacy or hand-written `_relations` document naming a
+            // cross-project parent would arm a bare parent collection name
+            // that then gets resolved inside the CHILD's own project env
+            // (armRelationsForChild/applyRelationDeclarations only ever
+            // resolve `r.parent` against the child's project) - so
+            // validate_on_write would check an unrelated, wrong collection
+            // for existence, silently accepting or rejecting for the wrong
+            // reason. Skip and log rather than fail the whole load: one bad
+            // record must not stop every other relation from arming, same
+            // reasoning as applyRelationDeclarations' per-relation try/catch.
+            try {
+                const auto rn = resolveCollection(r.name);
+                const auto rc = resolveCollection(r.child);
+                const auto rp = resolveCollection(r.parent);
+                if (rc.project != rp.project || rc.project != rn.project) {
+                    spdlog::error(
+                        "RelationManager: skipping relation '{}' - name/child/parent "
+                        "span more than one project ({}/{}/{}), which createRelation() "
+                        "refuses today; this record predates that check or was written "
+                        "by hand",
+                        r.name, rn.project, rc.project, rp.project);
+                    continue;
+                }
+            } catch (const std::exception& e) {
+                spdlog::error("RelationManager: skipping unparseable relation '{}': {}",
+                              r.name, e.what());
+                continue;
+            }
+
             cache_[r.name] = std::move(r);
         }
         if (res.documents.size() < kPage) break;

+ 11 - 1
service/src/storage/document_store_lmdb.cpp

@@ -832,7 +832,17 @@ void LmdbDocumentStore::maintainRelations(
         // Absent/null references never reach here: extract_relation_ids /
         // resolveFilterValue never put them in `new_ids`, so they can never
         // land in `to_add`.
-        if (r.validateOnWrite && !to_add.empty()) {
+        //
+        // r.relationsEnforced gates this exactly like it gates restrict/
+        // no_action in RelationEnforcer::canDelete: the documented escape
+        // hatch for a bulk import or a collection under write pressure is
+        // "enforcement off means every relation policy is skipped, not just
+        // restrict" - a validate_on_write rejection during a bulk import
+        // whose parents are not loaded yet is precisely the case the switch
+        // exists for. Before this, the only way to stop the rejection was to
+        // drop and re-declare the relation without validateOnWrite, which is
+        // not what the switch is for.
+        if (r.validateOnWrite && r.relationsEnforced && !to_add.empty()) {
             // open_for_write, not try_open_for_read: this call happens
             // inside our own write transaction (the only kind that may call
             // mdb_dbi_open - see try_open_for_read's file comment), and

+ 11 - 0
service/src/storage/document_store_lmdb.hpp

@@ -89,11 +89,22 @@ public:
 // `validateOnWrite` mirrors RelationInfo::validateOnWrite. Both stay
 // optional-by-default (empty / false) so every pre-T13 aggregate-init call
 // site (`RelationRef{name, childField}`) keeps compiling unchanged.
+// v2.11.0 T13 round 2 — `relationsEnforced` mirrors CollectionCfg::
+// relationsEnforced (config/collection_config_manager.hpp), snapshotted at
+// arming time (armRelationsForChild / applyRelationDeclarations) rather than
+// read live: the storage layer stays free of CollectionConfigManager the
+// same way it stays free of RelationManager (see the file comment above) -
+// that dependency is keyed by the QUALIFIED collection name, which this
+// per-project layer deliberately does not carry. Staying current after a
+// configureCollection flip is the caller's job: ConfigureCollection
+// re-arms the collection's relations after changing the flag, exactly like
+// any other change to a relation's declared shape.
 struct RelationRef {
     std::string name;
     std::string childField;
     std::string parent;
     bool validateOnWrite = false;
+    bool relationsEnforced = true;
 };
 
 class LmdbDocumentStore : public DocumentStore {

+ 148 - 24
tests/test_relation_enforcement.cpp

@@ -1440,33 +1440,155 @@ void test_validate_on_write_unrelated_update_not_rechecked() {
           "the update itself still applied");
 }
 
-// v2.11.0 T13 — proves the race is actually closed, not merely narrowed.
+// v2.11.0 T13 round 2 (review finding) — rollback of an ALREADY-APPLIED
+// index mutation within the same write, not merely "the mutation never
+// happened because validation ran first."
 //
-// Models the exact interleaving the plan describes: something (an
-// application-level pre-check, or the old restrict path's own read) observes
-// the parent present, and only AFTER that does the parent get deleted -
-// before the child's write actually lands. A stale check-then-act sequence
-// would let the child insert through anyway, because its answer was decided
-// against the state as of the check, not as of the write.
+// The earlier "left no posting behind" assertions are structurally weak on
+// their own: validate_on_write's check runs before that RELATION's own
+// index sub-db is even opened, so on rejection the sub-db often never
+// exists and relation_index_child_count returns 0 through the nullopt path
+// regardless of whether LMDB actually rolled anything back. This test
+// forces a real rollback to matter: `executions` declares TWO relations.
+// The first (wf_rel, validateOnWrite=false) has its posting WRITTEN - a
+// real mdb_put against a real sub-db, inside the loop, before the second
+// relation is even considered. The second (owner_rel, validateOnWrite=true)
+// then rejects. Both relations share the one write transaction `put()`
+// opens, so the abort must undo wf_rel's already-applied mdb_put along with
+// everything else - there is no sub-db-never-existed shortcut available
+// here, because it demonstrably did exist and did get written to.
+void test_validate_on_write_rolls_back_an_already_applied_sibling_relation() {
+    TmpEnv t("validate-rollback-sibling");
+    LmdbDocumentStore store(t.env);
+    store.set_relations("executions", {
+        {"wf_rel", "workflowId", "workflows", false},   // validateOnWrite=false
+        {"owner_rel", "ownerId", "users", true},         // validateOnWrite=true
+    });
+
+    Document w; w.id = "wf-1"; w.collection = "workflows";
+    w.set_data({{"name", "real workflow"}});
+    store.put("workflows", "wf-1", w);
+    // Deliberately no "users/u-ghost" - owner_rel's parent never exists.
+
+    Document e; e.id = "e1"; e.collection = "executions";
+    e.set_data({{"workflowId", "wf-1"}, {"ownerId", "u-ghost"}});
+    bool threw = false;
+    try {
+        store.put("executions", "e1", e);
+    } catch (const MissingParentReference&) {
+        threw = true;
+    }
+    check(threw, "owner_rel's missing parent rejects the write");
+    check(!store.get("executions", "e1").has_value(),
+          "the document itself was rolled back");
+    check(store.relation_index_child_count("owner_rel", "u-ghost") == 0,
+          "owner_rel (the relation that rejected) has no posting");
+    check(store.relation_index_child_count("wf_rel", "wf-1") == 0,
+          "wf_rel (the EARLIER, already-applied sibling relation) was rolled "
+          "back too - LMDB's transaction abort undid a real mdb_put, not "
+          "just 'the sub-db never got created'");
+
+    // Once owner_rel's parent exists, the identical write succeeds and BOTH
+    // relations end up with their postings - confirming the rollback above
+    // was real and not a side effect of some other bug losing wf_rel's
+    // posting permanently.
+    Document u; u.id = "u-ghost"; u.collection = "users";
+    u.set_data({{"name", "real user, now created"}});
+    store.put("users", "u-ghost", u);
+    threw = false;
+    try {
+        store.put("executions", "e1", e);
+    } catch (const MissingParentReference&) {
+        threw = true;
+    }
+    check(!threw, "once owner_rel's parent exists too, the write succeeds");
+    check(store.relation_index_child_count("wf_rel", "wf-1") == 1, "wf_rel posted");
+    check(store.relation_index_child_count("owner_rel", "u-ghost") == 1, "owner_rel posted");
+}
+
+// v2.11.0 T13 round 2 (review finding 1) — relationsEnforced is the
+// documented escape hatch (config/collection_config_manager.hpp) for "a
+// bulk import, or a collection under write pressure." Before this, the
+// only consumer was RelationEnforcer::canDelete (restrict/no_action on
+// delete); validate_on_write did not read it at all, so the only way to
+// stop a validate_on_write rejection was to drop and re-declare the
+// relation without validateOnWrite - not what the switch is for. Disabling
+// enforcement must also let a bulk-import-shaped write through even though
+// its parent is not loaded yet.
+void test_validate_on_write_relations_enforced_false_is_the_escape_hatch() {
+    TmpEnv t("validate-enforced-off");
+    LmdbDocumentStore store(t.env);
+    // relationsEnforced=false alongside validateOnWrite=true - the exact
+    // combination an operator reaches for mid-bulk-import.
+    store.set_relations("executions",
+                        {{"exec_wf", "workflowId", "workflows", true, false}});
+
+    Document d; d.id = "e1"; d.collection = "executions";
+    d.set_data({{"workflowId", "wf-ghost"}});
+    bool threw = false;
+    try {
+        store.put("executions", "e1", d);
+    } catch (const MissingParentReference&) {
+        threw = true;
+    }
+    check(!threw, "relationsEnforced=false lets a validateOnWrite=true write through");
+    check(store.get("executions", "e1").has_value(),
+          "the escape-hatch write actually landed");
+
+    // Re-enabling enforcement does not retroactively touch what was already
+    // written (documented behaviour, mirrors canDelete's own log message) -
+    // but a NEW write with a missing parent is rejected again.
+    store.set_relations("executions",
+                        {{"exec_wf", "workflowId", "workflows", true, true}});
+    Document d2; d2.id = "e2"; d2.collection = "executions";
+    d2.set_data({{"workflowId", "wf-ghost-2"}});
+    threw = false;
+    try {
+        store.put("executions", "e2", d2);
+    } catch (const MissingParentReference&) {
+        threw = true;
+    }
+    check(threw, "re-enabling enforcement rejects a new write with a missing parent");
+    check(store.get("executions", "e1").has_value(),
+          "the earlier escape-hatch write was not retroactively undone");
+}
+
+// v2.11.0 T13 — demonstrates validation answers against write-time state,
+// not a stale earlier observation.
+//
+// Models one honest slice of the interleaving the plan describes: something
+// (an application-level pre-check, or the old restrict path's own read)
+// observes the parent present, and only AFTER that does the parent get
+// deleted - in its own committed transaction - before the child's write
+// happens. A check that trusted the earlier observation would let the
+// child insert through anyway.
 //
-// LMDB is single-writer (see try_open_for_read's file comment and
-// document_store_lmdb.cpp's env setup): every write transaction begins only
-// after the previous one has fully committed, so the child's write
-// transaction here necessarily starts strictly after the parent-delete
-// transaction commits. Because validate_on_write's mdb_get runs INSIDE the
-// child's own write transaction rather than in a separate, earlier read, it
-// sees the parent's true state as of the write, not as of whatever was
-// observed before. That is the whole mechanism this task adds: no separate
-// transaction, no interval, nothing that can go stale.
+// What this test does NOT establish: it does NOT distinguish "the mdb_get
+// runs inside the child's own write transaction" (the actual fix - see
+// maintainRelations, before the child's own mdb_put, same transaction the
+// caller commits) from "the mdb_get runs in a separate read transaction
+// opened immediately before the child's write transaction" (a narrowed
+// window, not a closed one). This test's sequence - get, then a
+// committed del, then put - would reject identically either way, because
+// the del is fully committed before either kind of check would run. That
+// placement is inside the transaction, not merely adjacent to it, is
+// established by reading the code (document_store_lmdb.cpp: the mdb_get is
+// at maintainRelations, called from put() before put()'s own mdb_put,
+// against the same wtxn the caller commits), not by this test.
 //
-// What this test does NOT establish: it does not exercise real multi-thread
-// scheduling or prove there is no OTHER race at the LMDB layer. It does not
-// need to - LMDB's single-writer guarantee means transaction ORDER is the
-// only thing that can vary under concurrency, never interleaving within a
-// transaction, so serialising the two operations in program order is the
-// honest, deterministic equivalent of "the delete's transaction commits
-// before the child insert's transaction begins," which is the only
-// interleaving the race actually depends on.
+// What this test DOES show: the validation's answer tracks the parent's
+// state as of when the check actually runs, not whatever an earlier,
+// separate read happened to observe - which is the necessary condition for
+// the fix to work at all, even though it is not sufficient to prove
+// placement by itself. It is deliberately single-threaded and
+// deterministic, not a real multi-thread stress test: LMDB is single-writer
+// (see try_open_for_read's file comment and document_store_lmdb.cpp's env
+// setup), so under real concurrency the only thing that can vary is
+// transaction ORDER, never interleaving within a transaction - serialising
+// "the delete's transaction commits, then the child write's transaction
+// begins" in program order is the deterministic equivalent of that
+// ordering, which is as much of the race as a single-process test can
+// exercise.
 void test_validate_on_write_closes_the_stale_check_race() {
     TmpEnv t("validate-race");
     LmdbDocumentStore store(t.env);
@@ -1534,6 +1656,8 @@ int main() {
     test_validate_on_write_never_rejects_absent_or_null();
     test_validate_on_write_array_any_missing_rejects_whole_write();
     test_validate_on_write_unrelated_update_not_rechecked();
+    test_validate_on_write_rolls_back_an_already_applied_sibling_relation();
+    test_validate_on_write_relations_enforced_false_is_the_escape_hatch();
     test_validate_on_write_closes_the_stale_check_race();
 
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";

+ 55 - 0
tests/test_relation_manager.cpp

@@ -13,6 +13,7 @@
 
 #include <nlohmann/json.hpp>
 
+#include "document.hpp"
 #include "memory_store.hpp"
 #include "relations/relation_manager.hpp"
 
@@ -98,6 +99,59 @@ void test_more_than_one_page_of_relations_loads() {
           "policies and collection configs");
 }
 
+// v2.11.0 T13 round 2 (review finding 3) — createRelation() refuses a
+// cross-project declaration (test_cross_project_relation_is_refused,
+// above), but loadFromStore() is the OTHER way a RelationInfo enters the
+// cache and did not re-check it. A legacy or hand-written `_relations`
+// document naming a cross-project parent must be skipped at load time too -
+// arming it would resolve the bare parent name inside the CHILD's own
+// project env (validate_on_write/RelationRef only ever resolve `parent`
+// against the child's project), silently checking the wrong collection.
+void test_load_skips_a_cross_project_relation_written_by_hand() {
+    Fixture f;
+
+    // Bypass createRelation()'s own guard entirely - write the raw document
+    // straight into `_relations`, the way a legacy record or a hand-edited
+    // one would exist on disk. `parent` names a DIFFERENT project than
+    // `child`, which createRelation() would refuse today.
+    nlohmann::json bad = {
+        {"name", "default:bad_cross"},
+        {"child", "default:executions"},
+        {"child_field", "workflowId"},
+        {"parent", "acme:workflows"},
+        {"on_delete", "restrict"},
+        {"validate_on_write", true},
+        {"created_at", 0},
+        {"updated_at", 0},
+    };
+    Document d;
+    d.id = "default:bad_cross";
+    d.collection = RelationManager::SYSTEM_COLLECTION;
+    d.set_data(bad);
+    f.store.insert(RelationManager::SYSTEM_COLLECTION, d);
+
+    // A well-formed, same-project relation alongside it, to confirm one bad
+    // record does not stop the rest of the load.
+    RelationInfo good;
+    good.name = "default:exec_wf";
+    good.child = "default:executions";
+    good.childField = "ownerId";
+    good.parent = "default:users";
+    {
+        RelationManager rm(f.store);
+        rm.loadFromStore();
+        std::string err;
+        check(rm.createRelation(good, err), "the well-formed sibling declares fine");
+    }
+
+    RelationManager fresh(f.store);
+    fresh.loadFromStore();
+    check(!fresh.getRelation("default:bad_cross").has_value(),
+          "the cross-project relation was skipped, not armed with the wrong parent");
+    check(fresh.getRelation("default:exec_wf").has_value(),
+          "the well-formed sibling still loaded - one bad record did not stop the rest");
+}
+
 }  // namespace
 
 int main() {
@@ -105,6 +159,7 @@ int main() {
     test_relations_are_project_scoped_and_survive_reload();
     test_cross_project_relation_is_refused();
     test_more_than_one_page_of_relations_loads();
+    test_load_skips_a_cross_project_relation_written_by_hand();
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;
 }