Pārlūkot izejas kodu

feat(relations): T9 - end-to-end over the real RPC boundary

Extends the relations e2e coverage with everything only a live service
boundary can prove:

- tests/load_test/test_relations.{cpp,sh} (new) - a real restart, proving
  a declaration survives AND a child written after the restart is indexed
  (the write path is armed, not merely remembered); a manufactured
  boot-time self-heal (the relidx sub-db is dropped offline via a new
  standalone tool while the service is stopped, then restarted, then
  enforcement is asserted to work again on the ORIGINAL pre-drop child,
  proving the rebuild scanned existing rows); and per-project isolation
  of two identically-named relations' children (the v2.4.2 createView
  bug's shape).
- tests/load_test/relidx_drop_tool.cpp (new) - offline LMDB sub-db
  dropper, modeled on subdb_placement.cpp's EnvHandle, CLI-only /
  service-stopped per the documented POSIX-lock hazard.
- tests/load_test/test_relations_client_e2e.cpp - DescribeDelete over the
  boundary with project qualification (Part 3b).
- tests/load_test/test_policy_enforcement.cpp - Phase 11, admin gating on
  CreateRelation/DropRelation/ListRelations/GetRelationInfo under a
  secured project, the four RPCs T6b left untested.

Refusal text asserted in the new tests is read from
relation_enforcement.cpp's formatRelationBlockError, not transcribed from
a report (a prior report mis-transcribed "child document(s)" as
"document(s)").

test_relations_client_e2e.sh 33/33, test_relations.sh 29/29 across two
restarts, test_policy_enforcement.sh 34/34, test_relation_enforcement
45/45, test_relation_index 43/43, test_relation_manager 9/9,
test_subdb_identity 236/236. No defects found.
fszontagh 1 mēnesi atpakaļ
vecāks
revīzija
fffc1bfe66

+ 93 - 0
tests/load_test/relidx_drop_tool.cpp

@@ -0,0 +1,93 @@
+// v2.11.0 T9 — standalone LMDB sub-db removal, used ONLY to manufacture the
+// "relation declared but its reverse index sub-db is missing" state that
+// DatabaseService::applyRelationDeclarations()'s boot-time self-heal (T7)
+// exists to recover from. That self-heal path had no test anywhere in the
+// suite (see task-9-brief.md) because producing the precondition needs a
+// service-level fixture: the sub-db has to actually be gone, not simulated.
+//
+// This is the exact "audit(path, ...) is CLI-only, separate process" pattern
+// already established by storage/subdb_placement.cpp — it opens its own
+// MDB_env directly on the data directory and must ONLY run while the real
+// service process has that path closed. Running it against a live service's
+// env would violate the documented POSIX-lock hazard (closing any fd on a
+// file drops every lock the process holds on it) and corrupt the running
+// service's LMDB state. test_relations.sh enforces this by only invoking the
+// tool between stop_server and the next start_server.
+//
+// Usage: relidx_drop_tool <env_dir> <subdb_name>
+//   Drops the named sub-db (mdb_drop with del=1, which removes both its
+//   contents and its entry in the env's name table) and exits 0. Exits 1 if
+//   the sub-db does not exist or on any LMDB error.
+
+#include <lmdb.h>
+
+#include <cstdint>
+#include <iostream>
+#include <string>
+
+int main(int argc, char** argv) {
+    if (argc != 3) {
+        std::cerr << "usage: relidx_drop_tool <env_dir> <subdb_name>\n";
+        return 2;
+    }
+    const std::string envDir = argv[1];
+    const std::string subdb = argv[2];
+
+    MDB_env* env = nullptr;
+    if (mdb_env_create(&env) != MDB_SUCCESS) {
+        std::cerr << "mdb_env_create failed\n";
+        return 1;
+    }
+    // Same offline-tool ceiling as subdb_placement.cpp's EnvHandle.
+    mdb_env_set_maxdbs(env, 512);
+    mdb_env_set_mapsize(env, 2ULL << 30);
+
+    int rc = mdb_env_open(env, envDir.c_str(), 0, 0664);
+    if (rc != MDB_SUCCESS) {
+        std::cerr << "mdb_env_open('" << envDir << "') failed: " << mdb_strerror(rc)
+                  << " — this tool must run only while the service holding "
+                     "this env is stopped\n";
+        mdb_env_close(env);
+        return 1;
+    }
+
+    MDB_txn* txn = nullptr;
+    rc = mdb_txn_begin(env, nullptr, 0, &txn);
+    if (rc != MDB_SUCCESS) {
+        std::cerr << "mdb_txn_begin failed: " << mdb_strerror(rc) << "\n";
+        mdb_env_close(env);
+        return 1;
+    }
+
+    MDB_dbi dbi = 0;
+    rc = mdb_dbi_open(txn, subdb.c_str(), 0, &dbi);
+    if (rc != MDB_SUCCESS) {
+        std::cerr << "mdb_dbi_open('" << subdb << "') failed: " << mdb_strerror(rc)
+                  << " (sub-db already absent?)\n";
+        mdb_txn_abort(txn);
+        mdb_env_close(env);
+        return 1;
+    }
+
+    // del=1 removes the sub-database's contents AND its handle/name-table
+    // entry — the exact state applyRelationDeclarations()'s self-heal must
+    // recover from (relation_index_exists() returns false).
+    rc = mdb_drop(txn, dbi, 1);
+    if (rc != MDB_SUCCESS) {
+        std::cerr << "mdb_drop('" << subdb << "') failed: " << mdb_strerror(rc) << "\n";
+        mdb_txn_abort(txn);
+        mdb_env_close(env);
+        return 1;
+    }
+
+    rc = mdb_txn_commit(txn);
+    if (rc != MDB_SUCCESS) {
+        std::cerr << "mdb_txn_commit failed: " << mdb_strerror(rc) << "\n";
+        mdb_env_close(env);
+        return 1;
+    }
+
+    mdb_env_close(env);
+    std::cout << "dropped sub-db '" << subdb << "' in " << envDir << "\n";
+    return 0;
+}

+ 41 - 0
tests/load_test/test_policy_enforcement.cpp

@@ -159,6 +159,47 @@ int main(int argc, char** argv) {
     ck(!adminMissing.success && adminMissing.error.find("does not exist") != std::string::npos,
        "admin still gets a real 'does not exist' for a bogus name");
 
+    // ---- Phase 11: v2.11.0 T9 — admin gating on the other four
+    // relation-management RPCs. T6b (CreateRelation/DropRelation/
+    // ListRelations/GetRelationInfo) shipped `requireAnyAdmin()` gates but
+    // left no e2e proving a non-admin is actually refused under a secured
+    // project - the exact miss class the v2.7.0 audit caught for
+    // Upsert/Batch*/Subscribe (coverage established by walking every
+    // handler, not from a list). `docs_rel` (created by ops in Phase 10)
+    // is reused here as something real to probe: each check proves BOTH
+    // that the stranger is refused AND that the refusal is really the
+    // admin gate (ops can still do it), the same shape as Phase 10.
+    ck(!stranger.createRelation("stranger_rel", "docs", "x", "docs"),
+       "stranger is denied on CreateRelation");
+    ck(!ops.getRelationInfo("stranger_rel").has_value(),
+       "and nothing was actually created by the denied attempt");
+
+    ck(!stranger.dropRelation("docs_rel"),
+       "stranger is denied on DropRelation");
+    ck(ops.getRelationInfo("docs_rel").has_value(),
+       "and 'docs_rel' still exists - the denied attempt did not drop it");
+
+    {
+        auto strangerList = stranger.listRelations();
+        bool sawDocsRel = false;
+        for (const auto& r : strangerList) if (r.name == "docs_rel") sawDocsRel = true;
+        ck(!sawDocsRel, "stranger's ListRelations does not include 'docs_rel'");
+
+        auto opsList = ops.listRelations();
+        bool adminSawDocsRel = false;
+        for (const auto& r : opsList) if (r.name == "docs_rel") adminSawDocsRel = true;
+        ck(adminSawDocsRel, "admin's ListRelations DOES include it - proving the "
+                            "stranger's empty-of-it result is the gate, not the "
+                            "relation being gone");
+    }
+
+    ck(!stranger.getRelationInfo("docs_rel").has_value(),
+       "stranger is denied on GetRelationInfo for a relation that DOES "
+       "exist (proven by the admin listing above) - so this is a real "
+       "denial, not 'not found'");
+    ck(ops.getRelationInfo("docs_rel").has_value(),
+       "admin's GetRelationInfo still works for the same name");
+
     std::cout << "\npassed=" << pass << " failed=" << fail << "\n";
     return fail == 0 ? 0 : 1;
 }

+ 268 - 0
tests/load_test/test_relations.cpp

@@ -0,0 +1,268 @@
+// v2.11.0 T9 — relations end to end, over gRPC, through the real client,
+// including a real service restart and a manufactured self-heal fixture.
+//
+// tests/load_test/test_relations_client_e2e.cpp already covers the basic
+// client/server-boundary shape (non-default project, createRelation/
+// listRelations/getRelationInfo qualify-by-name correctly, restrict blocks
+// a delete, cross-project relation names don't collide). This driver adds
+// exactly what that suite structurally cannot, because it never restarts
+// the service:
+//
+//   1. A relation declaration SURVIVES a restart, and — the load-bearing
+//      half — a child document written AFTER the restart is indexed. That
+//      proves DatabaseService::applyRelationDeclarations() re-armed the
+//      write-path maintenance hook, not merely that RelationManager
+//      remembered the row. (Mirrors test_indexes.cpp's identical proof for
+//      secondary indexes.)
+//   2. Boot-time self-heal: with the relation's `_relidx1_*` reverse-index
+//      sub-db manufactured missing (via relidx_drop_tool, run by
+//      test_relations.sh while the service is stopped), a restart must
+//      rebuild the index from the CHILD ROWS THAT ALREADY EXISTED before
+//      the sub-db was dropped — not just index rows written afterward —
+//      and enforcement on the original relationship must work again.
+//   3. Per-project isolation: two projects each declare a relation under
+//      the identical bare name; a project must not see the other
+//      project's children when evaluating restrict/DescribeDelete. Same
+//      failure shape as the v2.4.2 createView bug.
+//
+// Usage: test_relations <address> <projectA> <projectB> <phase>
+//   phase=setup           - declare relations in both projects, insert data,
+//                            assert restrict/no_action/DescribeDelete/isolation
+//   phase=verify           - AFTER restart #1 (plain restart, no sub-db
+//                            tampering): declaration survived, restrict still
+//                            blocks the pre-restart relationship, and a NEW
+//                            child written post-restart is indexed
+//   phase=verify_selfheal - AFTER restart #2 (with projectA's relidx sub-db
+//                            dropped between stop and start): restrict still
+//                            blocks the ORIGINAL pre-drop relationship
+//                            (proving the rebuild recovered existing rows,
+//                            not just future writes), and post-restart
+//                            writes are still indexed too
+
+#include <smartbotic/database/client.hpp>
+
+#include <iostream>
+#include <string>
+
+using namespace smartbotic::database;
+
+namespace {
+
+int g_pass = 0;
+int g_fail = 0;
+
+void check(bool cond, const std::string& msg) {
+    if (cond) {
+        ++g_pass;
+        std::cout << "  ok   " << msg << "\n";
+    } else {
+        ++g_fail;
+        std::cout << "  FAIL " << msg << "\n";
+    }
+}
+
+constexpr const char* kRelation = "wf_exec";
+constexpr const char* kParent = "workflows";
+constexpr const char* kChild = "executions";
+constexpr const char* kChildField = "workflowId";
+
+// Refusal text is pinned from service/src/relations/relation_enforcement.cpp
+// (formatRelationBlockError) — NOT transcribed from any report. A prior task
+// report mis-transcribed "child document(s)" as "document(s)"; assert the
+// real string so that mistake can't repeat silently here.
+bool looksLikeRestrictRefusal(const std::string& err, const std::string& qualifiedRelation,
+                              const std::string& qualifiedChildColl, uint64_t count) {
+    if (err.find("cannot delete '") == std::string::npos) return false;
+    if (err.find("relation '" + qualifiedRelation + "' has " + std::to_string(count) +
+                  " child document(s) in '" + qualifiedChildColl + "' referencing it") ==
+        std::string::npos) {
+        return false;
+    }
+    if (err.find("or change the relation's on_delete to no_action to permit the "
+                 "dangling reference.") == std::string::npos) {
+        return false;
+    }
+    return true;
+}
+
+void setup(Client& a, Client& b, const std::string& projA, const std::string& projB) {
+    std::cout << "-- setup --\n";
+
+    // Identically-named relation declared independently in two projects.
+    // RelationManager keys by the project-qualified name, so this must NOT
+    // collide — same shape the v2.4.2 createView bug got wrong.
+    check(a.createRelation(kRelation, kChild, kChildField, kParent),
+          "projectA declares 'wf_exec'");
+    check(b.createRelation(kRelation, kChild, kChildField, kParent),
+          "projectB declares the SAME bare name 'wf_exec' without collision");
+
+    // A no_action relation in project A, to prove restrict isn't the only
+    // policy path this suite exercises.
+    check(a.createRelation("wf_logs_na", "logs", kChildField, kParent, "no_action"),
+          "projectA declares a no_action relation");
+
+    // Parent + child in A only.
+    a.upsert(kParent, nlohmann::json{{"name", "wf-1"}}, "wf-1");
+    a.upsert(kChild, nlohmann::json{{kChildField, "wf-1"}}, "ex-a1");
+
+    // Same parent id exists in B, but with NO referencing child — this is
+    // the isolation probe: B's identically-named relation must report zero
+    // children for a parent id that DOES have a child in A.
+    b.upsert(kParent, nlohmann::json{{"name", "wf-1"}}, "wf-1");
+
+    // ---- DescribeDelete, non-destructive, before any delete attempt ------
+    auto descA = a.describeDelete(kParent, "wf-1");
+    check(descA.success, "projectA describeDelete succeeds");
+    bool foundA = false;
+    for (const auto& imp : descA.impacts) {
+        if (imp.relation != kRelation) continue;
+        foundA = true;
+        check(imp.childCollection == kChild, "impact reports the bare child collection");
+        check(imp.childField == kChildField, "impact reports the child field");
+        check(imp.onDelete == "restrict", "impact reports the on_delete policy");
+        check(imp.childCount == 1, "projectA sees exactly its own 1 child");
+        check(imp.blocks, "impact.blocks is true (relations_enforced defaults on)");
+    }
+    check(foundA, "projectA describeDelete includes the wf_exec impact");
+
+    auto descB = b.describeDelete(kParent, "wf-1");
+    check(descB.success, "projectB describeDelete succeeds");
+    bool foundB = false;
+    for (const auto& imp : descB.impacts) {
+        if (imp.relation != kRelation) continue;
+        foundB = true;
+        check(imp.childCount == 0,
+              "projectB sees ZERO children for the SAME parent id 'wf-1' - "
+              "A's child does not leak across the project boundary");
+        check(!imp.blocks, "projectB's describeDelete is not blocked");
+    }
+    check(foundB, "projectB describeDelete includes its own wf_exec impact");
+
+    // ---- restrict actually blocks in A, and the message names the relation
+    std::string err;
+    bool deleted = a.remove(kParent, "wf-1", err);
+    check(!deleted, "projectA: restrict blocks deleting a referenced parent");
+    check(looksLikeRestrictRefusal(err, projA + ":" + kRelation, projA + ":" + kChild, 1),
+          "the refusal message matches formatRelationBlockError's real text "
+          "('child document(s)', not 'document(s)')");
+
+    // ---- restrict does NOT block in B: no children reference wf-1 there --
+    err.clear();
+    deleted = b.remove(kParent, "wf-1", err);
+    check(deleted, "projectB: the identically-named relation does not block - "
+                   "isolation holds for the actual delete, not just DescribeDelete");
+    check(err.empty(), "no error on projectB's successful delete");
+
+    // ---- no_action permits a dangling reference -----------------------
+    a.upsert(kParent, nlohmann::json{{"name", "wf-2"}}, "wf-2");
+    a.upsert("logs", nlohmann::json{{kChildField, "wf-2"}}, "log-a1");
+    err.clear();
+    deleted = a.remove(kParent, "wf-2", err);
+    check(deleted, "no_action permits deleting a parent with a dangling child");
+    check(err.empty(), "no error on the no_action delete");
+}
+
+void verify(Client& a, const std::string& projA) {
+    std::cout << "-- verify after restart #1 (plain restart) --\n";
+
+    auto listed = a.listRelations();
+    bool found = false;
+    for (const auto& r : listed) if (r.name == kRelation) found = true;
+    check(found, "the declaration SURVIVED the restart - listRelations still "
+                 "reports 'wf_exec'");
+
+    const std::string qualRelation = projA + ":" + kRelation;
+    const std::string qualChild = projA + ":" + kChild;
+
+    // The pre-restart relationship must still block.
+    std::string err;
+    bool deleted = a.remove(kParent, "wf-1", err);
+    check(!deleted, "restrict still blocks the pre-restart relationship "
+                    "(wf-1/ex-a1) after the restart");
+    check(looksLikeRestrictRefusal(err, qualRelation, qualChild, 1),
+          "the refusal still matches the pinned restrict-block text");
+
+    // THE load-bearing check: a child written AFTER the restart must be
+    // indexed. If applyRelationDeclarations() had merely restored the
+    // in-memory RelationManager cache without calling set_relations() on
+    // the LMDB store, writes after this point would go unmaintained and
+    // this would silently NOT block.
+    a.upsert(kParent, nlohmann::json{{"name", "wf-3"}}, "wf-3");
+    a.upsert(kChild, nlohmann::json{{kChildField, "wf-3"}}, "ex-a3");
+    err.clear();
+    deleted = a.remove(kParent, "wf-3", err);
+    check(!deleted, "a child written AFTER the restart IS indexed - the "
+                    "write path is armed, not merely remembered");
+    check(looksLikeRestrictRefusal(err, qualRelation, qualChild, 1),
+          "the refusal for the post-restart relationship matches the "
+          "pinned text too");
+}
+
+void verifySelfHeal(Client& a, const std::string& projA) {
+    std::cout << "-- verify after restart #2 (relidx sub-db manufactured "
+                 "missing, then restarted) --\n";
+
+    auto listed = a.listRelations();
+    bool found = false;
+    for (const auto& r : listed) if (r.name == kRelation) found = true;
+    check(found, "the declaration still exists after the self-heal restart "
+                 "(dropping the INDEX sub-db must not touch the declaration "
+                 "in RelationManager's own store)");
+
+    const std::string qualRelation = projA + ":" + kRelation;
+    const std::string qualChild = projA + ":" + kChild;
+
+    // THE point of this phase: wf-1/ex-a1 was written and indexed BEFORE
+    // the sub-db was dropped. If the self-heal rebuild only covered new
+    // writes (or didn't run at all), this relationship would no longer
+    // block, and a real delete of wf-1 would silently succeed, leaving
+    // ex-a1 dangling with nothing to say so.
+    std::string err;
+    bool deleted = a.remove(kParent, "wf-1", err);
+    check(!deleted, "restrict STILL blocks the ORIGINAL relationship after "
+                    "the sub-db was dropped and the service restarted - "
+                    "the boot-time self-heal rebuilt it from the existing "
+                    "child row, not merely from future writes");
+    check(looksLikeRestrictRefusal(err, qualRelation, qualChild, 1),
+          "the refusal matches the pinned restrict-block text");
+
+    // And maintenance is still armed post-self-heal for brand new writes.
+    a.upsert(kParent, nlohmann::json{{"name", "wf-4"}}, "wf-4");
+    a.upsert(kChild, nlohmann::json{{kChildField, "wf-4"}}, "ex-a4");
+    err.clear();
+    deleted = a.remove(kParent, "wf-4", err);
+    check(!deleted, "a child written after the self-heal restart is indexed too");
+}
+
+}  // namespace
+
+int main(int argc, char** argv) {
+    if (argc < 5) {
+        std::cerr << "usage: test_relations <address> <projectA> <projectB> "
+                     "<setup|verify|verify_selfheal>\n";
+        return 2;
+    }
+    const std::string address = argv[1];
+    const std::string projA = argv[2];
+    const std::string projB = argv[3];
+    const std::string phase = argv[4];
+
+    Client a({.address = address, .project = projA});
+    a.connect();
+
+    if (phase == "setup") {
+        Client b({.address = address, .project = projB});
+        b.connect();
+        setup(a, b, projA, projB);
+    } else if (phase == "verify") {
+        verify(a, projA);
+    } else if (phase == "verify_selfheal") {
+        verifySelfHeal(a, projA);
+    } else {
+        std::cerr << "unknown phase '" << phase << "'\n";
+        return 2;
+    }
+
+    std::cout << "\npassed=" << g_pass << " failed=" << g_fail << "\n";
+    return g_fail == 0 ? 0 : 1;
+}

+ 116 - 0
tests/load_test/test_relations.sh

@@ -0,0 +1,116 @@
+#!/usr/bin/env bash
+# v2.11.0 T9 — relations end to end, including a real restart and a
+# manufactured boot-time self-heal fixture. See test_relations.cpp for what
+# each phase asserts and why. Companion to test_relations_client_e2e.sh
+# (basic client/server-boundary shape, no restart) — this script exists
+# because the restart + sub-db-tampering machinery needed here has no
+# equivalent in that single-boot harness. Modeled on test_indexes.sh, which
+# established the setup/verify restart pattern for secondary indexes.
+
+set -uo pipefail
+
+BUILD="${1:-build}"
+REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
+PORT=19413
+PROJECT_A=relproj_a
+PROJECT_B=relproj_b
+WORK="$(mktemp -d)"
+DRIVER="$WORK/test_relations"
+DROP_TOOL="$WORK/relidx_drop_tool"
+SERVER_PID=""
+
+cleanup() {
+    [[ -n "$SERVER_PID" ]] && kill "$SERVER_PID" 2>/dev/null
+    wait "$SERVER_PID" 2>/dev/null
+    rm -rf "$WORK"
+}
+trap cleanup EXIT
+
+fail() { echo "FAIL: $*" >&2; exit 1; }
+
+echo "=== building the driver and the offline relidx-drop tool ==="
+g++ -std=c++20 -O1 -o "$DRIVER" "$REPO/tests/load_test/test_relations.cpp" \
+    -I"$REPO/client/include" -I"$REPO/$BUILD/client" \
+    -L"$REPO/$BUILD/client" -lsmartbotic-db-client -lspdlog -lfmt \
+    || fail "driver did not compile"
+g++ -std=c++20 -O1 -o "$DROP_TOOL" "$REPO/tests/load_test/relidx_drop_tool.cpp" \
+    -llmdb \
+    || fail "relidx_drop_tool did not compile"
+
+mkdir -p "$WORK/data"
+cat > "$WORK/config.json" <<EOF
+{
+  "storage": {
+    "data_directory": "$WORK/data",
+    "bind_address": "127.0.0.1",
+    "rpc_port": $PORT,
+    "encryption": { "enabled": false, "key_file": "$WORK/data/storage.key" },
+    "migrations": { "enabled": false },
+    "replication": { "enabled": false }
+  }
+}
+EOF
+
+start_server() {
+    local log="$1"
+    "$REPO/$BUILD/service/smartbotic-database" --config "$WORK/config.json" > "$log" 2>&1 &
+    SERVER_PID=$!
+    for _ in $(seq 1 60); do
+        grep -q "Notified systemd: READY" "$log" 2>/dev/null && return 0
+        kill -0 "$SERVER_PID" 2>/dev/null || { cat "$log"; fail "server exited during startup"; }
+        sleep 0.5
+    done
+    cat "$log"
+    fail "server never became ready"
+}
+
+stop_server() {
+    [[ -n "$SERVER_PID" ]] || return 0
+    kill "$SERVER_PID" 2>/dev/null
+    wait "$SERVER_PID" 2>/dev/null
+    SERVER_PID=""
+}
+
+export LD_LIBRARY_PATH="$REPO/$BUILD/client:${LD_LIBRARY_PATH:-}"
+
+echo
+echo "=== boot 1: declare relations, insert data, assert restrict/no_action/isolation ==="
+start_server "$WORK/boot1.log"
+"$DRIVER" "127.0.0.1:$PORT" "$PROJECT_A" "$PROJECT_B" setup || fail "setup phase"
+
+echo
+echo "=== restarting the service (plain restart, no tampering) ==="
+stop_server
+start_server "$WORK/boot2.log"
+
+grep -q "v2.11 relations" "$WORK/boot2.log" \
+    || { grep -iE "relation|error" "$WORK/boot2.log" | tail -20
+         fail "the service did not re-apply relation declarations at boot"; }
+echo "  boot log: $(grep 'v2.11 relations' "$WORK/boot2.log" | tail -1 | sed 's/.*\] //')"
+
+echo
+echo "=== phase: verify declaration + write-path survived the restart ==="
+"$DRIVER" "127.0.0.1:$PORT" "$PROJECT_A" "$PROJECT_B" verify || fail "verify phase"
+
+echo
+echo "=== manufacturing a missing reverse-index sub-db (service stopped) ==="
+stop_server
+ENV_A="$WORK/data/projects/$PROJECT_A/env"
+[[ -d "$ENV_A" ]] || fail "expected project A env dir at $ENV_A"
+"$DROP_TOOL" "$ENV_A" "_relidx1_wf_exec" || fail "relidx_drop_tool could not drop the sub-db"
+
+echo
+echo "=== restarting the service (self-heal must rebuild the dropped sub-db) ==="
+start_server "$WORK/boot3.log"
+
+grep -q "rebuilt 'wf_exec'" "$WORK/boot3.log" \
+    || { grep -iE "relation|error" "$WORK/boot3.log" | tail -20
+         fail "boot did not self-heal the missing relidx sub-db - the T7 self-heal path did not run"; }
+echo "  boot log: $(grep "rebuilt 'wf_exec'" "$WORK/boot3.log" | tail -1 | sed 's/.*\] //')"
+
+echo
+echo "=== phase: verify the self-heal rebuilt the index and enforcement works again ==="
+"$DRIVER" "127.0.0.1:$PORT" "$PROJECT_A" "$PROJECT_B" verify_selfheal || fail "verify_selfheal phase"
+
+echo
+echo "ALL RELATIONS E2E CHECKS PASSED"

+ 43 - 0
tests/load_test/test_relations_client_e2e.cpp

@@ -85,6 +85,49 @@ int main() {
 
     ck(c.setRelationsEnforced("workflows", true), "re-enable relations_enforced");
 
+    // ---- Part 3b: DescribeDelete over the boundary, with project
+    // qualification. Only exercised in the `default` project by tests/T5's
+    // own unit coverage; this is the first time it goes through a real
+    // Client with a non-default project, so it is the first place a
+    // qualify/unqualify mismatch in the DescribeDelete RPC specifically
+    // would be observable (see client.cpp's describeDelete: `collection` is
+    // qualified like get()/find(), `relation`/`childCollection` in the
+    // response are unqualified back to bare names the same way
+    // listRelations() unqualifies).
+    c.upsert("workflows", nlohmann::json{{"name", "wf-2"}}, "wf-2");
+    c.upsert("executions", nlohmann::json{{"workflowId", "wf-2"}}, "ex-2");
+
+    auto desc = c.describeDelete("workflows", "wf-2");
+    ck(desc.success, "describeDelete succeeds for a non-default project");
+    ck(desc.wouldBeBlocked, "describeDelete reports the delete would be blocked");
+    bool sawImpact = false;
+    for (const auto& imp : desc.impacts) {
+        if (imp.relation != "exec_wf") continue;
+        sawImpact = true;
+        ck(imp.childCollection == "executions",
+           "describeDelete reports the BARE child collection, unqualified "
+           "back from 'acme:executions'");
+        ck(imp.childField == "workflowId", "describeDelete reports child_field verbatim");
+        ck(imp.onDelete == "restrict", "describeDelete reports the on_delete policy");
+        ck(imp.childCount == 1, "describeDelete counts exactly the one referencing child");
+        ck(imp.blocks, "describeDelete's per-impact blocks agrees with wouldBeBlocked");
+        ck(!imp.sampleChildIds.empty() && imp.sampleChildIds[0] == "ex-2",
+           "describeDelete samples the actual referencing child id");
+    }
+    ck(sawImpact, "describeDelete's impacts include the 'exec_wf' relation by "
+                  "its BARE name - the qualify/unqualify round trip works for "
+                  "this RPC specifically, not just listRelations/getRelationInfo");
+
+    // DescribeDelete must never actually delete anything.
+    ck(c.get("workflows", "wf-2").has_value(),
+       "describeDelete is read-only - the parent document still exists");
+
+    // And a real delete of wf-2 is in fact still blocked, matching what
+    // describeDelete predicted.
+    err.clear();
+    ck(!c.remove("workflows", "wf-2", err), "the real delete agrees with describeDelete");
+    (void)c.remove("executions", "ex-2");  // leave wf-2 behind; harmless
+
     ck(c.dropRelation("exec_wf"), "dropRelation succeeds");
     ck(!c.getRelationInfo("exec_wf").has_value(), "relation is gone after drop");
     ck(!c.dropRelation("exec_wf"), "dropping a gone relation fails (not idempotent-success)");