|
|
@@ -1,12 +1,71 @@
|
|
|
# Relations: referential integrity for a document store
|
|
|
|
|
|
-> **STATUS: not implemented.** No `_relations` collection, `RelationManager` or
|
|
|
-> relation RPC exists in the codebase. This design and its companion plan
|
|
|
-> (`docs/superpowers/plans/2026-08-04-relations-v2.5.0.md`) were never executed.
|
|
|
-> The "v2.5.0" target is wrong - v2.5.0 was skipped entirely (2.4.5 → 2.6.0).
|
|
|
-> Re-validate against the current write path before using either: both predate
|
|
|
-> the v2.8.0 fix for caching an `MDB_dbi` before commit.
|
|
|
-
|
|
|
+> **STATUS: not implemented. RE-VALIDATED 2026-08-09 against the code at v2.10.0.**
|
|
|
+>
|
|
|
+> The design's *reasoning* holds up. Several of its facts and one of its design
|
|
|
+> choices do not, and there are new hard constraints it could not have known
|
|
|
+> about. Read this section before the body; where they disagree, this section is
|
|
|
+> right.
|
|
|
+>
|
|
|
+> **Still true, checked:**
|
|
|
+> - No `_relations` collection, `RelationManager` or relation RPC exists.
|
|
|
+> - `FilterOp` still has exactly eleven values, all single-collection predicates.
|
|
|
+> - `MemoryStore::update` still forces `updated.id = id` (`memory_store.cpp:502`
|
|
|
+> and `:802`), so the "no `ON UPDATE`" non-goal still rests on solid ground.
|
|
|
+> - MemoryStore is still in the write path, and `_`-prefixed collections are
|
|
|
+> still excluded from the LMDB mirror (`dual_write_mirror.hpp:43`), so
|
|
|
+> `_relations` lives in MemoryStore and is WAL'd/snapshotted - exactly parallel
|
|
|
+> to `_views`, as the body claims.
|
|
|
+> - Test targets still build with `-UNDEBUG`, so assertions stay live in Release.
|
|
|
+> - The restrict-race analysis and `validate_on_write` as its remedy still stand.
|
|
|
+>
|
|
|
+> **Wrong or stale:**
|
|
|
+> 1. **Target version.** v2.5.0 was skipped entirely (2.4.5 → 2.6.0). The project
|
|
|
+> is at **2.10.0**. Every "v2.5.0" in this document and its plan is wrong.
|
|
|
+> 2. **"Phase C" collides.** The rollout's "Phase C (later)" now clashes with the
|
|
|
+> abandoned v2.0 storage Phase C. Do not reuse that vocabulary - see
|
|
|
+> `docs/ROADMAP.md`, "Storage engine history".
|
|
|
+> 3. **A line citation has rotted.** The body cites
|
|
|
+> `docs/2026-05-15-v2.0-storage-engine-design.md:316` for "SQL surface"; it is
|
|
|
+> now line **329**. Cite text, not line numbers.
|
|
|
+> 4. **The non-goal "no candidate-key or unique-constraint machinery" is out of
|
|
|
+> date.** v2.10.0 built it - `find_duplicate_values` plus enforcement inside
|
|
|
+> the document's own transaction - and then deliberately left it unreachable
|
|
|
+> because it cannot be enforced through the mirror (see below). So references
|
|
|
+> to a non-`_id` field become *reachable* once the write path is fixed, which
|
|
|
+> is the same fix relations needs. They are one piece of work, not two.
|
|
|
+> 5. **"Every `LmdbDocumentStore` operation opens its own `WriteTxn`" is now only
|
|
|
+> half true.** Internal helpers already take a `WriteTxn&` (`open_for_write`,
|
|
|
+> `maintainIndexes`, `markIndexMultiValued`); eight public operations still
|
|
|
+> open their own. The restructuring the body asks for is therefore smaller
|
|
|
+> than it was, and partly precedented.
|
|
|
+>
|
|
|
+> **New hard constraints (all post-date this spec):**
|
|
|
+> 6. **⚠ A sub-db created at runtime MUST be registered with
|
|
|
+> `cacheCommittedDbi()` after its transaction commits.** Since v2.8.1 the read
|
|
|
+> path never calls `mdb_dbi_open` - `try_open_for_read` serves only from the
|
|
|
+> primed cache, and a miss means "no such sub-db". So a `_relidx_` sub-db
|
|
|
+> created without that call would be **invisible to every later read**, and
|
|
|
+> relations would silently not enforce. This is the single most likely way to
|
|
|
+> get relations wrong now, and nothing in the body warns about it.
|
|
|
+> 7. **⚠ Never cache an `MDB_dbi` before its transaction commits** (v2.8.0), and
|
|
|
+> **never open a second `MDB_env` on a path this process already has open**
|
|
|
+> (v2.4.4 - POSIX locks are per-process). Both cost production outages.
|
|
|
+> 8. **Index sub-dbs now carry TWO reserved keys**, not one: the v2.4.4 identity
|
|
|
+> sentinel and the v2.10.0 multivalued marker. Use `is_index_meta_key()`;
|
|
|
+> checking only `is_identity_key()` will miscount and mis-walk.
|
|
|
+> 9. **Give `_relidx_` a key-format version in its name from the first commit**
|
|
|
+> (`_relidx1_`). v2.9.1 had to bump `_idx_` → `_idx2_` when an encoding
|
|
|
+> changed, precisely so a stale index is never read under new rules. Paying
|
|
|
+> that forward costs nothing now.
|
|
|
+>
|
|
|
+> **One design choice is now beaten - see "Reverse index" below for the
|
|
|
+> replacement.**
|
|
|
+>
|
|
|
+> The companion plan (`docs/superpowers/plans/2026-08-04-relations-v2.5.0.md`,
|
|
|
+> 3278 lines) is **superseded**: it was written against the pre-v2.8 write path
|
|
|
+> and its task list embeds the wrong version and the stale facts above. Treat it
|
|
|
+> as a design input and re-plan rather than executing it.
|
|
|
|
|
|
Status: approved design, not yet implemented
|
|
|
Target: v2.5.0 (Phases A + B together)
|
|
|
@@ -116,18 +175,36 @@ leaving operators to discover it.
|
|
|
|
|
|
### Reverse index
|
|
|
|
|
|
-One LMDB sub-db per relation, `_relidx_<name>`, inside the project env. Key is
|
|
|
-the composite `<parentId>\0<childId>`; value empty. LMDB orders keys, so a
|
|
|
-parent's children are a cursor `MDB_SET_RANGE` over the `<parentId>\0` prefix -
|
|
|
-O(children), not O(collection).
|
|
|
-
|
|
|
-A list-valued index (`parentId -> [childIds]`) is rejected: it turns every child
|
|
|
-insert into a read-modify-write on one key shared by all siblings.
|
|
|
+One LMDB sub-db per relation, `_relidx1_<name>`, inside the project env.
|
|
|
+
|
|
|
+> **RE-VALIDATED: use `MDB_DUPSORT`, not composite keys.** This section
|
|
|
+> originally specified key = `<parentId>\0<childId>` with an empty value, walked
|
|
|
+> with `MDB_SET_RANGE` over the `<parentId>\0` prefix. That was right when
|
|
|
+> nothing else existed. Since v2.9.0 the secondary-index machinery does exactly
|
|
|
+> this shape as **key = parent id, data = child id, DUPSORT** - built, tested and
|
|
|
+> measured on production-sized data.
|
|
|
+>
|
|
|
+> Three reasons to switch:
|
|
|
+> 1. **`mdb_cursor_count` gives a parent's child count without reading the
|
|
|
+> children.** That is precisely what `restrict` and `DescribeDelete` need, and
|
|
|
+> it is O(1)-ish. The composite-key form has to walk the range to count.
|
|
|
+> 2. Removing one id from a parent's set is `mdb_del(key, data)`, which deletes
|
|
|
+> just that pair - no read-modify-write, so the rejection of a list-valued
|
|
|
+> index below is satisfied without a bespoke encoding.
|
|
|
+> 3. The surrounding discipline already exists and is tested: reserved-key
|
|
|
+> handling (`is_index_meta_key`), handle caching after commit
|
|
|
+> (`cacheCommittedDbi`), boot priming, and the identity sentinel.
|
|
|
+>
|
|
|
+> A list-valued index (`parentId -> [childIds]` as one value) stays rejected, for
|
|
|
+> the reason the original text gives: it turns every child insert into a
|
|
|
+> read-modify-write on one key shared by all siblings. DUPSORT is not that - it
|
|
|
+> stores the set as separate data items.
|
|
|
|
|
|
Index sub-dbs carry the v2.4.4 identity sentinel like any other sub-db, and
|
|
|
-`count()` and `scan()` skip the sentinel key. Without this a stale `MDB_dbi`
|
|
|
-could write index entries into an unrelated sub-db, which is precisely the
|
|
|
-failure that misfiled 31 production rows.
|
|
|
+`count()` and `scan()` skip **every** reserved key - as of v2.10.0 there are two
|
|
|
+(identity, and the multivalued marker), so use `is_index_meta_key()`. Without
|
|
|
+this a stale `MDB_dbi` could write index entries into an unrelated sub-db, which
|
|
|
+is precisely the failure that misfiled 31 production rows.
|
|
|
|
|
|
### Write-path change
|
|
|
|
|
|
@@ -175,6 +252,61 @@ documented rather than hidden.
|
|
|
validate existing data against the policy. On a large collection this blocks, so
|
|
|
it belongs in a migration rather than a live call. The scan is idempotent.
|
|
|
|
|
|
+## Per-collection enable/disable
|
|
|
+
|
|
|
+**Requested 2026-08-09.** Both relation enforcement and uniqueness must be
|
|
|
+switchable per collection.
|
|
|
+
|
|
|
+Home: **`CollectionCfg`, in the `_collection_meta` system collection.** That is
|
|
|
+where `timestampPrecision`, `versioningEnabled` and (since v2.9.0)
|
|
|
+`indexedFields` already live, and the reason is durability: `_collection_meta` is
|
|
|
+an ordinary collection, so it is WAL'd and snapshotted for free. Putting these in
|
|
|
+`CollectionOptions` instead would need a new WAL op to survive a restart between
|
|
|
+snapshots - the same reasoning recorded for `versioningEnabled` in v2.4.5.
|
|
|
+
|
|
|
+```
|
|
|
+struct CollectionCfg {
|
|
|
+ std::string timestampPrecision = "ns";
|
|
|
+ bool versioningEnabled = true;
|
|
|
+ std::vector<std::string> indexedFields;
|
|
|
+ std::vector<std::string> uniqueFields; // per-collection by construction
|
|
|
+ bool relationsEnforced = true; // new
|
|
|
+};
|
|
|
+```
|
|
|
+
|
|
|
+**`uniqueFields`** is inherently per-collection - it names fields of one
|
|
|
+collection - so it needs no separate switch. It must be a **subset of
|
|
|
+`indexedFields`**: the check reads the index, so uniqueness without an index has
|
|
|
+nothing to read. Declaring uniqueness over data that already contains duplicates
|
|
|
+is **refused**, with examples, rather than accepted: a constraint that is false
|
|
|
+from the moment it is created would fail later writes for reasons the caller
|
|
|
+never caused. `find_duplicate_values()` already does this.
|
|
|
+
|
|
|
+**`relationsEnforced`** defaults to **true**, because declaring a relation names
|
|
|
+its child and parent collections explicitly - the declaration *is* the opt-in,
|
|
|
+and a declared constraint that silently does nothing would be worse than no
|
|
|
+constraint. The switch is an operator escape hatch for the cases that genuinely
|
|
|
+need one: a bulk import, or a collection under write pressure where the
|
|
|
+`restrict` check costs more than the integrity is worth.
|
|
|
+
|
|
|
+Two things this must get right, both learned the hard way here:
|
|
|
+
|
|
|
+- **Disabling is not retroactive and must say so.** Turning enforcement off then
|
|
|
+ deleting parents creates dangling references that turning it back on will not
|
|
|
+ detect - only the `relations check` command will. The RPC response and the CLI
|
|
|
+ must state that at the point of use, not only in documentation.
|
|
|
+- **Re-arm on boot.** `DatabaseService::applyIndexDeclarations()` already re-reads
|
|
|
+ `indexedFields` at startup and exists for exactly this reason: a declaration
|
|
|
+ that is persisted but not applied leaves the write path not maintaining
|
|
|
+ something the read path still trusts. Relation and uniqueness switches need the
|
|
|
+ same treatment in the same place, and it is load-bearing, not bookkeeping.
|
|
|
+
|
|
|
+Both switches belong on `ConfigureCollection`, which is already a **partial
|
|
|
+update**: absent means "leave unchanged". Use `optional bool` for
|
|
|
+`relations_enforced` - a plain proto3 bool defaults to false and would silently
|
|
|
+disable enforcement on any unrelated config call, which is the exact trap v2.4.5
|
|
|
+hit with `versioning_enabled`.
|
|
|
+
|
|
|
## Surface
|
|
|
|
|
|
**RPCs**, mirroring the view surface: `CreateRelation`, `DropRelation`,
|