Parcourir la source

docs: make the tree state what is actually true

The docs described a system that was never built. The v2.0 roadmap's "Phase C"
specified a homegrown engine - 16 KB slotted pages, BPlusTree, BufferPool,
page-level WAL redo - but the build-vs-buy spike chose LMDB and that is what
shipped as v2.0. None of those types exist. Source comments still promised a
v2.1 rename of max_memory_mb to buffer_pool_size_mb "per the Phase C plan";
that knob has never existed and the rename is not planned. Two BUG-*.md files
at the repo root still read "Open" for issues fixed in v1.6.1 and superseded
by v2.0. A 3278-line relations plan was labelled v2.5.0, a version that was
skipped entirely.

This is not cosmetic: reading those docs produced a wrong answer about where
the "binary frame" work sits (it is a Stage 4 note in document_store_lmdb.cpp,
unrelated to Phase C), and the same reading missed that encode_document still
serialises with nlohmann::dump().

- docs/ROADMAP.md is new and authoritative: current version, what shipped,
  what is pending, what is explicitly not planned. It states that if any other
  document disagrees with it, the other document is stale.
- Plans for shipped work move to docs/superpowers/plans/archive/ with a README
  mapping each to its release and a do-not-execute warning. plans/ now holds
  only the one plan that is genuinely pending.
- Superseded forward-looking docs deleted (v2.0-roadmap, the Phase C skeleton
  with its TBD tech stack, v2.1-backlog whose still-true items folded into
  ROADMAP.md; its D.5 and D.7 items turned out to be implemented).
- Resolved BUG-*.md and the completed root plan.md move to docs/incidents/
  with a README recording how each was actually resolved.
- Every spec and the storage design doc gain a status banner: implemented,
  superseded, or never implemented.
- Stale "Phase C" vocabulary removed from source comments, replaced with what
  the code actually does plus a pointer to ROADMAP.md. Five comments
  referencing the moved BUG-memory-leak-zoe.md repointed.
- README.md and CLAUDE.md now point at ROADMAP.md first.

Build clean; no behaviour change.
fszontagh il y a 1 mois
Parent
commit
60128349a7
37 fichiers modifiés avec 323 ajouts et 421 suppressions
  1. 8 0
      CLAUDE.md
  2. 4 0
      README.md
  3. 8 0
      docs/2026-03-26-vector-storage-design.md
  4. 13 0
      docs/2026-05-15-v2.0-storage-engine-design.md
  5. 161 0
      docs/ROADMAP.md
  6. 0 0
      docs/incidents/2026-04-19-snapshot-truncation.md
  7. 0 0
      docs/incidents/2026-04-22-eviction-induced-write-loss-plan.md
  8. 0 0
      docs/incidents/2026-04-22-zoe-rss-exceeds-budget.md
  9. 13 0
      docs/incidents/README.md
  10. 0 118
      docs/superpowers/plans/2026-05-15-v2.0-roadmap.md
  11. 0 176
      docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c.md
  12. 0 107
      docs/superpowers/plans/2026-05-16-v2.1-backlog.md
  13. 13 0
      docs/superpowers/plans/2026-08-04-relations-v2.5.0.md
  14. 0 0
      docs/superpowers/plans/archive/2026-03-26-vector-storage.md
  15. 0 0
      docs/superpowers/plans/archive/2026-04-12-packaging-consolidation.md
  16. 0 0
      docs/superpowers/plans/archive/2026-04-13-views-v1.5.0.md
  17. 1 1
      docs/superpowers/plans/archive/2026-04-19-snapshot-recovery-v1.6.1.md
  18. 0 0
      docs/superpowers/plans/archive/2026-05-15-v1.10-yyjson-phase-a.md
  19. 0 0
      docs/superpowers/plans/archive/2026-05-15-v1.11-binary-docs-phase-b.md
  20. 0 0
      docs/superpowers/plans/archive/2026-05-15-v1.9.5-backup-before-upgrade.md
  21. 0 0
      docs/superpowers/plans/archive/2026-05-15-v2.0-storage-engine-phase-c-execution.md
  22. 0 0
      docs/superpowers/plans/archive/2026-05-16-v2.3-multi-project.md
  23. 0 0
      docs/superpowers/plans/archive/2026-05-31-v2.4-listeners-tls-auth.md
  24. 0 0
      docs/superpowers/plans/archive/2026-08-08-project-scoped-files-v2.6.0.md
  25. 30 0
      docs/superpowers/plans/archive/README.md
  26. 6 0
      docs/superpowers/specs/2026-05-15-binary-doc-format-decision.md
  27. 7 0
      docs/superpowers/specs/2026-05-15-v2.0-storage-engine-architecture.md
  28. 8 0
      docs/superpowers/specs/2026-08-03-relations-design.md
  29. 9 0
      docs/superpowers/specs/2026-08-08-rls-cls-design.md
  30. 5 1
      docs/superpowers/spikes/2026-05-15-storage-engine-spike.md
  31. 12 7
      service/src/database_service.cpp
  32. 10 4
      service/src/database_service.hpp
  33. 3 3
      service/src/main.cpp
  34. 1 1
      service/src/memory_store.cpp
  35. 1 1
      service/src/memory_store.hpp
  36. 1 1
      service/src/persistence/history_store.hpp
  37. 9 1
      service/src/storage/document_store_lmdb.cpp

+ 8 - 0
CLAUDE.md

@@ -1,5 +1,13 @@
 # Smartbotic Database
 
+> **Status and remaining work live in `docs/ROADMAP.md`** - read it before planning
+> anything. This file records per-release behaviour and hard-won lessons; it does
+> not track what is left to do.
+>
+> Docs hygiene: plans under `docs/superpowers/plans/` are pending work; anything
+> in `plans/archive/` has already shipped and must not be executed. Specs carry a
+> status banner. Resolved incidents are in `docs/incidents/`.
+
 ## Build
 
 ```bash

+ 4 - 0
README.md

@@ -2,6 +2,10 @@
 
 A standalone key-value document database with gRPC API, designed for microservice architectures.
 
+> **Current status and remaining work: [`docs/ROADMAP.md`](docs/ROADMAP.md).**
+> Per-release behaviour and the failure modes each release taught us: [`CLAUDE.md`](CLAUDE.md).
+> Consumer API guide: [`docs/integration-guide.md`](docs/integration-guide.md).
+
 ## Features
 
 - **Document Storage** - JSON document store with collections, backed by LMDB (2.0+)

+ 8 - 0
docs/2026-03-26-vector-storage-design.md

@@ -1,5 +1,13 @@
 # Vector Storage for smartbotic-database
 
+> **STATUS: implemented.** Vector storage with SIMD-accelerated cosine similarity
+> shipped in the v1.4 line and moved onto the LMDB substrate in v2.0
+> (`_vectors_<collection>` sub-dbs). One caveat learned later: a collection's
+> `vector_dimension` is immutable once it holds vectors, and collections created
+> through the pre-v2.4.5 `createCollection` namespacing bug were stranded with
+> dimension 0. v2.8.0 allows setting it while a collection holds no vectors.
+
+
 ## Goal
 Add embedding vector storage and similarity search to smartbotic-database, enabling semantic search without external dependencies (sqlite-vec, FAISS, etc.).
 

+ 13 - 0
docs/2026-05-15-v2.0-storage-engine-design.md

@@ -1,5 +1,18 @@
 # v2.0 — Page-based disk-backed storage engine
 
+> **STATUS: partly superseded. Read this before the rest of the document.**
+>
+> Phases A and B shipped as described (v1.10.0, v1.11.0). **Phase C did not.**
+> The "Phase C" section below specifies a homegrown engine - 16 KB slotted
+> pages, `BPlusTree`, `BufferPool`, page-level WAL redo. None of it was built.
+> The build-vs-buy spike chose LMDB instead, which shipped as v2.0.0 and
+> provides those internally. There is no `Page`, `BufferPool` or `BPlusTree`
+> type in this codebase, and `buffer_pool_size_mb` does not exist.
+>
+> Kept because the problem analysis and the phase reasoning are still sound.
+> For current status see `docs/ROADMAP.md`.
+
+
 > Date: 2026-05-15
 > Author: design draft
 > Status: Proposed

+ 161 - 0
docs/ROADMAP.md

@@ -0,0 +1,161 @@
+# Smartbotic Database - Status and Roadmap
+
+**Current version: 2.8.0** (see `VERSION`). Last reviewed: 2026-08-09.
+
+This is the single authoritative statement of what exists and what does not.
+If any other document in this repository disagrees with this one, this one is
+right and the other is stale - report it.
+
+Feature-level detail lives in `CLAUDE.md`. This file covers only *status*:
+shipped, pending, or abandoned.
+
+---
+
+## How to read the `docs/` tree
+
+| Path | What it is | Trust it? |
+|------|-----------|-----------|
+| `docs/ROADMAP.md` | This file - current status | Yes |
+| `CLAUDE.md` | Per-release feature notes and hard-won lessons | Yes |
+| `docs/integration-guide.md` | Consumer-facing API guide | Yes |
+| `docs/superpowers/specs/` | Design decisions, each with a status banner | Read the banner first |
+| `docs/superpowers/plans/` | Implementation plans for work **not yet done** | Yes, but re-validate against current code |
+| `docs/superpowers/plans/archive/` | Plans for work already shipped | **Historical only** - do not execute |
+| `docs/incidents/` | Post-mortems of resolved production incidents | Historical, lessons folded into `CLAUDE.md` |
+
+---
+
+## Storage engine history: Phases A, B, C
+
+Older documents refer to "Phase A/B/C". That vocabulary comes from the 2026-05-15
+storage-engine roadmap and is **retained only for reading old comments**. Do not
+plan new work in these terms.
+
+| Phase | Shipped | What actually happened |
+|-------|---------|------------------------|
+| A | v1.10.0 | yyjson became the *parser* on hot paths. `Document::data` stayed `nlohmann::json`. Done. |
+| B | v1.11.0 | Documents hold a binary `doc_binary::Doc` (yyjson `mut_doc`) with a lazy nlohmann view. Done. |
+| C | v2.0.0 | **Superseded in flight.** Phase C as designed meant writing our own engine: 16 KB slotted pages, a `BPlusTree`, an LRU buffer pool, page-level WAL redo - 2-3 months at high risk. The build-vs-buy spike chose LMDB instead (~2 weeks). Phase C's *goal* was met (bounded RSS, disk-backed B+ tree) but **none of its named components were built.** |
+
+Consequences of that substitution, which trip up readers:
+
+- **There is no `BufferPool`, no `Page`, no `BPlusTree` in this codebase.** LMDB
+  provides all three internally.
+- **`buffer_pool_size_mb` does not exist.** A comment in `database_service.cpp`
+  once promised `max_memory_mb` would be renamed to it "in v2.1". That never
+  happened and is not planned; `max_memory_mb` still configures MemoryStore
+  eviction, which no longer bounds RSS. LMDB mapsize plus OS page cache does.
+- "Phase C binary frame" is **not** a Phase C item. The relevant note is
+  `document_store_lmdb.cpp:9`, "Stage 4 may optimise to a binary frame later" -
+  part of the write-handler migration below.
+
+---
+
+## Shipped
+
+Every item below is in the installed product as of 2.8.0. `CLAUDE.md` has the
+detail and the failure modes.
+
+- JSON document store: collections, version history, field-level encryption, TTL
+- LMDB storage substrate, one env per project, dual-write mirror from MemoryStore
+- Multi-project namespaces (`<project>:<collection>`)
+- Vector storage with SIMD cosine similarity search
+- Views (read-only projections with baked-in filters)
+- Row- and column-level access policy, off by default per project
+- Project-scoped file storage with content-addressed, refcounted, deduplicated blobs
+- Per-listener TLS and bearer-token auth
+- Durable snapshots with tiered recovery and read-only lockout
+- Replication, events/subscribe, migrations, set operations
+- Paging fast path (no filter, no sort) and the two-pass filtered/sorted scan
+
+---
+
+## Pending
+
+Ordered by what a reader is most likely to need next. Nothing here has a
+committed date.
+
+### 1. Indexing (not designed yet)
+
+The largest remaining performance gap, and the natural next piece of work.
+
+A filtered query is a **full collection scan**. The v2.8.0 two-pass scan removed
+document materialisation (2463 ms → 402 ms on `executions`, 9800 docs / 505 MB),
+but the remaining 402 ms is yyjson parsing every value in the collection to test
+the predicate. That is the floor for a scan, and it is paid per query - so a
+consumer polling a filtered view costs a large fraction of a core continuously.
+
+Getting below it requires not reading non-matching rows at all: a secondary index
+mapping field value → doc id, maintained on write. Nothing exists yet - no design
+doc, no schema, no decision on which fields get indexed or whether declaration is
+explicit or automatic.
+
+### 2. `encode_document` still serialises with nlohmann
+
+`storage/document_store_lmdb.cpp` writes documents via `doc.toJson().dump()`.
+Measured, `nlohmann::dump()` is ~11x slower than `yyjson_write` on a large
+document (27.5 ms vs 2.5 ms on 2.91 MB). Swapping it to
+`doc_binary::to_json_text` is one function and no API change, but it alters the
+bytes every write produces, so it needs round-trip equivalence tests against an
+existing corpus before it can be trusted.
+
+### 3. Sub-db migration remainder (was "Stage 5")
+
+Three subsystems still persist outside LMDB:
+
+| Subsystem | Where it lives now | Target |
+|-----------|-------------------|--------|
+| History | `persistence/history_store.cpp`, `.hlog` files | `_history_<coll>` sub-db keyed `<doc_id>:<version>`; delete `history_store.cpp` |
+| File metadata | on-disk JSON at `<filesDir>/records/<2-char-prefix>/<id>.meta.json` | `_files` sub-db (blobs stay on disk) |
+| View definitions | MemoryStore `_views` system collection | `_views` sub-db keyed by qualified view name |
+
+Each is independently shippable. History is the largest.
+
+### 4. Write-handler migration (was "Stage 4-finish", once labelled the v2.4 milestone)
+
+Handlers would call `doc_store_->put()` directly, with id generation, timestamps
+and versioning moved out of MemoryStore into a `WriteCoordinator`. On completion
+MemoryStore, `persistence/wal.cpp`, `persistence/snapshot.cpp` and the eviction
+config knobs all delete, and the recovery-mode enum collapses.
+
+This is the v2.x arc's architectural endpoint. It touches the write path that
+produced the v2.4.3, v2.4.4 and v2.8.0 LMDB handle incidents, so it wants to land
+in small reviewable pieces with the sub-db identity sentinel kept intact.
+
+### 5. Relations
+
+`docs/superpowers/plans/2026-08-04-relations-v2.5.0.md` is a complete 3278-line
+plan that was **never executed**. Its version label is wrong - v2.5.0 was skipped
+entirely (2.4.5 → 2.6.0). It was written before the v2.8.0 dbi-caching fix, so its
+write-path assumptions need re-validating before use. Treat it as a design input,
+not a script.
+
+### 6. Smaller known gaps
+
+- Policy management has no dedicated RPCs; `_policies` is edited through the
+  ordinary document API with server-side interception.
+- Service-wide operations (`GetStats`, `SetReadOnly`, `Create/DropProject`) can
+  only be gated coarsely - they have no project to evaluate a policy against.
+  Listener separation is the real boundary there.
+- `event.data->dump()` and filter-echo `f.value.dump()` in
+  `database_grpc_impl.cpp` are the last nlohmann serialisers on warm paths.
+- `_orphans_*` sub-dbs from the v2.4.4 placement repair still await manual
+  comparison.
+- TTL is not retro-applied: rows written before a default TTL existed keep no
+  expiry.
+
+---
+
+## Not planned
+
+- **v3.** v2.x is the terminal arc.
+- **Removing `nlohmann::json`.** It is the public client API vocabulary type -
+  `insert`, `get`, `Filter::value`, `QueryResult::documents`. Removing it breaks
+  every consumer call site and needs a soname bump. It is header-only, so it is
+  not a runtime dependency of any shipped package and costs nothing at runtime.
+  The remaining ~453 internal references are cold paths (config, migrations,
+  policies, views) where its exceptions and value semantics are an asset.
+- **A homegrown page engine or buffer pool.** Settled: LMDB.
+- **v1.x auto-migration as a separate package.** Was slated to extract to
+  `smartbotic-database-migrate-v1` in v2.5; v2.5 never shipped and the in-binary
+  path is harmless.

+ 0 - 0
BUG-snapshot-truncation.md → docs/incidents/2026-04-19-snapshot-truncation.md


+ 0 - 0
plan.md → docs/incidents/2026-04-22-eviction-induced-write-loss-plan.md


+ 0 - 0
BUG-memory-leak-zoe.md → docs/incidents/2026-04-22-zoe-rss-exceeds-budget.md


+ 13 - 0
docs/incidents/README.md

@@ -0,0 +1,13 @@
+# Production incident records
+
+**All incidents here are resolved.** They were previously top-level `BUG-*.md`
+files whose status headers still read "Open", which is no longer true. The
+lessons are folded into `CLAUDE.md`; these are the long-form originals.
+
+| File | Resolution |
+|------|-----------|
+| `2026-04-19-snapshot-truncation.md` | **Fixed in v1.6.1.** The snapshot writer became atomic - write `.tmp`, fsync, rename, then verify by re-reading. A loader fallback chain and the MySQL-style recovery modes landed with it. |
+| `2026-04-22-zoe-rss-exceeds-budget.md` | **Superseded by v2.0.** The original "memory leak" framing was wrong: v2.4.3 established that `estimatedMemoryBytes_` under-reports by roughly 5x, so eviction never believed it had made progress. Since v2.0 RSS is bounded by the LMDB mapsize plus OS page cache, not by MemoryStore eviction, and `max_memory_mb` only sizes the mirror cache. |
+| `2026-04-22-eviction-induced-write-loss-plan.md` | **Shipped in v1.7.x.** Became chunked pressure-aware eviction (four pressure levels, hot-write floor, quiesce on in-flight writes, no per-collection mutex held during disk I/O) plus the `conf.d` drop-in config mechanism. v2.4.3 later added the episode cap and no-progress detector after eviction drained a live store. |
+
+Anything still open lives in `docs/ROADMAP.md` under "Pending", not here.

+ 0 - 118
docs/superpowers/plans/2026-05-15-v2.0-roadmap.md

@@ -1,118 +0,0 @@
-# v2.0 Storage Engine Rewrite — Roadmap
-
-> **Status:** Roadmap stub. Phase A and beyond require their own detailed plan files before subagent dispatch. Do NOT try to execute this file directly — it is an index.
-
-**Source of design:** [`/data/smartbotic-database/docs/2026-05-15-v2.0-storage-engine-design.md`](../../2026-05-15-v2.0-storage-engine-design.md) (316 lines, already merged into main as the canonical design).
-
-**Why this exists:** The operator asked to "proceed with them all" — all three phases of the v2.0 design. Phases A/B/C span ~4–5 months of engineering and cannot execute in a single session. This file:
-1. Lists the order in which the phases ship
-2. Captures cross-cutting upgrade-safety requirements that span every phase
-3. Identifies where the WAL/snapshot format breaks (Phase C, not A/B)
-4. Points each phase at the plan file it needs before subagents can be dispatched
-
----
-
-## Cross-cutting upgrade-safety requirements (operator-mandated)
-
-These apply to **every** phase, not just Phase C:
-
-1. **Backup before upgrade** — already shipping in v1.9.5 (`2026-05-15-v1.9.5-backup-before-upgrade.md`). Every subsequent release inherits this safety net.
-2. **WAL/snapshot backward compatibility** — each release must boot cleanly on the previous release's data dir. Concretely:
-   - v1.10 (Phase A) must read v1.9.x WAL + snapshot bytes unchanged (yyjson swap is only the parser; on-disk format is identical).
-   - v1.11 (Phase B) introduces binary doc storage in memory but keeps the WAL/snapshot doc payload as JSON text on disk so v1.10 readers can still parse v1.11 snapshots. (Migration to a binary on-disk format is deferred to Phase C.)
-   - v2.0 (Phase C) is the format break: page-based on-disk layout, WAL becomes page-level redo. A one-shot `--migrate-from-v1` mode is required (see Phase C plan when written).
-3. **Pin-back path** — `apt install smartbotic-database=<old-version>` plus the v1.9.5 backup restore procedure must remain a working rollback for at least one release back.
-
----
-
-## Phase A → v1.10.0 — yyjson hot-path swap
-
-**Status:** Plan not yet written. **Blocked on:** writing `2026-05-15-v1.10-yyjson-phase-a.md`.
-
-**Goal:** Replace nlohmann::json *parsing* with yyjson at hot paths (WAL replay, snapshot deserialize, gRPC request bodies, history reads). Keep nlohmann::json as the in-memory document representation — Phase A is a parse-speed win, not a memory win.
-
-**Honest impact assessment:** With jemalloc already shipped in v1.9.4, RSS is bounded. Phase A buys faster boot (parse-heavy paths) and slightly lower gRPC request latency on large bodies, but does **not** reduce steady-state memory. The operator should know this before Phase A is scheduled — if memory is the priority, Phase B is the work that matters.
-
-**Files in scope (parse sites identified, see `grep` output in the design doc):**
-- `service/src/persistence/wal.cpp:233,249` — replay
-- `service/src/persistence/snapshot.cpp:820,929,963` — deserialize
-- `service/src/persistence/history_store.cpp:124` — version read
-- `service/src/database_grpc_impl.cpp` — Insert/Update/Patch/Find handlers (~10 call sites)
-- `service/src/database_service.cpp:722` — replication apply
-- `packaging/Dockerfile.base` — add `libyyjson-dev`
-- `packaging/deb/templates/control.server` — add `libyyjson0` runtime dep
-- `service/CMakeLists.txt` — `pkg_check_modules(yyjson REQUIRED)` or `find_package(yyjson)`
-
-**Task breakdown sketch (to be expanded in the Phase A plan file):**
-1. Add yyjson build + runtime deps, bump Dockerfile.base tag
-2. Helper: `yyjson_to_nlohmann(const yyjson_val*) -> nlohmann::json` — the bridge
-3. Swap WAL replay parse → re-run replica eviction load test
-4. Swap snapshot deserialize parse → re-run snapshot durability test
-5. Swap gRPC handler parses → re-run integration tests
-6. Bench: report parse-time delta on a Zoe-shape snapshot
-7. Bump VERSION to 1.10.0, build deb, sync
-
-**Time estimate:** 1–1.5 weeks of subagent-driven work across 2–3 sessions.
-
----
-
-## Phase B → v1.11.0 — binary lazy document storage
-
-**Status:** Plan not yet written. **Blocked on:** writing `2026-05-15-v1.11-binary-docs-phase-b.md` *and* settling the format choice.
-
-**Goal:** Document in-memory representation changes from heap-allocated nlohmann::json AST → `std::vector<uint8_t> binary` with lazy `.data()` accessor that parses on demand. This is the phase that actually reduces memory.
-
-**Format decision (open):** BSON vs custom packed-JSON tape vs yyjson mut_doc-as-storage. Each has trade-offs:
-- **BSON** — well-specified, has libbson; but spec includes types we don't use (Date, ObjectId, Decimal128) and field tagging adds overhead.
-- **Custom tape** — minimal, fast, but we own the spec forever.
-- **yyjson mut_doc** — keep the parser's own buffer as storage; clean but ties us to yyjson semantics.
-
-**Pre-execution work:** Operator should brainstorm format choice (`superpowers:brainstorming` skill) before this plan can be written.
-
-**Files in scope:** `service/src/document.hpp`, `service/src/memory_store.cpp/.hpp`, every callsite that does `doc.data["field"]` (large blast radius — ~80+ call sites across service + tests).
-
-**Wire compatibility:** gRPC `Document.data` stays JSON text on the wire. Conversion happens at the boundary.
-
-**Snapshot compat:** Snapshot format v6 emits documents as JSON text (same as v5) so v1.10 readers still work. Phase C is where the on-disk format breaks.
-
-**Time estimate:** 3–4 weeks across 5–8 sessions.
-
----
-
-## Phase C → v2.0 — page-based on-disk storage + buffer pool
-
-**Status:** Plan not yet written. **Blocked on:** the build-vs-buy spike (1 week) AND a brainstorm session on the architecture AND writing `2026-05-15-v2.0-storage-engine-phase-c.md`.
-
-**Goal:** Replace the "load everything into memory" architecture with bounded-memory disk-backed page storage. This is the real "act like MySQL/InnoDB" rewrite.
-
-**The build-vs-buy spike (must run first):**
-- Option 1: Homegrown 16 KB slotted pages + B+ tree index + LRU buffer pool + page-level WAL redo (~3 months engineering)
-- Option 2: RocksDB-backed prototype — let LSM do the page management, keep our document/vector/index semantics on top (~3 weeks engineering, ongoing operational dep on RocksDB)
-- Option 3: LMDB-backed — simpler, single-writer, mmap'd B+ tree (~2 weeks engineering, but write-concurrency limits)
-
-The spike picks one. Without that decision, the Phase C plan cannot be written.
-
-**Required cross-cutting work in Phase C:**
-- One-shot `--migrate-from-v1` boot mode: read v1.x snapshot + replay v1.x WAL into pages, then write a v2.0 checkpoint and start serving
-- New WAL format (page-level redo records), new snapshot format (page checkpoints), new on-disk layout under `/var/lib/smartbotic-database/pages/`
-- Eviction machinery from v1.7.0 removed — page LRU subsumes it; `max_memory_mb` → `buffer_pool_size_mb`; evicted stubs / WAL fallback / `docWalSeq_` all deleted
-
-**Time estimate:** ~13 weeks across many sessions, contingent on spike outcome.
-
----
-
-## What to do next session
-
-1. Confirm v1.9.5 backup mechanism is shipped (current session's deliverable).
-2. **If continuing into Phase A:** write `2026-05-15-v1.10-yyjson-phase-a.md` with full task breakdown (use `superpowers:writing-plans`). Then dispatch implementers per `superpowers:subagent-driven-development`.
-3. **If considering reordering:** the honest recommendation is to **skip Phase A and go directly to Phase B** *if* memory is the priority. Phase A is foundation work whose primary benefit (parser swap) is also achievable as a Phase B side-effect (yyjson would naturally back the binary doc storage in B). Discuss with operator.
-
----
-
-## What NOT to do without operator approval
-
-- **Do not start Phase A** without re-confirming memory vs latency priorities with the operator.
-- **Do not start Phase B** without settling the binary format choice (BSON vs tape vs yyjson mut_doc).
-- **Do not start Phase C** without running and writing up the build-vs-buy spike.
-- **Do not break WAL/snapshot backward compat outside of Phase C.**
-- **Do not skip the v1.9.5 backup mechanism** — it is the rollback safety net for everything that follows.

+ 0 - 176
docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c.md

@@ -1,176 +0,0 @@
-# Page-Based Storage Engine (Phase C → v2.0) Implementation Plan — SKELETON
-
-> **For agentic workers:** This plan is INTENTIONALLY INCOMPLETE. Only Tasks 1 and 2 are executable now. Tasks 3+ depend on the spike outcome from Task 1 and the architecture brainstorm in Task 2, and CANNOT be specified without those decisions. Do NOT attempt to write or execute Tasks 3+ until Task 2 produces a follow-up plan file.
->
-> When Task 2 completes, it MUST author `docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md` with the full TDD-style task breakdown. That follow-up plan is what subagent-driven-development will execute.
-
-**Goal:** Replace the "load everything into memory" v1.x architecture with a bounded-memory disk-backed page storage engine. RSS is governed by `buffer_pool_size_mb`, not by total dataset size. This is the "act like MySQL/InnoDB" rewrite that lets Smartbotic Database run a dataset much larger than RAM.
-
-**Architecture:** TBD by spike. Three candidate architectures are pre-identified in the v2.0 design doc (`docs/2026-05-15-v2.0-storage-engine-design.md`):
-
-1. **Homegrown** — 16 KB slotted pages, B+ tree primary index, LRU buffer pool, page-level WAL redo. We own everything. ~3 months engineering. No new runtime deps.
-2. **RocksDB-backed** — LSM-tree handles page management; we keep document/vector/index semantics on top of RocksDB column families. ~3 weeks engineering. Ongoing operational dep on RocksDB.
-3. **LMDB-backed** — mmap'd B+ tree, single-writer concurrency. ~2 weeks engineering. Simpler than RocksDB but with write-concurrency limits that may matter at Zoe's write rate.
-
-The spike (Task 1) picks one. Without that decision, the implementation plan cannot be written.
-
-**Tech Stack:** TBD by spike. The non-negotiables — C++20, gRPC/Protobuf wire format unchanged, nlohmann::json kept as the document AST surface, yyjson (from Phase A) as the parser — survive any choice.
-
----
-
-## Context
-
-**What changes in v2.0:**
-
-- **On-disk layout:** `/var/lib/smartbotic-database/pages/` (or backend-equivalent — RocksDB's `*.sst`, LMDB's `data.mdb`). The current `snapshots/`, `wal/`, `history/` tree from v1.x is gone after one-shot migration.
-- **WAL format:** page-level redo records replace document-level append. Format break — v1.x WAL is unreadable by v2.0 directly; the migration tool reads it.
-- **Snapshot format:** page checkpoints replace whole-store LZ4 dumps. Another format break.
-- **Eviction:** the v1.7.0 chunked eviction machinery (`hot_write_floor_ms`, `eviction_chunk_size`, four pressure levels, evicted-stub WAL fallback, `docWalSeq_`) is GONE. The buffer pool's LRU replacement is the only eviction mechanism. `max_memory_mb` → `buffer_pool_size_mb`.
-- **Recovery:** the v1.7.0 tiered recovery (snapshot → snapshot_fallback → WAL-only → best_effort → force_empty) collapses to standard checkpoint-and-redo. `recovery.mode` config keeps the same surface for operator continuity but maps to the new engine's terms.
-
-**Operator constraint (load-bearing):** "I will not upgrade services which use the database just when the new v2 finished." This means v2.0 is a single-jump release from v1.9.5 (or whatever the latest v1.x is when the spike completes). The `--migrate-from-v1` one-shot mode is REQUIRED — the operator needs `apt upgrade smartbotic-database` to:
-
-1. Stop the service (v1.9.5 preinst takes the auto-backup).
-2. Install v2.0.
-3. On first boot, detect a v1.x data dir and run the one-shot migration into the new engine's format.
-4. Start serving.
-
-If migration fails, the operator restores from the v1.9.5 backup and reinstalls v1.x — that's the rollback path. v1.9.5's backup-before-upgrade is the load-bearing safety net here; treat it as a hard requirement.
-
-**Cross-cutting changes in v2.0 (from the design doc):**
-- New `service/src/storage/` directory tree replaces `service/src/persistence/`.
-- `MemoryStore` is renamed (e.g., `DocumentStore` or `PageStore`) and gains a buffer pool.
-- All replication paths re-routed through the new WAL.
-- Vector storage (currently parallel float arrays alongside docs) needs a v2.0 home — page-based or alongside the index. Spike must address this.
-- Field-level encryption stays at the AES-256-GCM layer; only the storage substrate changes.
-
----
-
-## Tasks
-
-### Task 1 — Build-vs-buy spike (1 week)
-
-**Files:**
-- Create: `docs/superpowers/spikes/2026-05-15-storage-engine-spike.md` — the spike writeup.
-- Create: `service/spike/` — throwaway prototype code, gitignored or committed on a `spike/v2-storage` branch (operator's call).
-
-The spike is research, not production code. Its only deliverable is the writeup that justifies the architecture choice.
-
-- [ ] **Step 1: Define the spike scope.** Prototype the *minimum* path that exercises each candidate's failure modes:
-  - Load 1M Zoe-shape docs into the backend.
-  - Run a sustained 5k write/s + 50k read/s workload.
-  - Crash the process and verify recovery time + data integrity.
-  - Run a SimilaritySearch over a vector subset.
-
-- [ ] **Step 2: Prototype Option 3 first (LMDB)** — smallest scope, fastest to falsify or validate. Measure: write throughput ceiling under single-writer constraint, mmap RSS behaviour, recovery time.
-
-- [ ] **Step 3: Prototype Option 2 (RocksDB)** — if Option 3's write-concurrency ceiling is below Zoe's projected sustained write rate, RocksDB is the alternative. Measure: compaction overhead, RSS plateau, recovery time, ops experience (compaction tuning, blob storage for large docs).
-
-- [ ] **Step 4: Prototype Option 1 (homegrown) only if both 2 and 3 fail to meet requirements.** Document the requirement gap that justifies a 3-month build over a 2-3 week buy. Be honest — the homegrown option is the heaviest commitment and should require strong evidence to choose.
-
-- [ ] **Step 5: Write the spike report.** Required sections:
-  - **Scope** — what was measured, what wasn't.
-  - **Results** — numbers per candidate: write tps, read tps, recovery time, RSS curve, on-disk size, vector search latency.
-  - **Operational cost** — what's the ops burden of each choice over 5 years (compaction tuning for RocksDB, single-writer for LMDB, total ownership for homegrown).
-  - **Decision** — one of {homegrown, RocksDB, LMDB} with the reasoning anchored in numbers.
-  - **Rejected alternatives** — what we ruled out and why.
-  - **Risks** — the top 3 things that could derail the chosen path.
-  - **Migration sketch** — high-level mapping of v1.x data → v2.0 backend (the detail goes in Task 2's plan).
-
-- [ ] **Step 6: Operator review gate.** The spike report needs operator sign-off before Task 2. Per the brainstorming skill's review-gate pattern.
-
-- [ ] **Step 7: Commit the spike report** (keep the prototype code on its own branch).
-
-```bash
-git checkout main
-git add docs/superpowers/spikes/2026-05-15-storage-engine-spike.md
-git commit -m "spike: v2.0 storage engine build-vs-buy ($DECISION)"
-```
-
----
-
-### Task 2 — Architecture brainstorm + execution plan authoring
-
-**Files:**
-- Create: `docs/superpowers/specs/2026-05-15-v2.0-storage-engine-architecture.md` — the architecture decision record.
-- Create: `docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md` — the full execution plan.
-
-Use `superpowers:brainstorming` to drive Step 1; use `superpowers:writing-plans` for Step 4.
-
-- [ ] **Step 1: Run a brainstorming session** scoped to the v2.0 architecture given the spike outcome. Key questions to drive the session:
-  - Buffer pool size policy — fixed at config, adaptive to system memory, both?
-  - Index strategy — what indexes ship in v2.0 (primary key + per-collection BTrees + vector index)? Which are deferred to v2.1?
-  - Migration UX — operator runs it explicitly, or auto-runs on first boot detecting a v1.x data dir? What does failure look like?
-  - Replication — does the v1.x WAL-streaming replication model survive, or does v2.0 replicate at the page level (Postgres physical replication style)? This is a big architecture call.
-  - Vector storage — co-located with docs in pages, or in a separate index?
-  - Schema/index migrations — v1.x's JSON migrations RPC needs a v2.0 equivalent.
-
-- [ ] **Step 2: Author the architecture decision record** with the brainstorm outcome.
-
-- [ ] **Step 3: Operator review gate** on the architecture doc.
-
-- [ ] **Step 4: Author the v2.0 execution plan** at `docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md`. This plan is full-detail (every TDD step, every file path, every command) at the same level as the Phase A plan in this same directory. Required sections:
-  - File-structure decomposition (what gets created in `service/src/storage/`, what gets deleted from `service/src/persistence/`, etc.)
-  - Migration tool task breakdown
-  - WAL format spec
-  - Snapshot/checkpoint format spec
-  - Buffer pool implementation tasks
-  - Index re-implementation tasks (primary index, vector index)
-  - Replication migration tasks
-  - Eviction-machinery deletion tasks (remove `hot_write_floor_ms`, evicted-stub WAL fallback, `docWalSeq_`, four-level pressure events from v1.7.0)
-  - Performance / soak gates
-  - VERSION bump to 2.0.0
-  - Operator-facing release notes
-
-- [ ] **Step 5: Operator review** of the execution plan.
-
-- [ ] **Step 6: Commit both docs.**
-
-```bash
-git add docs/superpowers/specs/2026-05-15-v2.0-storage-engine-architecture.md docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md
-git commit -m "spec+plan: v2.0 storage engine architecture and execution plan"
-```
-
----
-
-### Tasks 3+ — Execution
-
-**File:** see `docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md` (authored by Task 2).
-
-The detailed implementation tasks live in the execution plan. They CANNOT be written here because:
-
-1. The set of files to create/modify depends on the spike outcome (homegrown vs RocksDB vs LMDB).
-2. The TDD test design depends on the chosen storage primitives.
-3. The migration tool's structure depends on the source-to-target schema mapping decided in Task 2.
-
-Writing Tasks 3+ here without those decisions would produce exactly the kind of placeholder-laden plan that `writing-plans` warns against. Tasks 1 and 2 exist precisely to retire that uncertainty.
-
----
-
-## Out of scope (Phase C)
-
-- **Distributed scaling.** v2.0 is single-node + replication, same as v1.x. Sharding / Raft / multi-leader is post-v2.0.
-- **Online schema migration without rewriting docs.** v2.0 keeps the JSON-document-with-optional-schema model.
-- **Query language change.** Filters, views, vector search keep the v1.x gRPC API surface.
-- **Anything not flowing from the spike outcome.** No speculative work.
-
----
-
-## Success criteria
-
-1. The spike report exists, is reviewed, and names a chosen architecture with numerical justification.
-2. The architecture decision record exists and is reviewed.
-3. The execution plan exists, passes `writing-plans` self-review (no placeholders, no TBDs, every step has code/commands), and the operator approves it.
-4. Subagent-driven-development can execute the execution plan task-by-task with no ambiguity left in the plan.
-5. v2.0 ships with a working `--migrate-from-v1` mode that boots cleanly against a v1.9.x data dir (Zoe-shape, full size).
-6. v1.9.5 backup-before-upgrade still fires on v1.x → v2.0 upgrade — confirmed manually in the soak test.
-
----
-
-## Open questions (resolved by Task 1 + Task 2)
-
-- Build vs buy (Task 1).
-- Buffer pool sizing policy (Task 2).
-- Replication model — WAL streaming vs page-level (Task 2).
-- Vector index placement (Task 2).
-- Migration UX (Task 2).

+ 0 - 107
docs/superpowers/plans/2026-05-16-v2.1-backlog.md

@@ -1,107 +0,0 @@
-# v2.x backlog — shadowman audit follow-up + Stage 5-10 remainder
-
-Origin: shadowman team alignment audit `/tmp/sbdbv2.md` (2026-05-15) +
-the unfinished items from `2026-05-15-v2.0-storage-engine-phase-c-execution.md`.
-
-**v2.x is the project's terminal arc — no v3 is planned.** Every item
-below targets a v2.x.x slot. Items that don't fit the v2.x roadmap are
-simply not planned (no v3 backlog to escalate them to).
-
-This document supersedes the loose "Out of scope (deferred to v2.1+)"
-section in `2026-05-15-v2.0-roadmap.md` (lines 351-359). Each entry has
-an explicit drop trigger so the calendar of removals is observable.
-
-## Status of shadowman audit items
-
-### F-section (open questions) — answered
-
-| F# | Question | Answer (committed v2.1.1) |
-|----|----------|---------------------------|
-| F1 | When does the v1.x auto-migration path drop? | Stays in-binary through v2.4. Extracts to separate `smartbotic-database-migrate-v1` deb in v2.5 once every production host has booted v2.4 ≥ once. v2.x is the terminal arc — no v3 escalation. |
-| F2 | Is `setReadOnly` operator-only or client-callable? | Both. Operator via `smartbotic-db-cli unlock`; client via `Client::lock()` for self-defensive migrations. |
-| F3 | WAL-fallback fields — zero on healthy v2.0? | Yes, guaranteed zero on the LMDB-served path (Find handler explicitly leaves them at their default-constructed zero). Non-zero = real signal, treat as alert. |
-| F4 | `migrateCollectionTimestamps` cost on hot `messages`? | Deferred bench until shadowman runs D.1. Back-of-envelope: ~1ms / 1k docs, online-safe up to ~1M docs; >1M docs schedule a maintenance window. Will be re-measured against real traffic post-D.1. |
-
-### B-section (back-compat surfaces) — drop tracker
-
-| B# | Surface | Drop trigger | Target |
-|----|---------|--------------|--------|
-| B1 | In-binary v1.x → v2.0 migration (commit `15893b1`) | Every production host has booted v2.4 ≥ 1× | v2.5 (extract to separate deb) |
-| B2 | Defensive MemoryStore fallback on LMDB-throw | Shipped — closed v2.1.1 | ✅ done |
-| B3 | Stripped `QueryOptions` in document.hpp | Shipped — closed v2.1.0 | ✅ done |
-| B4 | Plain `find(collection)` overload | Shipped — closed v2.1.0 | ✅ done |
-| B5 | Default `"ms"` timestamp precision | Every rapid-write collection migrated via shadowman D.1 | v2.2 (flip default to `"ns"`) |
-| B6 | `update()` as RMW verb | Every RMW callsite converted to `patch()` via shadowman D.2 | v2.x.next (mark `update()` as "replace whole doc only" in docs) |
-| B7 | 6-value `recoveryOutcome` enum | Write-handler migration ships → MemoryStore decommission → LMDB-only recovery semantics | v2.4 (collapse to single `lmdb_ready`, paired with the write-flip) |
-
-### Shadowman D-section (their migration items)
-
-External dependency — not tracked here. The cookbook from their audit:
-
-* D.1 ns-precision timestamps on `messages` / `metrics_events` / `tool_call_blobs` / `agent_task_messages` — **highest leverage**, unblocks their T6.6 timeline refactor
-* D.2 `patch()` adoption for 71 RMW callsites
-* D.3 `count(filters)` for "any-matching-row" patterns (~10-15 sites)
-* D.4 `findWithMetrics()` for slow-query logging
-* D.5 Set operations for `conversation_tool_state`
-* D.6 Server-side FilterOp on cold paths (their T7.1-T7.3)
-* D.7 Diagnostic `createView()` (their T7.4)
-
-## Unfinished items from the v2.0 Phase C plan
-
-These were architectural endpoints not reached in v2.0:
-
-### Stage 5 — sub-db migration remainder
-
-| Sub | Module | Current state in v2.1 | v2.x target |
-|-----|--------|----------------------|-------------|
-| 5.2 | History `_history_<coll>` | Still on `service/src/persistence/history_store.cpp` .hlog files | LMDB sub-db keyed by `<doc_id>:<version>` storing DocumentVersion JSON. Delete `history_store.cpp`. |
-| 5.3 | Files `_files` | Metadata still in MemoryStore + .file blob path | `_files` sub-db for metadata; blobs unchanged at `files/<id>`. |
-| 5.4 | Views `_views` | View defs stored in MemoryStore `_views` system collection | `_views` LMDB sub-db keyed by view name. Existing `createView` / `getViewInfo` handlers ported. |
-
-Estimated effort: 5.2 medium (HistoryStore deletion + test port), 5.3 + 5.4 small. Each is one focused PR.
-
-### Stage 4-finish + Stage 6+ — write-handler migration & beyond
-
-This is the **v2.4** milestone — the v2.x arc's architectural endpoint. Handlers call `doc_store_->put()` directly with id-gen / timestamp / version orchestration extracted from MemoryStore into a WriteCoordinator. Once landed, MemoryStore deletes; persistence/wal.cpp / persistence/snapshot.cpp delete; eviction config knobs delete; recovery enum collapses (B7).
-
-Subsequent stages from the original plan, mostly unchanged:
-
-* Stage 6 — replication preservation: validate WAL streaming still works through the write-handler-migration boundary
-* Stage 7 — migrations RPC: port `MigrationRunner` to call `DocumentStore`
-* Stage 8 — eviction + recovery cleanup audit
-* Stage 9 — performance gates + 24h soak (NOT done for v2.0 — `load_test_mixed` 30s was the proxy)
-* Stage 10 — v2.4 release (the v2.x arc's architectural endpoint)
-
-## Not planned
-
-The v2.0 roadmap's "Out of scope (deferred to v2.1+)" list contained
-items that, with no v3 in the plan, simply aren't scheduled:
-
-* Secondary user-defined indexes
-* Online schema evolution
-* Distributed clustering beyond leader+follower
-* LMDB writer-pool sharding
-* Adaptive `buffer_pool_size_mb` auto-tune
-* Vector index variants (IVF, HNSW)
-
-If any of these become required, they need their own roadmap document
-+ explicit operator decision to extend the v2.x arc. None of them is
-blocking the v2.x terminal milestone (v2.4 write-handler migration +
-MemoryStore decommission).
-
-## Sequencing
-
-The earliest valuable sub-PR sequence after v2.1.1:
-
-1. **5.4 views** (smallest, demonstrates the `_views` sub-db pattern)
-2. **5.3 files** (similar shape; metadata-only move)
-3. **Shadowman D.1** (their migration, unblocks B5 in v2.2)
-4. **5.2 history** (bigger refactor; touches history_store.cpp + test_snapshot_durability)
-5. **B5 flip default to ns** — v2.2 release
-6. Then write-handler migration / Stage 4-finish — **v2.4** (the v2.x terminal architectural milestone)
-
-## Open questions
-
-* Is there a consumer beyond shadowman that we need to coordinate with on B5's default flip? (callerai uses ms today.)
-* Does the `smartbotic-database-migrate-v1` deb (per F1) need to support v1.x → any v2.x.x, or only v1.x → v2.5 (its target)? Affects how much v1.x code we carry inside the migration deb.
-* For Stage 9, do we wait for a real Zoe-shape soak before shipping v2.4, or is the `load_test_mixed` 30s proxy acceptable as the gate to MemoryStore decommission?

+ 13 - 0
docs/superpowers/plans/2026-08-04-relations-v2.5.0.md

@@ -1,5 +1,18 @@
 # Relations (v2.5.0) Implementation Plan
 
+> **STATUS: NEVER EXECUTED. Do not run this plan as written.**
+>
+> - Nothing in it is implemented - no `_relations`, no `RelationManager`, no RPC.
+> - **The version target is wrong.** v2.5.0 was skipped entirely; the project went
+>   2.4.5 → 2.6.0 and is now at 2.8.0. Renumber before use.
+> - It was written before the v2.8.0 discovery that an `MDB_dbi` must not be
+>   cached until its transaction commits. Its write-path tasks touch exactly that
+>   code and need re-validating against `document_store_lmdb.cpp` as it stands.
+> - Treat it as a design input, not a script.
+>
+> See `docs/ROADMAP.md` for where relations sits in the real ordering.
+
+
 > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
 
 **Goal:** Add referential integrity to smartbotic-database - per-relation `on_delete` policies over a document model, with a reverse index, a dry-run API, and atomic cascade.

+ 0 - 0
docs/plans/2026-03-26-vector-storage.md → docs/superpowers/plans/archive/2026-03-26-vector-storage.md


+ 0 - 0
docs/superpowers/plans/2026-04-12-packaging-consolidation.md → docs/superpowers/plans/archive/2026-04-12-packaging-consolidation.md


+ 0 - 0
docs/superpowers/plans/2026-04-13-views-v1.5.0.md → docs/superpowers/plans/archive/2026-04-13-views-v1.5.0.md


+ 1 - 1
docs/superpowers/plans/2026-04-19-snapshot-recovery-v1.6.1.md → docs/superpowers/plans/archive/2026-04-19-snapshot-recovery-v1.6.1.md

@@ -10,7 +10,7 @@
 
 ---
 
-## Context — the bug (see `BUG-snapshot-truncation.md`)
+## Context — the bug (see `docs/incidents/2026-04-19-snapshot-truncation.md`)
 
 **Writer issues:**
 1. `file.write()` calls never check failbit — silent short writes pass through

+ 0 - 0
docs/superpowers/plans/2026-05-15-v1.10-yyjson-phase-a.md → docs/superpowers/plans/archive/2026-05-15-v1.10-yyjson-phase-a.md


+ 0 - 0
docs/superpowers/plans/2026-05-15-v1.11-binary-docs-phase-b.md → docs/superpowers/plans/archive/2026-05-15-v1.11-binary-docs-phase-b.md


+ 0 - 0
docs/superpowers/plans/2026-05-15-v1.9.5-backup-before-upgrade.md → docs/superpowers/plans/archive/2026-05-15-v1.9.5-backup-before-upgrade.md


+ 0 - 0
docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md → docs/superpowers/plans/archive/2026-05-15-v2.0-storage-engine-phase-c-execution.md


+ 0 - 0
docs/superpowers/plans/2026-05-16-v2.3-multi-project.md → docs/superpowers/plans/archive/2026-05-16-v2.3-multi-project.md


+ 0 - 0
docs/superpowers/plans/2026-05-31-v2.4-listeners-tls-auth.md → docs/superpowers/plans/archive/2026-05-31-v2.4-listeners-tls-auth.md


+ 0 - 0
docs/superpowers/plans/2026-08-08-project-scoped-files-v2.6.0.md → docs/superpowers/plans/archive/2026-08-08-project-scoped-files-v2.6.0.md


+ 30 - 0
docs/superpowers/plans/archive/README.md

@@ -0,0 +1,30 @@
+# Archived implementation plans
+
+**Every plan in this directory describes work that has already shipped.**
+
+They are kept as a record of how each feature was built and what was decided
+along the way. They are **not** current:
+
+- Their "next steps", version targets and open questions are all resolved or
+  abandoned.
+- Code they reference may have been renamed, moved or deleted since.
+- Several were written against architecture that has since been replaced -
+  most notably the pre-v2.0 MemoryStore-only storage model.
+
+**Do not execute a plan from this directory.** For current status and remaining
+work see `docs/ROADMAP.md`. For per-release behaviour and the failure modes each
+release taught us, see `CLAUDE.md`.
+
+| Plan | Shipped in |
+|------|-----------|
+| `2026-03-26-vector-storage.md` | v1.4.x - vector storage + SIMD similarity search |
+| `2026-04-12-packaging-consolidation.md` | v1.2.0 - the four-package Debian layout |
+| `2026-04-13-views-v1.5.0.md` | v1.5.0 - views |
+| `2026-04-19-snapshot-recovery-v1.6.1.md` | v1.6.1 - durable snapshots + tiered recovery |
+| `2026-05-15-v1.9.5-backup-before-upgrade.md` | v1.9.5 - auto-backup preinst hook |
+| `2026-05-15-v1.10-yyjson-phase-a.md` | v1.10.0 - yyjson hot-path parsing |
+| `2026-05-15-v1.11-binary-docs-phase-b.md` | v1.11.0 - binary-backed documents |
+| `2026-05-15-v2.0-storage-engine-phase-c-execution.md` | v2.0.0 - LMDB substrate |
+| `2026-05-16-v2.3-multi-project.md` | v2.3.0 - multi-project namespaces |
+| `2026-05-31-v2.4-listeners-tls-auth.md` | v2.4.0 - per-listener TLS + bearer auth |
+| `2026-08-08-project-scoped-files-v2.6.0.md` | v2.6.0 - project-scoped file storage |

+ 6 - 0
docs/superpowers/specs/2026-05-15-binary-doc-format-decision.md

@@ -1,5 +1,11 @@
 # Binary Document Format Decision — v1.11 Phase B
 
+> **STATUS: implemented, shipped as v1.11.0.** The post-implementation
+> amendment near the end of this document is the honest accounting: per-document
+> heap footprint improved only ~1.5%, not the large RSS win originally
+> projected. The real wins were field-access CPU and the accessor surface.
+
+
 **Status:** Decided. Operator approved 2026-05-15.
 
 **Choice:** **yyjson `mut_doc`** as the in-memory representation of `Document::data`.

+ 7 - 0
docs/superpowers/specs/2026-05-15-v2.0-storage-engine-architecture.md

@@ -1,5 +1,12 @@
 # v2.0 Storage Engine — Architecture Decision Record
 
+> **STATUS: implemented, shipped as v2.0.0.** The LMDB substrate this describes
+> is what runs today. Note that several forward-looking notes here have since
+> been overtaken by incidents - in particular LMDB `MDB_dbi` handle lifetime,
+> which caused faults in v2.4.3, v2.4.4 and v2.8.0. Read the LMDB entries in
+> `CLAUDE.md` before touching handle caching or env lifetime.
+
+
 **Status:** Decided 2026-05-15 by operator (fszontagh) following the LMDB spike (`docs/superpowers/spikes/2026-05-15-storage-engine-spike.md`).
 
 **Scope statement:** v2.0 = v1.x feature surface ported to LMDB substrate. Smallest viable architectural change; biggest possible substrate change. New features beyond v1.x parity are explicitly deferred to v2.1+ unless they fall out for free from the substrate switch.

+ 8 - 0
docs/superpowers/specs/2026-08-03-relations-design.md

@@ -1,5 +1,13 @@
 # Relations: referential integrity for a document store
 
+> **STATUS: not implemented.** No `_relations` collection, `RelationManager` or
+> relation RPC exists in the codebase. This design and its companion plan
+> (`docs/superpowers/plans/2026-08-04-relations-v2.5.0.md`) were never executed.
+> The "v2.5.0" target is wrong - v2.5.0 was skipped entirely (2.4.5 → 2.6.0).
+> Re-validate against the current write path before using either: both predate
+> the v2.8.0 fix for caching an `MDB_dbi` before commit.
+
+
 Status: approved design, not yet implemented
 Target: v2.5.0 (Phases A + B together)
 Date: 2026-08-03

+ 9 - 0
docs/superpowers/specs/2026-08-08-rls-cls-design.md

@@ -1,5 +1,14 @@
 # Row- and Column-Level Security — Design
 
+> **STATUS: implemented, shipped as v2.7.0.** Access policy is off by default per
+> project. The design held, but two things only surfaced during implementation
+> and are load-bearing: principal identity MUST be resolved per call rather than
+> from the gRPC `AuthContext` (which is per-connection, and channel pooling made
+> three distinct keys resolve as one), and system collections need a special case
+> in the gate because they are global rather than project-scoped. See the v2.7.0
+> entry in `CLAUDE.md`.
+
+
 **Status:** approved, not yet implemented
 **Date:** 2026-08-08
 **Supersedes:** the claim in `docs/integration-guide.md` that view filters are a

+ 5 - 1
docs/superpowers/spikes/2026-05-15-storage-engine-spike.md

@@ -1,5 +1,9 @@
 # v2.0 Storage Engine — Build-vs-Buy Spike
 
+> **STATUS: concluded. Outcome: LMDB.** This spike is why there is no homegrown
+> page engine. Its recommendation was accepted and shipped as v2.0.0.
+
+
 **Status:** First pass complete. LMDB candidate measured. RocksDB and homegrown deferred.
 
 **Recommendation:** Proceed with LMDB as the v2.0 storage substrate. Numbers below justify the choice; the residual risks (single-writer concurrency under sustained write rates, schema-on-top burden) are bounded and addressable in the execution plan, and the alternative options are strictly more work for unclear gain.
@@ -8,7 +12,7 @@
 
 ## What this spike answers
 
-Per `docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c.md` Task 1, the spike's job is to pick between three candidates:
+Per the (since-removed) Phase C plan skeleton, Task 1, the spike's job is to pick between three candidates:
 
 1. **Homegrown** — 16 KB slotted pages + B+ tree + LRU buffer pool + page-level WAL redo. ~3 months engineering.
 2. **RocksDB** — LSM-tree handles page management; document/vector/index semantics on top of column families. ~3 weeks engineering. Ongoing operational dep on RocksDB.

+ 12 - 7
service/src/database_service.cpp

@@ -93,7 +93,7 @@ bool DatabaseService::initialize() {
         // runs afterward (`persistence_->recover`'s `replayWal` step) also
         // burns through GBs of small Document JSON allocations whose pages
         // sit on the per-thread freelist with no subsequent allocation to
-        // shake them loose. On Zoe (BUG-memory-leak-zoe.md update 15:05)
+        // shake them loose. On Zoe (docs/incidents/2026-04-22-zoe-rss-exceeds-budget.md update 15:05)
         // this was 5.7 GB at 23 min uptime — fully reclaimable via
         // malloc_trim, just nothing called it. Paired with the periodic
         // every-5-minute trim in MemoryStore::logMemoryCheck, so any
@@ -701,12 +701,17 @@ DatabaseService::Config DatabaseService::parseConfig(const nlohmann::json& json)
         if (db.contains("memory")) {
             auto& memory = db["memory"];
 
-            // v2.0 Stage 8 — deprecation log. These knobs control MemoryStore
-            // eviction, which is still active for the MemoryStore-side mirror
-            // but does NOT bound RSS in v2.0 (LMDB mmap is the dominant RSS
-            // contributor). The substrate-level equivalent is the LMDB env
-            // mapsize and OS page cache. v2.1 will rename `max_memory_mb` to
-            // `buffer_pool_size_mb` per the Phase C plan.
+            // Deprecation log. These knobs control MemoryStore eviction, which
+            // is still active for the MemoryStore-side mirror but does NOT bound
+            // RSS since v2.0 (LMDB mmap is the dominant RSS contributor). The
+            // substrate-level equivalent is the LMDB env mapsize and the OS page
+            // cache.
+            //
+            // They disappear when the write-handler migration deletes MemoryStore
+            // (docs/ROADMAP.md, "Pending"). An earlier revision of this comment
+            // promised a v2.1 rename to `buffer_pool_size_mb` "per the Phase C
+            // plan"; that plan was superseded by LMDB before it was written, no
+            // such knob exists, and the rename is not planned.
             for (const char* deprecated : {"max_memory_mb",
                                            "eviction_threshold_percent",
                                            "eviction_target_percent",

+ 10 - 4
service/src/database_service.hpp

@@ -14,10 +14,16 @@
 #include "config/collection_config_manager.hpp"
 #include "security/policy_manager.hpp"
 
-// v2.0 storage substrate (Stage 4) — DocumentStore + LmdbEnv live alongside
-// the v1.x MemoryStore during the Phase C transition. Stage 4 opens both;
-// later stages migrate handlers one group at a time, deleting v1.x paths
-// once they go cold.
+// LMDB storage substrate (v2.0+). DocumentStore + LmdbEnv live alongside the
+// v1.x MemoryStore, which is still the write entry point and a bounded read
+// cache; writes are mirrored into LMDB under MemoryStore's per-collection lock
+// and reads are LMDB-first while the mirror is healthy.
+//
+// This coexistence is deliberate and still current. Retiring MemoryStore means
+// migrating the write handlers onto a WriteCoordinator - see the "Pending"
+// section of docs/ROADMAP.md. Older comments call that "Stage 4"; the "Phase C"
+// label they sometimes carry refers to a homegrown page engine that was never
+// built (LMDB replaced it).
 namespace smartbotic::db::storage {
 class LmdbEnv;
 class DocumentStore;

+ 3 - 3
service/src/main.cpp

@@ -38,7 +38,7 @@ void signalHandler(int signal) {
 }
 
 // v1.8.3: diagnostic handlers for hunting the untracked RSS leak observed
-// on Zoe (BUG-memory-leak-zoe.md, "Update 2026-05-13 10:05"). Both are
+// on Zoe (docs/incidents/2026-04-22-zoe-rss-exceeds-budget.md, "Update 2026-05-13 10:05"). Both are
 // async-signal-unsafe under the spec (they call into glibc allocator
 // state), but `malloc_info` / `malloc_trim` are well-known to work from
 // signal handlers in practice on glibc — and we'd rather have the
@@ -165,7 +165,7 @@ struct CliOverrides {
     bool readOnly = false;
     bool forceReadwrite = false;
     std::optional<smartbotic::database::RecoveryMode> recoveryMode;
-    // v2.0 storage Phase C Stage 3 — one-shot v1.x → v2.0 migration toggles.
+    // One-shot v1.x → v2.0 migration toggles.
     bool migrateFromV1 = false;     // --migrate-from-v1 (force, escape hatch)
     bool noAutoMigrate = false;     // --no-auto-migrate (refuse auto-detect)
 };
@@ -444,7 +444,7 @@ int main(int argc, char* argv[]) {
             config.persistenceConfig.recoveryMode = *cli.recoveryMode;
         }
 
-        // v2.0 storage Phase C Stage 3 — auto-detect-and-migrate before the
+        // Auto-detect-and-migrate v1.x data before the
         // service initialises. Runs against <dataDir>/env/ (the v2.0 LMDB
         // env). On success, sets the schema_version=2 marker so subsequent
         // boots short-circuit. On a refused migration (--no-auto-migrate +

+ 1 - 1
service/src/memory_store.cpp

@@ -1857,7 +1857,7 @@ void MemoryStore::saveToHistory(CollectionData& coll, const Document& currentDoc
     // no-op — history is discarded. Pre-v1.9 this method built a
     // DocumentVersion, pushed it onto an in-heap deque, and stamped
     // `estimatedMemoryBytes_` with the version's overhead. That was
-    // the dominant source of the Zoe RSS gap; see BUG-memory-leak-zoe.md.
+    // the dominant source of the Zoe RSS gap; see docs/incidents/2026-04-22-zoe-rss-exceeds-budget.md.
     if (!historyStore_) return;
 
     // v2.4.5 — per-collection versioning switch. Checked here rather than at

+ 1 - 1
service/src/memory_store.hpp

@@ -670,7 +670,7 @@ private:
         std::map<uint64_t, std::unordered_set<std::string>> expirationIndex;  // expiresAt -> ids
         // v1.9.0 — version history moved to disk (HistoryStore). The
         // in-memory `versionHistory` deque was the dominant cause of the
-        // ~3 GB untracked RSS gap on Zoe (BUG-memory-leak-zoe.md). The
+        // ~3 GB untracked RSS gap on Zoe (docs/incidents/2026-04-22-zoe-rss-exceeds-budget.md). The
         // data was already on disk in the WAL; the in-memory copy was
         // pure redundancy. See `historyStore_` in MemoryStore and the
         // per-collection .hlog files under <data_dir>/history/.

+ 1 - 1
service/src/persistence/history_store.hpp

@@ -23,7 +23,7 @@ namespace smartbotic::database {
  * with `max_versions=0` this accumulated GBs of in-memory state that
  * was also already on disk in the WAL — pure redundancy that no
  * allocator can shrink and no tracker correctly counts. The Zoe leak
- * (BUG-memory-leak-zoe.md) tracked down to this redundancy.
+ * (docs/incidents/2026-04-22-zoe-rss-exceeds-budget.md) tracked down to this redundancy.
  *
  * HistoryStore replaces the in-memory map with one append-only file
  * per collection at `<data_dir>/history/<collection>.hlog`. The format

+ 9 - 1
service/src/storage/document_store_lmdb.cpp

@@ -6,7 +6,15 @@
 //
 // Encode/decode cost is paid per put/get; the substrate-swap exchange is
 // "lose v1.x's in-RAM speed" for "gain on-disk durability and bounded RSS".
-// Stage 4 may optimise to a binary frame later.
+//
+// Two known costs here, both tracked in docs/ROADMAP.md under "Pending":
+//   - encode_document() serialises with nlohmann::dump(), which measured ~11x
+//     slower than yyjson_write on a large document. Swapping it changes the
+//     bytes every write produces, so it needs round-trip equivalence tests.
+//   - Storing metadata and payload as separate frames would let a read hand
+//     back payload bytes without round-tripping at all. Older comments call
+//     this the "Stage 4 binary frame"; it is unrelated to the abandoned
+//     "Phase C" page engine.
 //
 // Collection -> sub-db mapping:
 //   Each collection name is the sub-db name. Sub-db handles (MDB_dbi) are