|
@@ -1,176 +0,0 @@
|
|
|
-# Page-Based Storage Engine (Phase C → v2.0) Implementation Plan — SKELETON
|
|
|
|
|
-
|
|
|
|
|
-> **For agentic workers:** This plan is INTENTIONALLY INCOMPLETE. Only Tasks 1 and 2 are executable now. Tasks 3+ depend on the spike outcome from Task 1 and the architecture brainstorm in Task 2, and CANNOT be specified without those decisions. Do NOT attempt to write or execute Tasks 3+ until Task 2 produces a follow-up plan file.
|
|
|
|
|
->
|
|
|
|
|
-> When Task 2 completes, it MUST author `docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md` with the full TDD-style task breakdown. That follow-up plan is what subagent-driven-development will execute.
|
|
|
|
|
-
|
|
|
|
|
-**Goal:** Replace the "load everything into memory" v1.x architecture with a bounded-memory disk-backed page storage engine. RSS is governed by `buffer_pool_size_mb`, not by total dataset size. This is the "act like MySQL/InnoDB" rewrite that lets Smartbotic Database run a dataset much larger than RAM.
|
|
|
|
|
-
|
|
|
|
|
-**Architecture:** TBD by spike. Three candidate architectures are pre-identified in the v2.0 design doc (`docs/2026-05-15-v2.0-storage-engine-design.md`):
|
|
|
|
|
-
|
|
|
|
|
-1. **Homegrown** — 16 KB slotted pages, B+ tree primary index, LRU buffer pool, page-level WAL redo. We own everything. ~3 months engineering. No new runtime deps.
|
|
|
|
|
-2. **RocksDB-backed** — LSM-tree handles page management; we keep document/vector/index semantics on top of RocksDB column families. ~3 weeks engineering. Ongoing operational dep on RocksDB.
|
|
|
|
|
-3. **LMDB-backed** — mmap'd B+ tree, single-writer concurrency. ~2 weeks engineering. Simpler than RocksDB but with write-concurrency limits that may matter at Zoe's write rate.
|
|
|
|
|
-
|
|
|
|
|
-The spike (Task 1) picks one. Without that decision, the implementation plan cannot be written.
|
|
|
|
|
-
|
|
|
|
|
-**Tech Stack:** TBD by spike. The non-negotiables — C++20, gRPC/Protobuf wire format unchanged, nlohmann::json kept as the document AST surface, yyjson (from Phase A) as the parser — survive any choice.
|
|
|
|
|
-
|
|
|
|
|
----
|
|
|
|
|
-
|
|
|
|
|
-## Context
|
|
|
|
|
-
|
|
|
|
|
-**What changes in v2.0:**
|
|
|
|
|
-
|
|
|
|
|
-- **On-disk layout:** `/var/lib/smartbotic-database/pages/` (or backend-equivalent — RocksDB's `*.sst`, LMDB's `data.mdb`). The current `snapshots/`, `wal/`, `history/` tree from v1.x is gone after one-shot migration.
|
|
|
|
|
-- **WAL format:** page-level redo records replace document-level append. Format break — v1.x WAL is unreadable by v2.0 directly; the migration tool reads it.
|
|
|
|
|
-- **Snapshot format:** page checkpoints replace whole-store LZ4 dumps. Another format break.
|
|
|
|
|
-- **Eviction:** the v1.7.0 chunked eviction machinery (`hot_write_floor_ms`, `eviction_chunk_size`, four pressure levels, evicted-stub WAL fallback, `docWalSeq_`) is GONE. The buffer pool's LRU replacement is the only eviction mechanism. `max_memory_mb` → `buffer_pool_size_mb`.
|
|
|
|
|
-- **Recovery:** the v1.7.0 tiered recovery (snapshot → snapshot_fallback → WAL-only → best_effort → force_empty) collapses to standard checkpoint-and-redo. `recovery.mode` config keeps the same surface for operator continuity but maps to the new engine's terms.
|
|
|
|
|
-
|
|
|
|
|
-**Operator constraint (load-bearing):** "I will not upgrade services which use the database just when the new v2 finished." This means v2.0 is a single-jump release from v1.9.5 (or whatever the latest v1.x is when the spike completes). The `--migrate-from-v1` one-shot mode is REQUIRED — the operator needs `apt upgrade smartbotic-database` to:
|
|
|
|
|
-
|
|
|
|
|
-1. Stop the service (v1.9.5 preinst takes the auto-backup).
|
|
|
|
|
-2. Install v2.0.
|
|
|
|
|
-3. On first boot, detect a v1.x data dir and run the one-shot migration into the new engine's format.
|
|
|
|
|
-4. Start serving.
|
|
|
|
|
-
|
|
|
|
|
-If migration fails, the operator restores from the v1.9.5 backup and reinstalls v1.x — that's the rollback path. v1.9.5's backup-before-upgrade is the load-bearing safety net here; treat it as a hard requirement.
|
|
|
|
|
-
|
|
|
|
|
-**Cross-cutting changes in v2.0 (from the design doc):**
|
|
|
|
|
-- New `service/src/storage/` directory tree replaces `service/src/persistence/`.
|
|
|
|
|
-- `MemoryStore` is renamed (e.g., `DocumentStore` or `PageStore`) and gains a buffer pool.
|
|
|
|
|
-- All replication paths re-routed through the new WAL.
|
|
|
|
|
-- Vector storage (currently parallel float arrays alongside docs) needs a v2.0 home — page-based or alongside the index. Spike must address this.
|
|
|
|
|
-- Field-level encryption stays at the AES-256-GCM layer; only the storage substrate changes.
|
|
|
|
|
-
|
|
|
|
|
----
|
|
|
|
|
-
|
|
|
|
|
-## Tasks
|
|
|
|
|
-
|
|
|
|
|
-### Task 1 — Build-vs-buy spike (1 week)
|
|
|
|
|
-
|
|
|
|
|
-**Files:**
|
|
|
|
|
-- Create: `docs/superpowers/spikes/2026-05-15-storage-engine-spike.md` — the spike writeup.
|
|
|
|
|
-- Create: `service/spike/` — throwaway prototype code, gitignored or committed on a `spike/v2-storage` branch (operator's call).
|
|
|
|
|
-
|
|
|
|
|
-The spike is research, not production code. Its only deliverable is the writeup that justifies the architecture choice.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 1: Define the spike scope.** Prototype the *minimum* path that exercises each candidate's failure modes:
|
|
|
|
|
- - Load 1M Zoe-shape docs into the backend.
|
|
|
|
|
- - Run a sustained 5k write/s + 50k read/s workload.
|
|
|
|
|
- - Crash the process and verify recovery time + data integrity.
|
|
|
|
|
- - Run a SimilaritySearch over a vector subset.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 2: Prototype Option 3 first (LMDB)** — smallest scope, fastest to falsify or validate. Measure: write throughput ceiling under single-writer constraint, mmap RSS behaviour, recovery time.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 3: Prototype Option 2 (RocksDB)** — if Option 3's write-concurrency ceiling is below Zoe's projected sustained write rate, RocksDB is the alternative. Measure: compaction overhead, RSS plateau, recovery time, ops experience (compaction tuning, blob storage for large docs).
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 4: Prototype Option 1 (homegrown) only if both 2 and 3 fail to meet requirements.** Document the requirement gap that justifies a 3-month build over a 2-3 week buy. Be honest — the homegrown option is the heaviest commitment and should require strong evidence to choose.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 5: Write the spike report.** Required sections:
|
|
|
|
|
- - **Scope** — what was measured, what wasn't.
|
|
|
|
|
- - **Results** — numbers per candidate: write tps, read tps, recovery time, RSS curve, on-disk size, vector search latency.
|
|
|
|
|
- - **Operational cost** — what's the ops burden of each choice over 5 years (compaction tuning for RocksDB, single-writer for LMDB, total ownership for homegrown).
|
|
|
|
|
- - **Decision** — one of {homegrown, RocksDB, LMDB} with the reasoning anchored in numbers.
|
|
|
|
|
- - **Rejected alternatives** — what we ruled out and why.
|
|
|
|
|
- - **Risks** — the top 3 things that could derail the chosen path.
|
|
|
|
|
- - **Migration sketch** — high-level mapping of v1.x data → v2.0 backend (the detail goes in Task 2's plan).
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 6: Operator review gate.** The spike report needs operator sign-off before Task 2. Per the brainstorming skill's review-gate pattern.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 7: Commit the spike report** (keep the prototype code on its own branch).
|
|
|
|
|
-
|
|
|
|
|
-```bash
|
|
|
|
|
-git checkout main
|
|
|
|
|
-git add docs/superpowers/spikes/2026-05-15-storage-engine-spike.md
|
|
|
|
|
-git commit -m "spike: v2.0 storage engine build-vs-buy ($DECISION)"
|
|
|
|
|
-```
|
|
|
|
|
-
|
|
|
|
|
----
|
|
|
|
|
-
|
|
|
|
|
-### Task 2 — Architecture brainstorm + execution plan authoring
|
|
|
|
|
-
|
|
|
|
|
-**Files:**
|
|
|
|
|
-- Create: `docs/superpowers/specs/2026-05-15-v2.0-storage-engine-architecture.md` — the architecture decision record.
|
|
|
|
|
-- Create: `docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md` — the full execution plan.
|
|
|
|
|
-
|
|
|
|
|
-Use `superpowers:brainstorming` to drive Step 1; use `superpowers:writing-plans` for Step 4.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 1: Run a brainstorming session** scoped to the v2.0 architecture given the spike outcome. Key questions to drive the session:
|
|
|
|
|
- - Buffer pool size policy — fixed at config, adaptive to system memory, both?
|
|
|
|
|
- - Index strategy — what indexes ship in v2.0 (primary key + per-collection BTrees + vector index)? Which are deferred to v2.1?
|
|
|
|
|
- - Migration UX — operator runs it explicitly, or auto-runs on first boot detecting a v1.x data dir? What does failure look like?
|
|
|
|
|
- - Replication — does the v1.x WAL-streaming replication model survive, or does v2.0 replicate at the page level (Postgres physical replication style)? This is a big architecture call.
|
|
|
|
|
- - Vector storage — co-located with docs in pages, or in a separate index?
|
|
|
|
|
- - Schema/index migrations — v1.x's JSON migrations RPC needs a v2.0 equivalent.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 2: Author the architecture decision record** with the brainstorm outcome.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 3: Operator review gate** on the architecture doc.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 4: Author the v2.0 execution plan** at `docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md`. This plan is full-detail (every TDD step, every file path, every command) at the same level as the Phase A plan in this same directory. Required sections:
|
|
|
|
|
- - File-structure decomposition (what gets created in `service/src/storage/`, what gets deleted from `service/src/persistence/`, etc.)
|
|
|
|
|
- - Migration tool task breakdown
|
|
|
|
|
- - WAL format spec
|
|
|
|
|
- - Snapshot/checkpoint format spec
|
|
|
|
|
- - Buffer pool implementation tasks
|
|
|
|
|
- - Index re-implementation tasks (primary index, vector index)
|
|
|
|
|
- - Replication migration tasks
|
|
|
|
|
- - Eviction-machinery deletion tasks (remove `hot_write_floor_ms`, evicted-stub WAL fallback, `docWalSeq_`, four-level pressure events from v1.7.0)
|
|
|
|
|
- - Performance / soak gates
|
|
|
|
|
- - VERSION bump to 2.0.0
|
|
|
|
|
- - Operator-facing release notes
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 5: Operator review** of the execution plan.
|
|
|
|
|
-
|
|
|
|
|
-- [ ] **Step 6: Commit both docs.**
|
|
|
|
|
-
|
|
|
|
|
-```bash
|
|
|
|
|
-git add docs/superpowers/specs/2026-05-15-v2.0-storage-engine-architecture.md docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md
|
|
|
|
|
-git commit -m "spec+plan: v2.0 storage engine architecture and execution plan"
|
|
|
|
|
-```
|
|
|
|
|
-
|
|
|
|
|
----
|
|
|
|
|
-
|
|
|
|
|
-### Tasks 3+ — Execution
|
|
|
|
|
-
|
|
|
|
|
-**File:** see `docs/superpowers/plans/2026-05-15-v2.0-storage-engine-phase-c-execution.md` (authored by Task 2).
|
|
|
|
|
-
|
|
|
|
|
-The detailed implementation tasks live in the execution plan. They CANNOT be written here because:
|
|
|
|
|
-
|
|
|
|
|
-1. The set of files to create/modify depends on the spike outcome (homegrown vs RocksDB vs LMDB).
|
|
|
|
|
-2. The TDD test design depends on the chosen storage primitives.
|
|
|
|
|
-3. The migration tool's structure depends on the source-to-target schema mapping decided in Task 2.
|
|
|
|
|
-
|
|
|
|
|
-Writing Tasks 3+ here without those decisions would produce exactly the kind of placeholder-laden plan that `writing-plans` warns against. Tasks 1 and 2 exist precisely to retire that uncertainty.
|
|
|
|
|
-
|
|
|
|
|
----
|
|
|
|
|
-
|
|
|
|
|
-## Out of scope (Phase C)
|
|
|
|
|
-
|
|
|
|
|
-- **Distributed scaling.** v2.0 is single-node + replication, same as v1.x. Sharding / Raft / multi-leader is post-v2.0.
|
|
|
|
|
-- **Online schema migration without rewriting docs.** v2.0 keeps the JSON-document-with-optional-schema model.
|
|
|
|
|
-- **Query language change.** Filters, views, vector search keep the v1.x gRPC API surface.
|
|
|
|
|
-- **Anything not flowing from the spike outcome.** No speculative work.
|
|
|
|
|
-
|
|
|
|
|
----
|
|
|
|
|
-
|
|
|
|
|
-## Success criteria
|
|
|
|
|
-
|
|
|
|
|
-1. The spike report exists, is reviewed, and names a chosen architecture with numerical justification.
|
|
|
|
|
-2. The architecture decision record exists and is reviewed.
|
|
|
|
|
-3. The execution plan exists, passes `writing-plans` self-review (no placeholders, no TBDs, every step has code/commands), and the operator approves it.
|
|
|
|
|
-4. Subagent-driven-development can execute the execution plan task-by-task with no ambiguity left in the plan.
|
|
|
|
|
-5. v2.0 ships with a working `--migrate-from-v1` mode that boots cleanly against a v1.9.x data dir (Zoe-shape, full size).
|
|
|
|
|
-6. v1.9.5 backup-before-upgrade still fires on v1.x → v2.0 upgrade — confirmed manually in the soak test.
|
|
|
|
|
-
|
|
|
|
|
----
|
|
|
|
|
-
|
|
|
|
|
-## Open questions (resolved by Task 1 + Task 2)
|
|
|
|
|
-
|
|
|
|
|
-- Build vs buy (Task 1).
|
|
|
|
|
-- Buffer pool sizing policy (Task 2).
|
|
|
|
|
-- Replication model — WAL streaming vs page-level (Task 2).
|
|
|
|
|
-- Vector index placement (Task 2).
|
|
|
|
|
-- Migration UX (Task 2).
|
|
|