2026-05-15-v2.0-roadmap.md 8.1 KB

v2.0 Storage Engine Rewrite — Roadmap

Status: Roadmap stub. Phase A and beyond require their own detailed plan files before subagent dispatch. Do NOT try to execute this file directly — it is an index.

Source of design: /data/smartbotic-database/docs/2026-05-15-v2.0-storage-engine-design.md (316 lines, already merged into main as the canonical design).

Why this exists: The operator asked to "proceed with them all" — all three phases of the v2.0 design. Phases A/B/C span ~4–5 months of engineering and cannot execute in a single session. This file:

  1. Lists the order in which the phases ship
  2. Captures cross-cutting upgrade-safety requirements that span every phase
  3. Identifies where the WAL/snapshot format breaks (Phase C, not A/B)
  4. Points each phase at the plan file it needs before subagents can be dispatched

Cross-cutting upgrade-safety requirements (operator-mandated)

These apply to every phase, not just Phase C:

  1. Backup before upgrade — already shipping in v1.9.5 (2026-05-15-v1.9.5-backup-before-upgrade.md). Every subsequent release inherits this safety net.
  2. WAL/snapshot backward compatibility — each release must boot cleanly on the previous release's data dir. Concretely:
    • v1.10 (Phase A) must read v1.9.x WAL + snapshot bytes unchanged (yyjson swap is only the parser; on-disk format is identical).
    • v1.11 (Phase B) introduces binary doc storage in memory but keeps the WAL/snapshot doc payload as JSON text on disk so v1.10 readers can still parse v1.11 snapshots. (Migration to a binary on-disk format is deferred to Phase C.)
    • v2.0 (Phase C) is the format break: page-based on-disk layout, WAL becomes page-level redo. A one-shot --migrate-from-v1 mode is required (see Phase C plan when written).
  3. Pin-back path — apt install smartbotic-database=<old-version> plus the v1.9.5 backup restore procedure must remain a working rollback for at least one release back.

Phase A → v1.10.0 — yyjson hot-path swap

Status: Plan not yet written. Blocked on: writing 2026-05-15-v1.10-yyjson-phase-a.md.

Goal: Replace nlohmann::json parsing with yyjson at hot paths (WAL replay, snapshot deserialize, gRPC request bodies, history reads). Keep nlohmann::json as the in-memory document representation — Phase A is a parse-speed win, not a memory win.

Honest impact assessment: With jemalloc already shipped in v1.9.4, RSS is bounded. Phase A buys faster boot (parse-heavy paths) and slightly lower gRPC request latency on large bodies, but does not reduce steady-state memory. The operator should know this before Phase A is scheduled — if memory is the priority, Phase B is the work that matters.

Files in scope (parse sites identified, see grep output in the design doc):

  • service/src/persistence/wal.cpp:233,249 — replay
  • service/src/persistence/snapshot.cpp:820,929,963 — deserialize
  • service/src/persistence/history_store.cpp:124 — version read
  • service/src/database_grpc_impl.cpp — Insert/Update/Patch/Find handlers (~10 call sites)
  • service/src/database_service.cpp:722 — replication apply
  • packaging/Dockerfile.base — add libyyjson-dev
  • packaging/deb/templates/control.server — add libyyjson0 runtime dep
  • service/CMakeLists.txt — pkg_check_modules(yyjson REQUIRED) or find_package(yyjson)

Task breakdown sketch (to be expanded in the Phase A plan file):

  1. Add yyjson build + runtime deps, bump Dockerfile.base tag
  2. Helper: yyjson_to_nlohmann(const yyjson_val*) -> nlohmann::json — the bridge
  3. Swap WAL replay parse → re-run replica eviction load test
  4. Swap snapshot deserialize parse → re-run snapshot durability test
  5. Swap gRPC handler parses → re-run integration tests
  6. Bench: report parse-time delta on a Zoe-shape snapshot
  7. Bump VERSION to 1.10.0, build deb, sync

Time estimate: 1–1.5 weeks of subagent-driven work across 2–3 sessions.


Phase B → v1.11.0 — binary lazy document storage

Status: Plan not yet written. Blocked on: writing 2026-05-15-v1.11-binary-docs-phase-b.md and settling the format choice.

Goal: Document in-memory representation changes from heap-allocated nlohmann::json AST → std::vector<uint8_t> binary with lazy .data() accessor that parses on demand. This is the phase that actually reduces memory.

Format decision (open): BSON vs custom packed-JSON tape vs yyjson mut_doc-as-storage. Each has trade-offs:

  • BSON — well-specified, has libbson; but spec includes types we don't use (Date, ObjectId, Decimal128) and field tagging adds overhead.
  • Custom tape — minimal, fast, but we own the spec forever.
  • yyjson mut_doc — keep the parser's own buffer as storage; clean but ties us to yyjson semantics.

Pre-execution work: Operator should brainstorm format choice (superpowers:brainstorming skill) before this plan can be written.

Files in scope: service/src/document.hpp, service/src/memory_store.cpp/.hpp, every callsite that does doc.data["field"] (large blast radius — ~80+ call sites across service + tests).

Wire compatibility: gRPC Document.data stays JSON text on the wire. Conversion happens at the boundary.

Snapshot compat: Snapshot format v6 emits documents as JSON text (same as v5) so v1.10 readers still work. Phase C is where the on-disk format breaks.

Time estimate: 3–4 weeks across 5–8 sessions.


Phase C → v2.0 — page-based on-disk storage + buffer pool

Status: Plan not yet written. Blocked on: the build-vs-buy spike (1 week) AND a brainstorm session on the architecture AND writing 2026-05-15-v2.0-storage-engine-phase-c.md.

Goal: Replace the "load everything into memory" architecture with bounded-memory disk-backed page storage. This is the real "act like MySQL/InnoDB" rewrite.

The build-vs-buy spike (must run first):

  • Option 1: Homegrown 16 KB slotted pages + B+ tree index + LRU buffer pool + page-level WAL redo (~3 months engineering)
  • Option 2: RocksDB-backed prototype — let LSM do the page management, keep our document/vector/index semantics on top (~3 weeks engineering, ongoing operational dep on RocksDB)
  • Option 3: LMDB-backed — simpler, single-writer, mmap'd B+ tree (~2 weeks engineering, but write-concurrency limits)

The spike picks one. Without that decision, the Phase C plan cannot be written.

Required cross-cutting work in Phase C:

  • One-shot --migrate-from-v1 boot mode: read v1.x snapshot + replay v1.x WAL into pages, then write a v2.0 checkpoint and start serving
  • New WAL format (page-level redo records), new snapshot format (page checkpoints), new on-disk layout under /var/lib/smartbotic-database/pages/
  • Eviction machinery from v1.7.0 removed — page LRU subsumes it; max_memory_mb → buffer_pool_size_mb; evicted stubs / WAL fallback / docWalSeq_ all deleted

Time estimate: ~13 weeks across many sessions, contingent on spike outcome.


What to do next session

  1. Confirm v1.9.5 backup mechanism is shipped (current session's deliverable).
  2. If continuing into Phase A: write 2026-05-15-v1.10-yyjson-phase-a.md with full task breakdown (use superpowers:writing-plans). Then dispatch implementers per superpowers:subagent-driven-development.
  3. If considering reordering: the honest recommendation is to skip Phase A and go directly to Phase B if memory is the priority. Phase A is foundation work whose primary benefit (parser swap) is also achievable as a Phase B side-effect (yyjson would naturally back the binary doc storage in B). Discuss with operator.

What NOT to do without operator approval

  • Do not start Phase A without re-confirming memory vs latency priorities with the operator.
  • Do not start Phase B without settling the binary format choice (BSON vs tape vs yyjson mut_doc).
  • Do not start Phase C without running and writing up the build-vs-buy spike.
  • Do not break WAL/snapshot backward compat outside of Phase C.
  • Do not skip the v1.9.5 backup mechanism — it is the rollback safety net for everything that follows.