# Smartbotic Database - Status and Roadmap **Current version: 2.10.0** (see `VERSION`). Last reviewed: 2026-08-09. This is the single authoritative statement of what exists and what does not. If any other document in this repository disagrees with this one, this one is right and the other is stale - report it. Feature-level detail lives in `CLAUDE.md`. This file covers only *status*: shipped, pending, or abandoned. --- ## How to read the `docs/` tree | Path | What it is | Trust it? | |------|-----------|-----------| | `docs/ROADMAP.md` | This file - current status | Yes | | `CLAUDE.md` | Per-release feature notes and hard-won lessons | Yes | | `docs/integration-guide.md` | Consumer-facing API guide | Yes | | `docs/superpowers/specs/` | Design decisions, each with a status banner | Read the banner first | | `docs/superpowers/plans/` | Implementation plans for work **not yet done** | Yes, but re-validate against current code | | `docs/superpowers/plans/archive/` | Plans for work already shipped | **Historical only** - do not execute | | `docs/incidents/` | Post-mortems of resolved production incidents | Historical, lessons folded into `CLAUDE.md` | --- ## Storage engine history: Phases A, B, C Older documents refer to "Phase A/B/C". That vocabulary comes from the 2026-05-15 storage-engine roadmap and is **retained only for reading old comments**. Do not plan new work in these terms. | Phase | Shipped | What actually happened | |-------|---------|------------------------| | A | v1.10.0 | yyjson became the *parser* on hot paths. `Document::data` stayed `nlohmann::json`. Done. | | B | v1.11.0 | Documents hold a binary `doc_binary::Doc` (yyjson `mut_doc`) with a lazy nlohmann view. Done. | | C | v2.0.0 | **Superseded in flight.** Phase C as designed meant writing our own engine: 16 KB slotted pages, a `BPlusTree`, an LRU buffer pool, page-level WAL redo - 2-3 months at high risk. The build-vs-buy spike chose LMDB instead (~2 weeks). Phase C's *goal* was met (bounded RSS, disk-backed B+ tree) but **none of its named components were built.** | Consequences of that substitution, which trip up readers: - **There is no `BufferPool`, no `Page`, no `BPlusTree` in this codebase.** LMDB provides all three internally. - **`buffer_pool_size_mb` does not exist.** A comment in `database_service.cpp` once promised `max_memory_mb` would be renamed to it "in v2.1". That never happened and is not planned; `max_memory_mb` still configures MemoryStore eviction, which no longer bounds RSS. LMDB mapsize plus OS page cache does. - "Phase C binary frame" is **not** a Phase C item. The relevant note is `document_store_lmdb.cpp:9`, "Stage 4 may optimise to a binary frame later" - part of the write-handler migration below. --- ## Shipped Every item below is in the installed product as of 2.10.0. `CLAUDE.md` has the detail and the failure modes. - JSON document store: collections, version history, field-level encryption, TTL - LMDB storage substrate, one env per project, dual-write mirror from MemoryStore - Multi-project namespaces (`:`) - Vector storage with SIMD cosine similarity search - Views (read-only projections with baked-in filters) - Row- and column-level access policy, off by default per project - Project-scoped file storage with content-addressed, refcounted, deduplicated blobs - Per-listener TLS and bearer-token auth - Durable snapshots with tiered recovery and read-only lockout - Replication, events/subscribe, migrations, set operations - Paging fast path (no filter, no sort) and the two-pass filtered/sorted scan - Secondary indexes on declared fields: equality, IN, CONTAINS, EXISTS, ranges, intersection, result ordering, filtered totals, and distinct values / min / max. A measured selectivity guard means declaring an index cannot make a query slower. CLI: `indexes`, `index-create`, `index-drop`, `index-values` --- ## Pending Ordered by what a reader is most likely to need next. Nothing here has a committed date. ### 1. Indexing - what is left Served from an index: equality, `IN`, `CONTAINS`, `EXISTS=true`, numeric and string ranges, two-list intersection, result ORDERING, filtered TOTALS, and distinct values / min / max. Remaining: - **Unique constraints.** Built and deliberately not shipped - see the v2.10.0 entry in `CLAUDE.md`. The check is correct; it cannot be enforced while `applyDualWriteMirror` swallows every exception and MemoryStore mutates before the mirror runs. Needs the write path restructured, which wants its own change. - **`NE` and `REGEX`** would need a full walk either way. A prefix-anchored `REGEX` (`^abc`) could become a range over the string tag - the one real opportunity here. - **`SEARCH`** reads whole documents by definition. - **Sort + unselective filter** stays a full pass, and this one is structural: an exact `total_matched` requires visiting every match, which is precisely what stopping early avoids. Changing it means making the total approximate - a contract decision, not an optimisation. - **Covering reads** are closed as not viable: the response always retains six metadata fields and only `_id` exists in a posting. Would need an ids-only response mode. - **Compound (multi-column) indexes.** Intersection covers much of the benefit; a real compound index would beat it for a pair queried constantly. ### 2. `encode_document` still serialises with nlohmann `storage/document_store_lmdb.cpp` writes documents via `doc.toJson().dump()`. Measured, `nlohmann::dump()` is ~11x slower than `yyjson_write` on a large document (27.5 ms vs 2.5 ms on 2.91 MB). Swapping it to `doc_binary::to_json_text` is one function and no API change, but it alters the bytes every write produces, so it needs round-trip equivalence tests against an existing corpus before it can be trusted. ### 3. Sub-db migration remainder (was "Stage 5") Three subsystems still persist outside LMDB: | Subsystem | Where it lives now | Target | |-----------|-------------------|--------| | History | `persistence/history_store.cpp`, `.hlog` files | `_history_` sub-db keyed `:`; delete `history_store.cpp` | | File metadata | on-disk JSON at `/records/<2-char-prefix>/.meta.json` | `_files` sub-db (blobs stay on disk) | | View definitions | MemoryStore `_views` system collection | `_views` sub-db keyed by qualified view name | Each is independently shippable. History is the largest. ### 4. Write-handler migration (was "Stage 4-finish", once labelled the v2.4 milestone) Handlers would call `doc_store_->put()` directly, with id generation, timestamps and versioning moved out of MemoryStore into a `WriteCoordinator`. On completion MemoryStore, `persistence/wal.cpp`, `persistence/snapshot.cpp` and the eviction config knobs all delete, and the recovery-mode enum collapses. This is the v2.x arc's architectural endpoint. It touches the write path that produced the v2.4.3, v2.4.4 and v2.8.0 LMDB handle incidents, so it wants to land in small reviewable pieces with the sub-db identity sentinel kept intact. ### 5. Relations, and unique constraints - one piece of work Referential integrity does not exist: deleting a parent leaves children pointing at nothing, silently, with no way to detect it. `smartbotic-automation` has `workflows` referenced by `executions`, `users` by `sessions`, and `credentials` from node configuration. **The design is re-validated and current** as of 2026-08-09: `docs/superpowers/specs/2026-08-03-relations-design.md`. Read its status section first - it lists what survived re-validation, what was wrong, and the constraints that post-date it. The 3278-line plan beside it is **superseded and must not be executed**; re-plan from the design. **Relations and unique constraints share one blocker**, so schedule them together. Both need `LmdbDocumentStore` operations to accept a caller's `WriteTxn` - relations for atomic cascade across parent, children and index sub-dbs; uniqueness so a rejection can propagate instead of being swallowed by `applyDualWriteMirror`, which currently catches every exception, bumps mirror drift and flips `mirror_healthy_` (the v2.8.1 fault). MemoryStore also mutates before the mirror runs, so a clean rejection needs the in-memory write rolled back. The uniqueness check itself is already built and tested, sitting unreachable behind that. Also requested and specified: **per-collection enable/disable** for relation enforcement and uniqueness, persisted in `CollectionCfg` alongside `versioningEnabled` and `indexedFields`, and re-armed at boot the way `applyIndexDeclarations()` already does. ### 6. Smaller known gaps - Policy management has no dedicated RPCs; `_policies` is edited through the ordinary document API with server-side interception. - Service-wide operations (`GetStats`, `SetReadOnly`, `Create/DropProject`) can only be gated coarsely - they have no project to evaluate a policy against. Listener separation is the real boundary there. - `event.data->dump()` and filter-echo `f.value.dump()` in `database_grpc_impl.cpp` are the last nlohmann serialisers on warm paths. - `_orphans_*` sub-dbs from the v2.4.4 placement repair still await manual comparison. - TTL is not retro-applied: rows written before a default TTL existed keep no expiry. --- ## Not planned - **v3.** v2.x is the terminal arc. - **Removing `nlohmann::json`.** It is the public client API vocabulary type - `insert`, `get`, `Filter::value`, `QueryResult::documents`. Removing it breaks every consumer call site and needs a soname bump. It is header-only, so it is not a runtime dependency of any shipped package and costs nothing at runtime. The remaining ~453 internal references are cold paths (config, migrations, policies, views) where its exceptions and value semantics are an asset. - **A homegrown page engine or buffer pool.** Settled: LMDB. - **v1.x auto-migration as a separate package.** Was slated to extract to `smartbotic-database-migrate-v1` in v2.5; v2.5 never shipped and the in-binary path is harmless.