STATUS: partly superseded. Read this before the rest of the document.
Phases A and B shipped as described (v1.10.0, v1.11.0). Phase C did not. The "Phase C" section below specifies a homegrown engine - 16 KB slotted pages,
BPlusTree,BufferPool, page-level WAL redo. None of it was built. The build-vs-buy spike chose LMDB instead, which shipped as v2.0.0 and provides those internally. There is noPage,BufferPoolorBPlusTreetype in this codebase, andbuffer_pool_size_mbdoes not exist.Kept because the problem analysis and the phase reasoning are still sound. For current status see
docs/ROADMAP.md.Date: 2026-05-15 Author: design draft Status: Proposed Supersedes: nothing yet — v1.x in-memory-everything model still in production
v1.x is a memory database with WAL/snapshot persistence. It assumes the entire working set fits in RAM. v1.9.x patched the symptoms (jemalloc to bound allocator retention, history on disk to stop tracker inflation) but the architecture still requires RAM proportional to dataset size.
For ShadowMan-Zoe today: 2.1 M docs / 1.6 GB tracker / 3-4 GB RSS — fits comfortably in 11 GB host. For ShadowMan-at-Battery scale (target: 20 M+ docs, files, telemetry rollups), we'll be back at the same wall.
The constraint we're hitting: RAM is the cap. Eviction is a fallback, not a design feature. WAL fallback for evicted docs is O(WAL size) per read — works for tens of evictions, falls over at thousands.
MySQL/InnoDB's design choice: RAM is a configured budget (the buffer pool). Disk is the canonical store. The buffer pool caches hot pages. A 100 GB database on a 4 GB pool just means more cache misses, not OOM.
v2.0 adopts the buffer-pool model.
┌─────────────────────────────────────────────────────────────────┐
│ gRPC handlers │
│ (Insert / Get / Update / Find / Subscribe etc.) │
│ unchanged wire API │
└──────────────────────────────┬──────────────────────────────────┘
│
┌──────────────────────────────▼──────────────────────────────────┐
│ Document API layer │
│ - parses JSON wire bytes once with yyjson (replaces nlohmann) │
│ - converts to compact binary doc record on write │
│ - converts binary back to JSON on read │
└──────────────────────────────┬──────────────────────────────────┘
│
┌──────────────────────────────▼──────────────────────────────────┐
│ Storage engine (NEW in v2.0) │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ BufferPool (bounded LRU page cache) │ │
│ │ buffer_pool_size_mb (default 1024) — hard RAM cap │ │
│ └─────────────────────────────────────────────────────────────┘ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ B+ tree indexes (one per collection) │ │
│ │ key: doc_id → value: page_offset │ │
│ └─────────────────────────────────────────────────────────────┘ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ Page files: {data_dir}/pages/{collection}.dat │ │
│ │ 16 KB pages, slotted-page format │ │
│ └─────────────────────────────────────────────────────────────┘ │
│ ┌─────────────────────────────────────────────────────────────┐ │
│ │ WAL (mostly unchanged) │ │
│ │ now records page-level redo records │ │
│ └─────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
This is a multi-release effort. Three phases that each ship value without depending on the next one.
yyjson swap (option 1)nlohmann::json with yyjson on the parse hot path
(gRPC request body parsing, WAL deserialize, snapshot deserialize)nlohmann::json in the API for now (it's used everywhere)yyjson's flatter representationShip target: v1.10.0 Effort: 1-2 weeks Risk: low — drop-in API change in a few hot files; tests cover regressions
Document storage: std::vector<uint8_t> binary (BSON or
custom packed format) instead of nlohmann::json dataDocument::data() is a lazy accessor that parses the binary on
demand, used only by paths that need a JSON view (gRPC response
serialization, projection in views, etc.)yyjson, convert to
binary, storeShip target: v1.11.0 or v2.0-alpha
Effort: 3-5 weeks
Risk: medium — touches Document, every read-write path, snapshot
format. Snapshot format v6 needed (read v3-v5 on the way in, migrate).
Expected memory drop on Zoe-shape workload:
The big rewrite. The Document API stays binary (from Phase B); we swap out the storage layer underneath.
Ship target: v2.0 Effort: 2-3 months Risk: high — new on-disk format, B+ tree implementation, recovery semantics, migration tooling
Subphases inside Phase C:
Page struct: 16 KB, fixed header + slotted-page bodyPageFile: one file per collection, mmap'd or read via positional preadBufferPool: bounded LRU map, page_id → in-memory page buffer,
configurable buffer_pool_size_mb, dirty-page write-back on evictionBPlusTree<Key, Value> parameterized on key type (string for doc_id)MemoryStore::insert/update/get/find route to the new enginecoll->documents hash map goes awaycoll->vectors may stay in RAM for vector search (hot, small) or
move to page-stored vectors per doc(page_id, offset, before, after)max_memory_mb becomes buffer_pool_size_mb+--------------------------------------------------------------+
| Page header (32 B): |
| magic(4) + page_id(8) + lsn(8) + slot_count(2) + |
| free_space_offset(2) + checksum(8) |
+--------------------------------------------------------------+
| Slot directory (4 B per slot): |
| [offset(2) + length(2)] × slot_count |
+--------------------------------------------------------------+
| ... free space grows down from free_space_offset ... |
| |
| |
| |
| ... record data grows up from end ... |
| [record N data: doc_id_len + doc_id + bin_doc_len + bin_doc] |
| [record N-1 data: ...] |
+--------------------------------------------------------------+
Slotted pages are battle-tested (PostgreSQL, InnoDB, SQLite use variants). Records inside a page can be variable length. Slot directory at the top is sorted by slot_id so a doc_id lookup costs one binary search within the page.
+----------------------------------------------------+
| Header (24 B): |
| version(8) + createdAt(8) + updatedAt(8) |
+----------------------------------------------------+
| Metadata: encrypted(1) + node_id_len(2) + node_id |
+----------------------------------------------------+
| Data: bin_doc_len(4) + bin_doc |
+----------------------------------------------------+
bin_doc is either BSON or a packed-JSON we define. BSON is well-known
but heavyweight in places; let's spec a small custom format that's a
1:1 mapping of JSON values to length-tagged bytes (close to "what
yyjson outputs as a binary tape"). Decoding back to JSON is just a
serializer.
Vectors are dense float[dim]. They get their own page section
attached to the doc record or, for vector_dimension > 4096, a
separate "vector page" referenced by the doc's vector_offset field.
Unchanged from v1.9.0. The HistoryStore design we just shipped is already disk-backed and page-friendly. v2.0 keeps it as-is.
v1.x → v2.0 needs a one-time data migration:
--migrate-from-v1 flagRollback path: rename data dir back, downgrade deb, restart.
We pin v1.11.x and v2.0 in apt so operators can choose; we don't auto-migrate. ShadowMan-cpp gets a migration runbook.
cp -r of the page files +
WAL).getAllDocuments scans the B+ tree, no longer returns a vector
of fully-loaded Documents. Becomes an iterator-style API. Most
callers want pagination anyway.| Property | v1.9.x | v2.0 |
|---|---|---|
| Dataset size limit | bounded by RAM | bounded by disk |
| RAM usage | grows with dataset | bounded by buffer_pool_size_mb |
| Read latency (cache hit) | hash map lookup, µs | B+ tree lookup, low µs |
| Read latency (cache miss) | WAL fallback, ms-s | page read, sub-ms |
| Recovery time | parse all snapshot data | replay WAL forward |
| Memory predictability | rough | exact |
| Document tree heap overhead | 4-5× | ~1.1× (binary record) |
| Eviction work under pressure | O(N docs) scan + stub | O(1) page replacement |
| Compatible with current API | yes | yes |
| Compatible with current data | yes (in-place) | one-time migration |
find with sort, hash gives faster point lookups. Default
to B+ tree (sortable); add hash as opt-in via collection option
(index_type: hash)._views, _collection_meta,
_migrations are small (<100 docs); keeping them in RAM is fine.
Add a cache_all: true option per collection that pins everything
in pool.Document JSON in
ReplicationEntry. v2.0 could send the binary record (smaller,
faster on both sides). Wire compatibility matters; need to
version-tag the payload.SimilaritySearch brute-force scans
the entire coll->vectors map. For v2.0 we either keep vectors
pinned in RAM (works for current scale) or build a proper ANN index
(HNSW, IVF). Decision deferred — keep RAM-pinned vectors in v2.0,
add ANN in v2.1+.buffer_pool_size_mb × 1.1 under sustained load. No SIGUSR2
intervention. No OOM kills.tests/test_* pass against v2.0 (some need updates for
removed APIs, but no test logic should change).| Phase | Releases | Calendar | Engineer-weeks |
|---|---|---|---|
| Phase A | v1.10.0 | week 1-2 | 1.5 |
| Phase B | v1.11.0 | week 3-7 | 4 |
| C.1-C.2 | v2.0-alpha1 | week 8-14 | 6 |
| C.3-C.4 | v2.0-alpha2 | week 15-20 | 5 |
| C.5 | v2.0-beta | week 21-22 | 1.5 |
| Bake | v2.0 | week 23-24 | 1.5 |
| Total | — | ~6 mo | ~19 wks |
Aggressive but each phase is independently shippable. If C is paused at any point, we still have the (1)+(2) gains from A+B.