Current version: 2.11.3 (see VERSION). Last reviewed: 2026-08-10.
This is the single authoritative statement of what exists and what does not. If any other document in this repository disagrees with this one, this one is right and the other is stale - report it.
Feature-level detail lives in CLAUDE.md. This file covers only status:
shipped, pending, or abandoned.
docs/ tree| Path | What it is | Trust it? |
|---|---|---|
docs/ROADMAP.md |
This file - current status | Yes |
CLAUDE.md |
Per-release feature notes and hard-won lessons | Yes |
docs/integration-guide.md |
Consumer-facing API guide | Yes |
docs/superpowers/specs/ |
Design decisions, each with a status banner | Read the banner first |
docs/superpowers/plans/ |
Implementation plans for work not yet done | Yes, but re-validate against current code |
docs/superpowers/plans/archive/ |
Plans for work already shipped | Historical only - do not execute |
docs/incidents/ |
Post-mortems of resolved production incidents | Historical, lessons folded into CLAUDE.md |
Older documents refer to "Phase A/B/C". That vocabulary comes from the 2026-05-15 storage-engine roadmap and is retained only for reading old comments. Do not plan new work in these terms.
| Phase | Shipped | What actually happened |
|---|---|---|
| A | v1.10.0 | yyjson became the parser on hot paths. Document::data stayed nlohmann::json. Done. |
| B | v1.11.0 | Documents hold a binary doc_binary::Doc (yyjson mut_doc) with a lazy nlohmann view. Done. |
| C | v2.0.0 | Superseded in flight. Phase C as designed meant writing our own engine: 16 KB slotted pages, a BPlusTree, an LRU buffer pool, page-level WAL redo - 2-3 months at high risk. The build-vs-buy spike chose LMDB instead (~2 weeks). Phase C's goal was met (bounded RSS, disk-backed B+ tree) but none of its named components were built. |
Consequences of that substitution, which trip up readers:
BufferPool, no Page, no BPlusTree in this codebase. LMDB
provides all three internally.buffer_pool_size_mb does not exist. A comment in database_service.cpp
once promised max_memory_mb would be renamed to it "in v2.1". That never
happened and is not planned; max_memory_mb still configures MemoryStore
eviction, which no longer bounds RSS. LMDB mapsize plus OS page cache does.document_store_lmdb.cpp:9, "Stage 4 may optimise to a binary frame later" -
part of the write-handler migration below.Every item below is in the installed product as of 2.11.3. CLAUDE.md has the
detail and the failure modes.
<project>:<collection>)DescribeDelete, relations check, per-collection
enforcement switches, and TTL expiry running the same policy as a manual deleteindexes, index-create, index-drop,
index-valuesOrdered by what a reader is most likely to need next. Nothing here has a committed date.
Served from an index: equality, IN, CONTAINS, EXISTS=true, numeric and
string ranges, two-list intersection, result ORDERING, filtered TOTALS, and
distinct values / min / max. Remaining:
CLAUDE.md. The check is correct; it cannot be enforced while
applyDualWriteMirror swallows every exception and MemoryStore mutates before
the mirror runs. Needs the write path restructured, which wants its own change.NE and REGEX would need a full walk either way. A prefix-anchored
REGEX (^abc) could become a range over the string tag - the one real
opportunity here.SEARCH reads whole documents by definition.total_matched requires visiting every match, which is precisely what
stopping early avoids. Changing it means making the total approximate - a
contract decision, not an optimisation._id exists in a posting. Would need an ids-only
response mode.encode_document still serialises with nlohmannstorage/document_store_lmdb.cpp writes documents via doc.toJson().dump().
Measured, nlohmann::dump() is ~11x slower than yyjson_write on a large
document (27.5 ms vs 2.5 ms on 2.91 MB). Swapping it to
doc_binary::to_json_text is one function and no API change, but it alters the
bytes every write produces, so it needs round-trip equivalence tests against an
existing corpus before it can be trusted.
Three subsystems still persist outside LMDB:
| Subsystem | Where it lives now | Target |
|---|---|---|
| History | persistence/history_store.cpp, .hlog files |
_history_<coll> sub-db keyed <doc_id>:<version>; delete history_store.cpp |
| File metadata | on-disk JSON at <filesDir>/records/<2-char-prefix>/<id>.meta.json |
_files sub-db (blobs stay on disk) |
| View definitions | MemoryStore _views system collection |
_views sub-db keyed by qualified view name |
Each is independently shippable. History is the largest.
Handlers would call doc_store_->put() directly, with id generation, timestamps
and versioning moved out of MemoryStore into a WriteCoordinator. On completion
MemoryStore, persistence/wal.cpp, persistence/snapshot.cpp and the eviction
config knobs all delete, and the recovery-mode enum collapses.
This is the v2.x arc's architectural endpoint. It touches the write path that produced the v2.4.3, v2.4.4 and v2.8.0 LMDB handle incidents, so it wants to land in small reviewable pieces with the sub-db identity sentinel kept intact.
Shipped in v2.11.0; see the CLAUDE.md entry for the full surface. Remaining,
all deliberate and none blocking:
restrict is checked one level deep. A cascade that would destroy a
restrict-protected grandchild is refused rather than recursing. Recursion
needs an unbounded transaction and its own WAL-first story.relation_cascade.hpp and the error.cascade/set_null relation whose
parent collection carries a TTL and is renewed after expiry. The fix is a
per-document claim flag that makes the concurrent write lose visibly - not a
versioned delete, which does not close it (no version-checked delete primitive
exists, and the cascade never goes through remove())._relations has no direct-document-write interception, unlike _policies, so
an admin editing the declaration by hand has no effect until restart. A
replication follower likewise does not re-arm until restart._-prefix check on relation names.Raised by a consumer running against 2.11.0 and NOT fixed in 2.11.1, with the reason each was left:
nodes[].config.credentialId). Path resolution descends objects only, so
there is no way to name "this field of every element". An array of scalars is
supported and contributes one posting per element. Supporting the object case
needs a path syntax and a reverse-index maintenance story for element-level
changes; nobody has asked for it twice yet.validate_on_write is all-or-nothing per relation. null and an absent
field are already exempt (they are not references), so "no parent" IS
expressible - but an empty STRING is rejected, and a consumer with legitimately
empty-string references cannot opt those rows out while validating the rest.
The clean answer is for the consumer to write null; a per-relation
"treat empty string as absent" flag is the fallback if that is impossible.RestoreVersion does not reject a write that touches a masked column,
unlike Insert/Update/Upsert (rejectMaskedWrite). Same family as the two
version-read holes 2.11.1 closed, and latent for the same reason: no deployment
has enabled security. Fix when security is first armed for real.The Subscribe unwind is instant (0 ms measured). What remains in a stop, neither caused nor worsened by v2.11.2, both bounded, neither yet investigated:
grpc::Server::Shutdown() when a client is still connected,
draining that connection's transport. With the client killed first, Shutdown
returns in 1 ms. Lowering the 5 s grace would cut the tail, but it would also
cancel legitimately long streaming calls sooner (DownloadFile chunks a whole
file), so it is not a free change.stop() appears to run some component stops twice ("FileManager stopped" is
logged again after "Database service exited cleanly"). Harmless, noisy, worth
a look when either of the above is picked up.Verifying the v2.11.3 unwind on zeus showed docker stop taking 10.5 s, i.e.
hitting Docker's default 10 s timeout and SIGKILLing the process mid
teardown - Database service stopped and exited cleanly were absent from the
log. The final snapshot happened to land first, so nothing was lost, but that is
luck: a slower snapshot would be killed before it renamed into place, leaving the
next boot on WAL-only replay and read-only mode - the exact failure v2.11.2 was
fixing, reintroduced by the container runtime rather than by the code.
A full teardown is ~10 s today (~4 s gRPC transport drain + ~4-5 s component
joins; see 5c), which does not fit in 10 s. The container is now run with
--stop-timeout 120, matching the unit file's TimeoutStopSec=300 in spirit.
Any future docker run for this service must carry --stop-timeout, and any
docker stop in a script wants -t to match. The deb/systemd path is unaffected.
_policies is edited through the
ordinary document API with server-side interception.GetStats, SetReadOnly, Create/DropProject) can
only be gated coarsely - they have no project to evaluate a policy against.
Listener separation is the real boundary there.event.data->dump() and filter-echo f.value.dump() in
database_grpc_impl.cpp are the last nlohmann serialisers on warm paths._orphans_* sub-dbs from the v2.4.4 placement repair still await manual
comparison.nlohmann::json. It is the public client API vocabulary type -
insert, get, Filter::value, QueryResult::documents. Removing it breaks
every consumer call site and needs a soname bump. It is header-only, so it is
not a runtime dependency of any shipped package and costs nothing at runtime.
The remaining ~453 internal references are cold paths (config, migrations,
policies, views) where its exceptions and value semantics are an asset.smartbotic-database-migrate-v1 in v2.5; v2.5 never shipped and the in-binary
path is harmless.