Kaynağa Gözat

feat(index): secondary indexes in the storage layer

A filtered query was a full collection scan. v2.8.0 removed document
materialisation from it (2463ms -> 402ms on 505 MB), but what remained is
yyjson parsing every value in the collection to test the predicate, and that
is the floor for a scan. This stops reading non-matching rows at all.

Layout: one _idx_<collection>#<field> sub-db per indexed field, MDB_DUPSORT,
key = order-preserving encoded value, data = document id. DUPSORT gives the
posting list for a value as a sorted set at no extra structural cost.

Measured on a copy of the live executions collection (10,089 rows / 505 MB):
  EQ workflowId, 1 match (0.01%)          372ms -> <1ms  (~1000x)
  EQ status="completed", 6,665 (66%)      347ms -> 344ms, index declined
Both returned identical results to the unindexed plan.

Three things this had to get right, each a silent-wrong-data risk:

1. Index equality must mean what the scan's equality means. The scan's EQ is
   nlohmann::json::operator==, which compares numbers across subtypes, so
   json(5) == json(5.0). encode_index_key canonicalises integral values to the
   integer form or an indexed EQ would miss rows a scan finds.
   test_secondary_index_keys checks agreement over every pair in a corpus.
   Integers are not routed through double either: ns timestamps exceed 2^53,
   where neighbouring values collide as doubles.

2. Maintenance runs in the SAME write transaction as the document, so a write
   that throws rolls the index back with it. Deletes remove the specific id
   from the value's posting list rather than the whole list.

3. An index must not be able to make a query slower. The guard uses it only
   while matches are under rows/10 - derived from measurement, since a full
   decode_document costs ~8.6x a yyjson parse of the same row, so an index
   stops paying above ~11.6% selectivity. status="completed" at 66% is exactly
   the case that would regress.

The index changes only WHICH rows pass 1 visits. Predicate evaluation, sort-key
extraction, sorting and pagination are shared verbatim with the scan path, and
every filter is still applied to every candidate, so the index is allowed to be
a superset - it narrows work, it does not decide the answer.
test_indexed_and_unindexed_plans_agree runs 13 query shapes both ways and
compares ids, order, total_matched and has_more. index_plan_stats() proves the
plan was taken rather than silently declined - without it those cases would be
comparing the scan against itself.

NOT usable yet, and docs/ROADMAP.md says so: no RPC, no client API, and
set_indexed_fields is in-memory so nothing survives a restart.

ctest 20/20.
fszontagh 1 ay önce
ebeveyn
işleme
c4e0c72d25

+ 41 - 14
docs/ROADMAP.md

@@ -75,20 +75,47 @@ detail and the failure modes.
 Ordered by what a reader is most likely to need next. Nothing here has a
 committed date.
 
-### 1. Indexing (not designed yet)
-
-The largest remaining performance gap, and the natural next piece of work.
-
-A filtered query is a **full collection scan**. The v2.8.0 two-pass scan removed
-document materialisation (2463 ms → 402 ms on `executions`, 9800 docs / 505 MB),
-but the remaining 402 ms is yyjson parsing every value in the collection to test
-the predicate. That is the floor for a scan, and it is paid per query - so a
-consumer polling a filtered view costs a large fraction of a core continuously.
-
-Getting below it requires not reading non-matching rows at all: a secondary index
-mapping field value → doc id, maintained on write. Nothing exists yet - no design
-doc, no schema, no decision on which fields get indexed or whether declaration is
-explicit or automatic.
+### 1. Indexing (storage layer done, NOT yet reachable by clients)
+
+**What exists and works**, with tests, in `LmdbDocumentStore`:
+
+- `_idx_<collection>#<field>` sub-dbs, `MDB_DUPSORT`, key = order-preserving
+  encoded value, data = document id (`storage/secondary_index.{hpp,cpp}`).
+- Maintenance inside the **same write transaction** as the document, so an
+  aborted write cannot leave a stale index. Unindexed collections pay nothing.
+- `build_index` backfills existing rows, idempotently; `drop_index` removes one.
+- A **selectivity guard**: an index is used only while matches are under
+  `rows / 10`. Derived from measurement, not taste - a full `decode_document`
+  costs ~8.6x a yyjson parse of the same row, so an index stops paying above
+  ~11.6% selectivity. Declaring an index therefore cannot pessimise a query.
+- `index_plan_stats()` reports indexed vs full scans vs declined-as-unselective.
+
+Measured on a copy of the live `smartbotic-automation` `executions` collection
+(10,089 rows / 505 MB):
+
+| query | before | after |
+|---|---|---|
+| EQ on `workflowId` (1 match, 0.01%) | 372 ms | **<1 ms** (~1000x) |
+| EQ on `status="completed"` (6,665 matches, 66%) | 347 ms | 344 ms, index **declined** |
+
+Both returned identical results to the unindexed plan.
+
+**What is missing - it cannot be used yet:**
+
+- No RPC. No `CreateIndex` / `DropIndex` / `ListIndexes`.
+- No persistence of the declaration. `set_indexed_fields()` is in-memory, so
+  nothing survives a restart. The declaration belongs in the `_collection_meta`
+  system collection (`CollectionCfg`), which is an ordinary collection and so is
+  WAL'd and snapshotted for free - the same reasoning that put
+  `versioningEnabled` there.
+- No client API. It must be **methods** (`createIndex`/`dropIndex`/
+  `listIndexes`), never new members on a public struct, since changing a struct's
+  size under an unchanged soname crashes installed consumers.
+- Only `EQ` is served. Ranges need the numeric encoding unified first: integral
+  and non-integral numbers currently sit under different type tags, which is
+  correct for equality but means they order independently.
+- `CONTAINS` (array membership) would need a different index shape - one posting
+  per element rather than per value.
 
 ### 2. `encode_document` still serialises with nlohmann
 

+ 1 - 0
service/CMakeLists.txt

@@ -63,6 +63,7 @@ set(DATABASE_SERVICE_SOURCES
     src/storage/lmdb_txn.cpp
     src/storage/lmdb_dbi.cpp
     src/storage/document_store_lmdb.cpp
+    src/storage/secondary_index.cpp
     src/storage/subdb_identity.cpp
     src/storage/subdb_placement.cpp
     src/storage/migrate_v1_to_v2.cpp

+ 416 - 30
service/src/storage/document_store_lmdb.cpp

@@ -53,6 +53,7 @@
 #include "storage/lmdb_dbi.hpp"
 #include "storage/lmdb_env.hpp"
 #include "storage/lmdb_txn.hpp"
+#include "storage/secondary_index.hpp"
 #include "storage/subdb_identity.hpp"
 
 namespace smartbotic::db::storage {
@@ -307,7 +308,8 @@ LmdbDocumentStore::LmdbDocumentStore(LmdbEnv& env) : env_(env) {
 }
 
 unsigned int LmdbDocumentStore::open_for_write(WriteTxn& wtxn,
-                                                std::string_view collection) {
+                                                std::string_view collection,
+                                                unsigned int extra_flags) {
     std::string key(collection);
     {
         std::lock_guard<std::mutex> lock(cache_mutex_);
@@ -326,7 +328,8 @@ unsigned int LmdbDocumentStore::open_for_write(WriteTxn& wtxn,
     }
     MDB_dbi raw_dbi = 0;
     std::string name = to_cstr(collection);
-    mdb_check(mdb_dbi_open(wtxn.raw(), name.c_str(), MDB_CREATE, &raw_dbi),
+    mdb_check(mdb_dbi_open(wtxn.raw(), name.c_str(), MDB_CREATE | extra_flags,
+                            &raw_dbi),
               "dbi_open (collection create)");
     // Stamp on every fresh open, not only on creation, so sub-dbs that
     // pre-date v2.4.4 acquire a sentinel on their next write without a
@@ -444,7 +447,10 @@ size_t LmdbDocumentStore::prime_dbi_cache() {
 
     for (const auto& n : names) {
         MDB_dbi dbi = 0;
-        int rc = mdb_dbi_open(wtxn.raw(), n.c_str(), 0, &dbi);
+        // Index sub-dbs are MDB_DUPSORT; pass the flag on reopen so the handle
+        // agrees with how the sub-db was created.
+        const unsigned int flags = is_index_subdb(n) ? MDB_DUPSORT : 0u;
+        int rc = mdb_dbi_open(wtxn.raw(), n.c_str(), flags, &dbi);
         if (rc == MDB_NOTFOUND) continue;          // vanished under us; ignore
         if (rc != MDB_SUCCESS) throw_mdb(rc, "dbi_open (prime)");
         opened[n] = dbi;
@@ -468,10 +474,276 @@ void LmdbDocumentStore::put(std::string_view collection,
     WriteTxn wtxn(env_);
     unsigned int dbi = open_for_write(wtxn, collection);
     MDB_val k = to_val(id);
+
+    // v2.9.0 — index maintenance runs in THIS transaction, so a write that
+    // throws after this point rolls the index back with the document. An
+    // asynchronously-maintained index would let a query read entries for a row
+    // that was never stored, and return silently wrong rows rather than an error.
+    std::vector<std::pair<std::string, unsigned int>> index_dbis;
+    if (!indexed_fields(collection).empty()) {
+        MDB_val old{0, nullptr};
+        const int rc = mdb_get(wtxn.raw(), dbi, &k, &old);
+        if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) throw_mdb(rc, "get (pre-index)");
+        const std::string_view old_payload =
+            rc == MDB_SUCCESS ? to_sv(old) : std::string_view{};
+        maintainIndexes(wtxn, collection, id, old_payload, &doc, index_dbis);
+    }
+
     MDB_val v = to_val(payload);
     mdb_check(mdb_put(wtxn.raw(), dbi, &k, &v, 0), "put");
     wtxn.commit();
     cacheCommittedDbi(collection, dbi);
+    for (const auto& [sub, d] : index_dbis) cacheCommittedDbi(sub, d);
+}
+
+
+// -------------------------------------------------------------------------
+// v2.9.0 secondary index maintenance
+// -------------------------------------------------------------------------
+
+void LmdbDocumentStore::set_indexed_fields(std::string_view collection,
+                                            std::vector<std::string> fields) {
+    std::lock_guard<std::mutex> lock(index_mutex_);
+    if (fields.empty()) {
+        indexed_fields_.erase(std::string(collection));
+    } else {
+        indexed_fields_[std::string(collection)] = std::move(fields);
+    }
+}
+
+std::vector<std::string>
+LmdbDocumentStore::indexed_fields(std::string_view collection) {
+    std::lock_guard<std::mutex> lock(index_mutex_);
+    auto it = indexed_fields_.find(std::string(collection));
+    if (it == indexed_fields_.end()) return {};
+    return it->second;
+}
+
+void LmdbDocumentStore::maintainIndexes(
+    WriteTxn& wtxn,
+    std::string_view collection,
+    std::string_view id,
+    std::string_view old_payload,
+    const smartbotic::database::Document* new_doc,
+    std::vector<std::pair<std::string, unsigned int>>& to_cache) {
+
+    const auto fields = indexed_fields(collection);
+    if (fields.empty()) return;   // unindexed collections pay nothing
+
+    // Old keys come from the STORED bytes, parsed with yyjson and resolved one
+    // field at a time - never materialised into a Document. On a 505 MB
+    // collection a full decode_document costs ~1ms per write; resolving just the
+    // indexed fields off the yyjson tree is a fraction of that.
+    std::unordered_map<std::string, std::optional<std::string>> old_keys;
+    if (!old_payload.empty()) {
+        yyjson_doc* d = yyjson_read(old_payload.data(), old_payload.size(), 0);
+        if (d) {
+            yyjson_val* root = yyjson_doc_get_root(d);
+            for (const auto& f : fields) {
+                auto v = resolve_from_yyjson(root, f);
+                old_keys[f] = v ? encode_index_key(*v) : std::nullopt;
+            }
+            yyjson_doc_free(d);
+        }
+    }
+
+    for (const auto& f : fields) {
+        std::optional<std::string> new_key;
+        if (new_doc != nullptr) {
+            // resolveFilterValue is the SAME resolver the scan uses, so an
+            // indexed lookup and a scan cannot disagree about which field a
+            // filter names.
+            auto v = filter_eval::resolveFilterValue(*new_doc, f);
+            if (v) new_key = encode_index_key(*v);
+        }
+        std::optional<std::string> old_key;
+        if (auto it = old_keys.find(f); it != old_keys.end()) old_key = it->second;
+
+        // An unchanged value needs no index write at all - the common case for
+        // an update that touches other fields.
+        if (old_key == new_key) continue;
+
+        const std::string sub = index_subdb_name(collection, f);
+        const unsigned int dbi = open_for_write(wtxn, sub, MDB_DUPSORT);
+        to_cache.emplace_back(sub, dbi);
+
+        if (old_key) {
+            MDB_val k = to_val(*old_key);
+            MDB_val v = to_val(id);
+            // DUPSORT: passing the data removes just THIS id from the value's
+            // posting list, not every id under that value.
+            const int rc = mdb_del(wtxn.raw(), dbi, &k, &v);
+            if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) {
+                throw_mdb(rc, "index del");
+            }
+        }
+        if (new_key) {
+            MDB_val k = to_val(*new_key);
+            MDB_val v = to_val(id);
+            // MDB_NODUPDATA makes a repeat put a no-op rather than an error, so
+            // re-indexing the same row twice is safe (build_index re-runs).
+            const int rc = mdb_put(wtxn.raw(), dbi, &k, &v, MDB_NODUPDATA);
+            if (rc != MDB_SUCCESS && rc != MDB_KEYEXIST) {
+                throw_mdb(rc, "index put");
+            }
+        }
+    }
+}
+
+std::optional<uint64_t>
+LmdbDocumentStore::index_count_eq(std::string_view collection,
+                                   const std::string& field,
+                                   const nlohmann::json& value) {
+    ReadTxn rtxn(env_);
+    return countIndexEqTxn(rtxn, collection, field, value);
+}
+
+std::optional<uint64_t>
+LmdbDocumentStore::countIndexEqTxn(ReadTxn& rtxn,
+                                    std::string_view collection,
+                                    const std::string& field,
+                                    const nlohmann::json& value) {
+    const auto key = encode_index_key(value);
+    if (!key) return std::nullopt;      // unindexable value; caller must scan
+
+    auto dbi_opt = try_open_for_read(rtxn, index_subdb_name(collection, field));
+    if (!dbi_opt) return std::nullopt;   // no such index
+
+    MDB_cursor* cur = nullptr;
+    mdb_check(mdb_cursor_open(rtxn.raw(), *dbi_opt, &cur), "cursor_open (index count)");
+    struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+
+    MDB_val k = to_val(*key);
+    MDB_val v{0, nullptr};
+    const int rc = mdb_cursor_get(cur, &k, &v, MDB_SET);
+    if (rc == MDB_NOTFOUND) return uint64_t{0};
+    if (rc != MDB_SUCCESS) throw_mdb(rc, "cursor_get (index count)");
+
+    // mdb_cursor_count reports the size of this key's duplicate set without
+    // reading the items - which is what makes the selectivity guard cheap.
+    size_t n = 0;
+    mdb_check(mdb_cursor_count(cur, &n), "cursor_count (index)");
+    return static_cast<uint64_t>(n);
+}
+
+std::optional<std::vector<std::string>>
+LmdbDocumentStore::index_lookup_eq(std::string_view collection,
+                                    const std::string& field,
+                                    const nlohmann::json& value) {
+    ReadTxn rtxn(env_);
+    return lookupIndexEqTxn(rtxn, collection, field, value);
+}
+
+std::optional<std::vector<std::string>>
+LmdbDocumentStore::lookupIndexEqTxn(ReadTxn& rtxn,
+                                     std::string_view collection,
+                                     const std::string& field,
+                                     const nlohmann::json& value) {
+    const auto key = encode_index_key(value);
+    if (!key) return std::nullopt;
+
+    auto dbi_opt = try_open_for_read(rtxn, index_subdb_name(collection, field));
+    if (!dbi_opt) return std::nullopt;
+
+    MDB_cursor* cur = nullptr;
+    mdb_check(mdb_cursor_open(rtxn.raw(), *dbi_opt, &cur), "cursor_open (index)");
+    struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+
+    std::vector<std::string> ids;
+    MDB_val k = to_val(*key);
+    MDB_val v{0, nullptr};
+    int rc = mdb_cursor_get(cur, &k, &v, MDB_SET);
+    if (rc == MDB_NOTFOUND) return ids;      // indexed, genuinely no matches
+    if (rc != MDB_SUCCESS) throw_mdb(rc, "cursor_get (index)");
+
+    rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST_DUP);
+    while (rc == MDB_SUCCESS) {
+        ids.emplace_back(to_sv(v));
+        rc = mdb_cursor_get(cur, &k, &v, MDB_NEXT_DUP);
+    }
+    if (rc != MDB_NOTFOUND) throw_mdb(rc, "cursor next_dup (index)");
+    return ids;
+}
+
+LmdbDocumentStore::IndexPlanStats LmdbDocumentStore::index_plan_stats() const {
+    IndexPlanStats st;
+    st.indexed_scans = indexed_scans_.load(std::memory_order_relaxed);
+    st.full_scans = full_scans_.load(std::memory_order_relaxed);
+    st.declined_unselective =
+        declined_unselective_.load(std::memory_order_relaxed);
+    return st;
+}
+
+void LmdbDocumentStore::reset_index_plan_stats() {
+    indexed_scans_.store(0, std::memory_order_relaxed);
+    full_scans_.store(0, std::memory_order_relaxed);
+    declined_unselective_.store(0, std::memory_order_relaxed);
+}
+
+uint64_t LmdbDocumentStore::build_index(std::string_view collection,
+                                         const std::string& field) {
+    // Walk the collection once and index every row. Idempotent thanks to
+    // MDB_NODUPDATA, so a re-run over an existing index is harmless.
+    std::vector<std::pair<std::string, std::string>> rows;   // id -> payload
+    {
+        ReadTxn rtxn(env_);
+        auto dbi_opt = try_open_for_read(rtxn, collection);
+        if (!dbi_opt) return 0;
+        MDB_cursor* cur = nullptr;
+        mdb_check(mdb_cursor_open(rtxn.raw(), *dbi_opt, &cur), "cursor_open (build)");
+        struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+        MDB_val k{0, nullptr};
+        MDB_val v{0, nullptr};
+        int rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST);
+        while (rc == MDB_SUCCESS) {
+            if (!is_identity_key(to_sv(k))) {
+                rows.emplace_back(std::string(to_sv(k)), std::string(to_sv(v)));
+            }
+            rc = mdb_cursor_get(cur, &k, &v, MDB_NEXT);
+        }
+    }
+
+    const std::string sub = index_subdb_name(collection, field);
+    uint64_t indexed = 0;
+    {
+        WriteTxn wtxn(env_);
+        const unsigned int dbi = open_for_write(wtxn, sub, MDB_DUPSORT);
+        for (const auto& [id, payload] : rows) {
+            yyjson_doc* d = yyjson_read(payload.data(), payload.size(), 0);
+            if (!d) continue;
+            auto val = resolve_from_yyjson(yyjson_doc_get_root(d), field);
+            yyjson_doc_free(d);
+            if (!val) continue;                       // field absent on this row
+            const auto key = encode_index_key(*val);
+            if (!key) continue;                       // unindexable value
+            MDB_val kk = to_val(*key);
+            MDB_val vv = to_val(id);
+            const int rc = mdb_put(wtxn.raw(), dbi, &kk, &vv, MDB_NODUPDATA);
+            if (rc != MDB_SUCCESS && rc != MDB_KEYEXIST) throw_mdb(rc, "index build put");
+            ++indexed;
+        }
+        wtxn.commit();
+        cacheCommittedDbi(sub, dbi);
+    }
+    return indexed;
+}
+
+bool LmdbDocumentStore::drop_index(std::string_view collection,
+                                    const std::string& field) {
+    const std::string sub = index_subdb_name(collection, field);
+    {
+        ReadTxn rtxn(env_);
+        if (!try_open_for_read(rtxn, sub)) return false;
+    }
+    WriteTxn wtxn(env_);
+    const unsigned int dbi = open_for_write(wtxn, sub, MDB_DUPSORT);
+    mdb_check(mdb_drop(wtxn.raw(), dbi, 1), "mdb_drop (index)");
+    wtxn.commit();
+    {
+        std::lock_guard<std::mutex> lock(cache_mutex_);
+        dbi_cache_.erase(sub);
+    }
+    return true;
 }
 
 std::optional<smartbotic::database::Document>
@@ -502,6 +774,20 @@ bool LmdbDocumentStore::del(std::string_view collection, std::string_view id) {
     WriteTxn wtxn(env_);
     unsigned int dbi = open_for_write(wtxn, collection);
     MDB_val k = to_val(id);
+
+    // Remove index entries before the row goes, while its stored bytes are
+    // still readable - they are the only record of which index keys it owns.
+    std::vector<std::pair<std::string, unsigned int>> index_dbis;
+    if (!indexed_fields(collection).empty()) {
+        MDB_val old{0, nullptr};
+        const int grc = mdb_get(wtxn.raw(), dbi, &k, &old);
+        if (grc == MDB_SUCCESS) {
+            maintainIndexes(wtxn, collection, id, to_sv(old), nullptr, index_dbis);
+        } else if (grc != MDB_NOTFOUND) {
+            throw_mdb(grc, "get (pre-index del)");
+        }
+    }
+
     int rc = mdb_del(wtxn.raw(), dbi, &k, nullptr);
     if (rc == MDB_NOTFOUND) {
         wtxn.commit();
@@ -511,6 +797,7 @@ bool LmdbDocumentStore::del(std::string_view collection, std::string_view id) {
     if (rc != MDB_SUCCESS) throw_mdb(rc, "del");
     wtxn.commit();
     cacheCommittedDbi(collection, dbi);
+    for (const auto& [sub, d] : index_dbis) cacheCommittedDbi(sub, d);
     return true;
 }
 
@@ -531,6 +818,70 @@ uint64_t LmdbDocumentStore::count(std::string_view collection) {
     return entries;
 }
 
+
+// v2.9.0 — pick an index for this query, or decline.
+//
+// Returns the candidate document ids when an index can serve one of the EQ
+// predicates cheaply, nullopt to mean "walk the whole collection".
+//
+// THE GUARD. An index is not automatically a win. Reading it costs one random
+// mdb_get + full decode_document per candidate, while the scan costs one yyjson
+// parse per row - and those are not the same price. Measured over the same 10,124
+// rows / 193 MB: full decode 3168ms against 369ms for the parse alone, a ratio of
+// ~8.6x. So the index wins only while
+//
+//     candidates * 8.6  <  total_rows
+//
+// i.e. below roughly 1/8.6 = 11.6% selectivity. kIndexSelectivityDivisor = 10
+// sits just inside that, measured rather than guessed.
+//
+// This matters on real data: on the live `executions` collection
+// status="completed" matches 6,676 of 10,101 rows (66%), so an index on `status`
+// would make that query SLOWER. Declaring an index must not be able to
+// pessimise a query, so the plan is re-decided per value, not per field.
+std::optional<std::vector<std::string>>
+LmdbDocumentStore::planIndexCandidates(ReadTxn& rtxn,
+                                        std::string_view collection,
+                                        const smartbotic::database::Query& query,
+                                        unsigned int coll_dbi) {
+    const auto fields = indexed_fields(collection);
+    if (fields.empty()) return std::nullopt;
+
+    MDB_stat st{};
+    if (mdb_stat(rtxn.raw(), coll_dbi, &st) != MDB_SUCCESS) return std::nullopt;
+    const uint64_t total = st.ms_entries;
+    if (total == 0) return std::nullopt;
+    const uint64_t budget = total / kIndexSelectivityDivisor;
+
+    const std::string* best_field = nullptr;
+    const nlohmann::json* best_value = nullptr;
+    uint64_t best_count = 0;
+    bool found = false;
+
+    for (const auto& f : query.filters) {
+        if (f.op != smartbotic::database::FilterOp::EQ) continue;
+        if (std::find(fields.begin(), fields.end(), f.field) == fields.end()) continue;
+        auto n = countIndexEqTxn(rtxn, collection, f.field, f.value);
+        if (!n) continue;                     // no index, or unindexable value
+        if (!found || *n < best_count) {
+            found = true;
+            best_count = *n;
+            best_field = &f.field;
+            best_value = &f.value;
+        }
+    }
+    if (!found) return std::nullopt;
+
+    // Zero candidates is a legitimate, extremely selective answer - the query
+    // matches nothing and we can say so without reading a single row.
+    if (best_count > budget) {
+        declined_unselective_.fetch_add(1, std::memory_order_relaxed);
+        return std::nullopt;
+    }
+
+    return lookupIndexEqTxn(rtxn, collection, *best_field, *best_value);
+}
+
 ScanResult LmdbDocumentStore::scan(std::string_view collection,
                                     const smartbotic::database::Query& query) {
     ScanResult result;
@@ -647,38 +998,73 @@ ScanResult LmdbDocumentStore::scan(std::string_view collection,
         };
         std::vector<Hit> hits;
 
-        int rc2 = mdb_cursor_get(cursor, &k, &v, MDB_FIRST);
-        while (rc2 == MDB_SUCCESS) {
-            if (!is_identity_key(to_sv(k))) {
-                const auto bytes = to_sv(v);
-                yyjson_doc* ydoc = yyjson_read(bytes.data(), bytes.size(), 0);
-                if (ydoc) {
-                    yyjson_val* root = yyjson_doc_get_root(ydoc);
-                    const bool ok =
-                        smartbotic::db::storage::filter_eval::matchesFiltersResolved(
-                            [root](const std::string& f) {
-                                return resolve_from_yyjson(root, f);
-                            },
-                            query.filters);
-                    if (ok) {
-                        Hit h;
-                        h.id = std::string(to_sv(k));
-                        if (sorting) {
-                            auto kv = resolve_from_yyjson(root, query.sort->field);
-                            if (kv) { h.key = *kv; h.hasKey = true; }
-                        }
-                        hits.push_back(std::move(h));
-                    }
-                    yyjson_doc_free(ydoc);
-                }
+        // One row body, two row sources. The predicate evaluation, sort-key
+        // extraction, sorting and pagination below are shared verbatim between
+        // the indexed and unindexed plans - an index changes only WHICH rows are
+        // visited, never how a row is judged. Two copies of the judging would
+        // drift, and the divergence would surface as wrong rows rather than an
+        // error.
+        const auto consider = [&](std::string_view row_id, std::string_view bytes) {
+            yyjson_doc* ydoc = yyjson_read(bytes.data(), bytes.size(), 0);
+            if (!ydoc) {
                 // A row that will not parse cannot be matched; skipping it is the
                 // same outcome the old path reached by throwing on decode, minus
                 // failing the whole query for one bad row.
+                return;
             }
-            rc2 = mdb_cursor_get(cursor, &k, &v, MDB_NEXT);
+            yyjson_val* root = yyjson_doc_get_root(ydoc);
+            const bool ok =
+                smartbotic::db::storage::filter_eval::matchesFiltersResolved(
+                    [root](const std::string& f) {
+                        return resolve_from_yyjson(root, f);
+                    },
+                    query.filters);
+            if (ok) {
+                Hit h;
+                h.id = std::string(row_id);
+                if (sorting) {
+                    auto kv = resolve_from_yyjson(root, query.sort->field);
+                    if (kv) { h.key = *kv; h.hasKey = true; }
+                }
+                hits.push_back(std::move(h));
+            }
+            yyjson_doc_free(ydoc);
+        };
+
+        auto candidates = planIndexCandidates(rtxn, collection, query, *dbi_opt);
+        if (candidates) {
+            indexed_scans_.fetch_add(1, std::memory_order_relaxed);
+        } else {
+            full_scans_.fetch_add(1, std::memory_order_relaxed);
         }
-        if (rc2 != MDB_NOTFOUND && rc2 != MDB_SUCCESS) {
-            throw_mdb(rc2, "cursor_get");
+
+        if (candidates) {
+            // Indexed plan: visit only the rows the index named. Every filter is
+            // still applied to each one, so the index is allowed to be a
+            // superset - it narrows work, it does not decide the answer.
+            for (const auto& cid : *candidates) {
+                MDB_val ck = to_val(cid);
+                MDB_val cv{0, nullptr};
+                if (mdb_get(rtxn.raw(), *dbi_opt, &ck, &cv) != MDB_SUCCESS) {
+                    // A posting with no row means the index is ahead of the
+                    // collection, which the same-transaction maintenance is meant
+                    // to prevent. Skipping is the safe reading: report what
+                    // exists, never a row that does not.
+                    continue;
+                }
+                consider(cid, to_sv(cv));
+            }
+        } else {
+            int rc2 = mdb_cursor_get(cursor, &k, &v, MDB_FIRST);
+            while (rc2 == MDB_SUCCESS) {
+                if (!is_identity_key(to_sv(k))) {
+                    consider(to_sv(k), to_sv(v));
+                }
+                rc2 = mdb_cursor_get(cursor, &k, &v, MDB_NEXT);
+            }
+            if (rc2 != MDB_NOTFOUND && rc2 != MDB_SUCCESS) {
+                throw_mdb(rc2, "cursor_get");
+            }
         }
 
         if (sorting) {

+ 86 - 1
service/src/storage/document_store_lmdb.hpp

@@ -12,6 +12,7 @@
 
 #pragma once
 
+#include <atomic>
 #include <mutex>
 #include <optional>
 #include <string>
@@ -19,6 +20,8 @@
 #include <unordered_map>
 #include <vector>
 
+#include <nlohmann/json.hpp>
+
 #include "storage/document_store.hpp"
 
 namespace smartbotic::db::storage {
@@ -46,6 +49,48 @@ public:
     // handle when a sub-db is already open.
     size_t prime_dbi_cache();
 
+    // v2.9.0 — declare which fields of `collection` carry a secondary index.
+    // Empty removes the declaration. Maintenance happens inside the same write
+    // transaction as the document, so an aborted write cannot leave a stale
+    // index. Collections with no declared fields pay nothing: the write path
+    // checks this map and returns immediately.
+    void set_indexed_fields(std::string_view collection,
+                            std::vector<std::string> fields);
+    std::vector<std::string> indexed_fields(std::string_view collection);
+
+    // Doc ids whose `field` equals `value`, straight from the index. nullopt
+    // means "no index for this field" - the caller must fall back to a scan
+    // rather than treat it as "no matches", which would be silently wrong.
+    std::optional<std::vector<std::string>>
+    index_lookup_eq(std::string_view collection,
+                    const std::string& field,
+                    const nlohmann::json& value);
+
+    // Number of ids the index holds for one exact value, without reading them.
+    // Used as the selectivity guard: an index over a low-cardinality field
+    // (status="completed" matches 66% of rows) is slower than scanning.
+    std::optional<uint64_t> index_count_eq(std::string_view collection,
+                                            const std::string& field,
+                                            const nlohmann::json& value);
+
+    // Observability: how many scans used an index plan versus walked the
+    // collection. An operator needs to know whether a declared index is
+    // actually being used, and a test needs it to prove the plan was taken
+    // rather than silently declined.
+    struct IndexPlanStats {
+        uint64_t indexed_scans = 0;     // an index served the query
+        uint64_t full_scans = 0;        // walked the collection
+        uint64_t declined_unselective = 0;  // an index existed but was too broad
+    };
+    IndexPlanStats index_plan_stats() const;
+    void reset_index_plan_stats();
+
+    // Populate an index over the rows already present. Returns rows indexed.
+    uint64_t build_index(std::string_view collection, const std::string& field);
+
+    // Remove an index entirely.
+    bool drop_index(std::string_view collection, const std::string& field);
+
     // Outcome of the constructor's priming pass, for the owner to report.
     // prime_error() is empty on success.
     size_t primed_count() const noexcept { return primed_count_; }
@@ -97,6 +142,14 @@ private:
     std::mutex cache_mutex_;
     size_t primed_count_ = 0;
     std::string prime_error_;
+
+    // collection -> declared indexed fields. Its own mutex: the write path
+    // consults it on every put and must not contend with dbi cache lookups.
+    std::mutex index_mutex_;
+    std::unordered_map<std::string, std::vector<std::string>> indexed_fields_;
+    mutable std::atomic<uint64_t> indexed_scans_{0};
+    mutable std::atomic<uint64_t> full_scans_{0};
+    mutable std::atomic<uint64_t> declined_unselective_{0};
     std::unordered_map<std::string, unsigned int> dbi_cache_;
 
     // Resolve a collection name to an MDB_dbi handle.
@@ -107,7 +160,39 @@ private:
     // try_open_for_read: opens the sub-db within the provided read txn IF
     // it already exists. Returns nullopt if the sub-db doesn't exist —
     // because read txns can't MDB_CREATE.
-    unsigned int open_for_write(class WriteTxn& wtxn, std::string_view collection);
+    // `extra_flags` is OR-ed into the mdb_dbi_open flags. Index sub-dbs pass
+    // MDB_DUPSORT so one value key holds a sorted set of document ids.
+    unsigned int open_for_write(class WriteTxn& wtxn, std::string_view collection,
+                                unsigned int extra_flags = 0);
+
+    // v2.9.0 — bring every declared index for `collection` into line with a
+    // write, inside `wtxn`. old_payload is the stored bytes being replaced
+    // (empty on insert); new_doc is null on delete. Handles opened here are
+    // appended to `to_cache` for the caller to cache AFTER its commit.
+    // Selectivity guard: use an index only while candidates <= rows/this.
+    // Derived from measurement - full decode_document costs ~8.6x a yyjson parse
+    // of the same row, so an index stops paying above ~11.6% selectivity.
+    static constexpr uint64_t kIndexSelectivityDivisor = 10;
+
+    std::optional<std::vector<std::string>>
+    lookupIndexEqTxn(class ReadTxn& rtxn, std::string_view collection,
+                     const std::string& field, const nlohmann::json& value);
+    std::optional<uint64_t>
+    countIndexEqTxn(class ReadTxn& rtxn, std::string_view collection,
+                    const std::string& field, const nlohmann::json& value);
+
+    // Choose an index for a query, or nullopt to walk the whole collection.
+    std::optional<std::vector<std::string>>
+    planIndexCandidates(class ReadTxn& rtxn, std::string_view collection,
+                        const smartbotic::database::Query& query,
+                        unsigned int coll_dbi);
+
+    void maintainIndexes(class WriteTxn& wtxn,
+                         std::string_view collection,
+                         std::string_view id,
+                         std::string_view old_payload,
+                         const smartbotic::database::Document* new_doc,
+                         std::vector<std::pair<std::string, unsigned int>>& to_cache);
 
     // v2.8.0 — record a handle in the cache, to be called ONLY after the
     // transaction that opened it has committed. LMDB closes a handle whose

+ 149 - 0
service/src/storage/secondary_index.cpp

@@ -0,0 +1,149 @@
+// v2.9.0 secondary indexes — key encoding and sub-db naming.
+
+#include "storage/secondary_index.hpp"
+
+#include <cmath>
+#include <cstdint>
+#include <cstring>
+#include <limits>
+
+namespace smartbotic::db::storage {
+
+namespace {
+
+constexpr char kSep = '#';
+
+constexpr unsigned char kTagNull = 0x10;
+constexpr unsigned char kTagFalse = 0x20;
+constexpr unsigned char kTagTrue = 0x21;
+constexpr unsigned char kTagInt = 0x30;
+constexpr unsigned char kTagReal = 0x31;
+constexpr unsigned char kTagStr = 0x40;
+
+// Append `v` big-endian so that memcmp order matches unsigned integer order.
+void append_be64(std::string& out, uint64_t v) {
+    for (int shift = 56; shift >= 0; shift -= 8) {
+        out.push_back(static_cast<char>((v >> shift) & 0xFF));
+    }
+}
+
+std::string encode_int(int64_t v) {
+    std::string out;
+    out.reserve(9);
+    out.push_back(static_cast<char>(kTagInt));
+    // Flipping the sign bit maps the signed range onto unsigned while
+    // preserving order: INT64_MIN -> 0, -1 -> 0x7FFF..., 0 -> 0x8000..., etc.
+    const uint64_t biased =
+        static_cast<uint64_t>(v) ^ (uint64_t{1} << 63);
+    append_be64(out, biased);
+    return out;
+}
+
+std::string encode_real(double d) {
+    std::string out;
+    out.reserve(9);
+    out.push_back(static_cast<char>(kTagReal));
+    uint64_t bits = 0;
+    static_assert(sizeof(double) == sizeof(uint64_t), "IEEE754 double expected");
+    std::memcpy(&bits, &d, sizeof(bits));
+    // The standard order-preserving transform for IEEE754: for negatives
+    // (sign bit set) invert every bit, which reverses their descending
+    // magnitude order; for non-negatives just set the sign bit so they sort
+    // above all negatives.
+    if (bits & (uint64_t{1} << 63)) {
+        bits = ~bits;
+    } else {
+        bits |= (uint64_t{1} << 63);
+    }
+    append_be64(out, bits);
+    return out;
+}
+
+}  // namespace
+
+std::string index_subdb_name(std::string_view collection, std::string_view field) {
+    std::string out;
+    out.reserve(kIndexSubdbPrefix.size() + collection.size() + 1 + field.size());
+    out.append(kIndexSubdbPrefix);
+    out.append(collection);
+    out.push_back(kSep);
+    out.append(field);
+    return out;
+}
+
+bool is_index_subdb(std::string_view name) {
+    return name.size() > kIndexSubdbPrefix.size() &&
+           name.compare(0, kIndexSubdbPrefix.size(), kIndexSubdbPrefix) == 0;
+}
+
+std::optional<IndexSubdb> parse_index_subdb(std::string_view name) {
+    if (!is_index_subdb(name)) return std::nullopt;
+    std::string_view rest = name.substr(kIndexSubdbPrefix.size());
+    const size_t sep = rest.find(kSep);
+    if (sep == std::string_view::npos) return std::nullopt;
+    IndexSubdb out;
+    out.collection = std::string(rest.substr(0, sep));
+    out.field = std::string(rest.substr(sep + 1));
+    if (out.collection.empty() || out.field.empty()) return std::nullopt;
+    return out;
+}
+
+std::optional<std::string> encode_index_key(const nlohmann::json& value) {
+    switch (value.type()) {
+        case nlohmann::json::value_t::null:
+            return std::string(1, static_cast<char>(kTagNull));
+
+        case nlohmann::json::value_t::boolean:
+            return std::string(
+                1, static_cast<char>(value.get<bool>() ? kTagTrue : kTagFalse));
+
+        case nlohmann::json::value_t::string: {
+            std::string out;
+            const auto& s = value.get_ref<const nlohmann::json::string_t&>();
+            out.reserve(1 + s.size());
+            out.push_back(static_cast<char>(kTagStr));
+            out.append(s);
+            return out;
+        }
+
+        case nlohmann::json::value_t::number_integer:
+            return encode_int(value.get<int64_t>());
+
+        case nlohmann::json::value_t::number_unsigned: {
+            const uint64_t u = value.get<uint64_t>();
+            // Above INT64_MAX there is no int64 form, so such a value can only
+            // be compared against another value stored the same way. Encode it
+            // as a real, which is what nlohmann's own comparison degrades to.
+            if (u > static_cast<uint64_t>(std::numeric_limits<int64_t>::max())) {
+                return encode_real(static_cast<double>(u));
+            }
+            return encode_int(static_cast<int64_t>(u));
+        }
+
+        case nlohmann::json::value_t::number_float: {
+            const double d = value.get<double>();
+            if (std::isnan(d)) return std::nullopt;   // never equals anything
+            // CANONICALISATION - the invariant in the header. nlohmann compares
+            // numbers across subtypes, so 5.0 must produce the same key as 5 or
+            // an indexed EQ would miss rows the scan finds.
+            if (std::isfinite(d) && d == std::floor(d) &&
+                d >= static_cast<double>(std::numeric_limits<int64_t>::min()) &&
+                d <= static_cast<double>(std::numeric_limits<int64_t>::max())) {
+                return encode_int(static_cast<int64_t>(d));
+            }
+            return encode_real(d);
+        }
+
+        // A whole object or array is not a lookup key. CONTAINS asks about an
+        // element rather than the container, so it needs its own index shape and
+        // is not served from here.
+        case nlohmann::json::value_t::object:
+        case nlohmann::json::value_t::array:
+        case nlohmann::json::value_t::binary:
+        case nlohmann::json::value_t::discarded:
+        default:
+            return std::nullopt;
+    }
+}
+
+}  // namespace smartbotic::db::storage

+ 87 - 0
service/src/storage/secondary_index.hpp

@@ -0,0 +1,87 @@
+// v2.9.0 secondary indexes — key encoding and sub-db naming.
+//
+// A filtered query is otherwise a full collection scan. v2.8.0 removed document
+// materialisation from that scan (2463ms -> 402ms on a 505 MB collection), but
+// what remains is yyjson parsing every value in the collection to test the
+// predicate, and that is the floor for a scan. Getting below it means not
+// reading non-matching rows at all.
+//
+// Layout: one LMDB sub-db per (collection, field), opened MDB_DUPSORT.
+//   key  = encoded field value  (see encode_index_key)
+//   data = document id
+// DUPSORT stores the ids for one value as a sorted, deduplicated set, which is
+// exactly "the posting list for this value" and costs no extra structure.
+//
+// ⚠ THE INVARIANT THIS FILE EXISTS TO PROTECT
+//
+// An index lookup must return exactly the rows a scan would return. The scan
+// compares with `compareJson`, whose EQ is `nlohmann::json::operator==` - and
+// that compares numbers ACROSS subtypes: json(5) == json(5.0) is true. So if
+// this encoder gave 5 and 5.0 different keys, an indexed EQ would silently miss
+// rows that the unindexed path finds. Two code paths answering one question must
+// agree, and when they do not the failure is wrong data, not an error.
+//
+// encode_index_key therefore CANONICALISES: any number whose value is integral
+// and fits int64 encodes to the integer form regardless of how it was stored.
+//
+// Known limitation, deliberate: integral and non-integral numbers land under
+// different type tags, so they order independently. That is fine for EQ (the
+// only operator wired to indexes today) but means a range query spanning both
+// cannot be served from an index yet. Ranges must keep using the scan until the
+// encoding is unified.
+
+#pragma once
+
+#include <optional>
+#include <string>
+#include <string_view>
+#include <vector>
+
+#include <nlohmann/json.hpp>
+
+namespace smartbotic::db::storage {
+
+// Prefix marking a sub-db as an index. Leading '_' means is_system_subdb()
+// already treats these as internal, so they stay out of list_collections().
+inline constexpr std::string_view kIndexSubdbPrefix = "_idx_";
+
+// Sub-db name for one indexed field.
+//
+// The field is appended after a '#' separator, which cannot appear in a
+// collection name (validated elsewhere) so "<coll>#<field>" is unambiguous even
+// though a field name may contain dots for a nested path.
+std::string index_subdb_name(std::string_view collection, std::string_view field);
+
+// True if a sub-db name is an index. Used to skip them when walking sub-dbs and
+// to open them with MDB_DUPSORT.
+bool is_index_subdb(std::string_view name);
+
+// Parse an index sub-db name back into (collection, field). nullopt if `name`
+// is not an index name.
+struct IndexSubdb {
+    std::string collection;
+    std::string field;
+};
+std::optional<IndexSubdb> parse_index_subdb(std::string_view name);
+
+// Encode a JSON scalar into an order-preserving index key.
+//
+// Returns nullopt for values that must not be indexed:
+//   - objects and arrays (a whole subtree is not a lookup key)
+//   - NaN (never equal to anything, so an entry could never be found anyway)
+// A row whose value does not encode simply has no index entry; queries that
+// would need it must fall back to a scan, which is why callers must never treat
+// "no index entry" as "no such row" for anything but exact-value lookups.
+//
+// Encoding is a one-byte type tag followed by the payload, chosen so that
+// memcmp order matches value order WITHIN a tag:
+//   0x10  null      (no payload)
+//   0x20  false     (no payload)
+//   0x21  true      (no payload)
+//   0x30  integer   8 bytes, big-endian int64 with the sign bit flipped
+//   0x31  real      8 bytes, big-endian IEEE754 with the standard
+//                   order-preserving transform
+//   0x40  string    raw UTF-8 bytes
+std::optional<std::string> encode_index_key(const nlohmann::json& value);
+
+}  // namespace smartbotic::db::storage

+ 29 - 0
tests/CMakeLists.txt

@@ -353,6 +353,7 @@ add_executable(test_document_store
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/lmdb_txn.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/lmdb_dbi.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/document_store_lmdb.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/secondary_index.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/subdb_identity.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/json_parse.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/doc_binary.cpp
@@ -391,6 +392,7 @@ add_executable(test_migrate_v1_to_v2
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/lmdb_txn.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/lmdb_dbi.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/document_store_lmdb.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/secondary_index.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/subdb_identity.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/migrate_v1_to_v2.cpp
 )
@@ -446,6 +448,7 @@ add_executable(test_dual_write_mirror
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/lmdb_txn.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/lmdb_dbi.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/document_store_lmdb.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/secondary_index.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/subdb_identity.cpp
 )
 
@@ -530,6 +533,7 @@ add_executable(test_subdb_identity
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/lmdb_dbi.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/subdb_identity.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/document_store_lmdb.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/secondary_index.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/json_parse.cpp
     ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/doc_binary.cpp
 )
@@ -551,6 +555,31 @@ endif()
 
 add_test(NAME test_subdb_identity COMMAND test_subdb_identity)
 
+# v2.9.0 — secondary index key encoding. Pins the one property that matters:
+# an index key comparison must mean the same thing as the scan's comparison,
+# because two paths answering one question that disagree return wrong data
+# rather than an error.
+add_executable(test_secondary_index_keys
+    test_secondary_index_keys.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/storage/secondary_index.cpp
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src/doc_binary.cpp
+)
+
+target_include_directories(test_secondary_index_keys PRIVATE
+    ${CMAKE_CURRENT_SOURCE_DIR}/../service/src
+    ${yyjson_INCLUDE_DIRS}
+)
+
+target_link_libraries(test_secondary_index_keys PRIVATE ${yyjson_LIBRARIES})
+
+if(TARGET nlohmann_json::nlohmann_json)
+    target_link_libraries(test_secondary_index_keys PRIVATE nlohmann_json::nlohmann_json)
+else()
+    target_include_directories(test_secondary_index_keys PRIVATE ${NLOHMANN_JSON_INCLUDE_DIRS})
+endif()
+
+add_test(NAME test_secondary_index_keys COMMAND test_secondary_index_keys)
+
 # v2.6.0 — project-scoped file storage. Covers cross-project id isolation and
 # the per-project deduplicated flag (global dedup would otherwise be a
 # cross-project existence oracle).

+ 165 - 0
tests/test_secondary_index_keys.cpp

@@ -0,0 +1,165 @@
+// v2.9.0 — index key encoding must agree with the scan's comparison semantics.
+//
+// An indexed lookup and an unindexed scan answer the same question by different
+// routes. If they disagree the result is silently wrong data, not an error, so
+// the agreement is the property worth testing rather than the encoding's
+// internal shape.
+//
+// The specific trap: the scan's EQ is nlohmann::json::operator==, which compares
+// numbers ACROSS subtypes - json(5) == json(5.0) is true. An encoder that gave
+// those different keys would make an indexed EQ miss rows the scan finds.
+
+#include <cmath>
+#include <iostream>
+#include <limits>
+#include <string>
+#include <vector>
+
+#include <nlohmann/json.hpp>
+
+#include "storage/filter_eval.hpp"
+#include "storage/secondary_index.hpp"
+
+using namespace smartbotic::db::storage;
+using smartbotic::database::FilterOp;
+
+namespace {
+
+int g_pass = 0;
+int g_fail = 0;
+
+void check(bool cond, const std::string& msg) {
+    if (cond) { ++g_pass; }
+    else { ++g_fail; std::cerr << "FAIL: " << msg << "\n"; }
+}
+
+// THE headline property. For every ordered pair in a representative corpus,
+// "the keys are equal" must mean the same thing as "the scan considers them
+// equal".
+void test_key_equality_matches_scan_equality() {
+    const std::vector<nlohmann::json> corpus = {
+        nullptr,
+        false, true,
+        0, 1, -1, 5, -5,
+        0.0, 1.0, 5.0, -5.0,          // integral floats - must match their ints
+        5.5, -5.5, 0.5,
+        std::numeric_limits<int64_t>::max(),
+        std::numeric_limits<int64_t>::min(),
+        "", "a", "b", "5", "abc",
+        1786263002080195076LL,        // an ns timestamp: > 2^53, needs exactness
+        1786263002080195077LL,        // its neighbour, must not collide
+    };
+
+    int compared = 0;
+    for (const auto& a : corpus) {
+        for (const auto& b : corpus) {
+            const auto ka = encode_index_key(a);
+            const auto kb = encode_index_key(b);
+            if (!ka || !kb) continue;          // unindexable, scan handles it
+            const bool keys_equal = (*ka == *kb);
+            const bool scan_equal = smartbotic::db::storage::filter_eval::compareJson(
+                a, FilterOp::EQ, b);
+            ++compared;
+            if (keys_equal != scan_equal) {
+                check(false, "key equality disagrees with scan EQ for " +
+                             a.dump() + " vs " + b.dump() +
+                             " (keys_equal=" + (keys_equal ? "1" : "0") +
+                             " scan_equal=" + (scan_equal ? "1" : "0") + ")");
+                return;
+            }
+        }
+    }
+    check(compared > 300, "compared a meaningful number of pairs");
+    check(true, "key equality matches scan EQ across the whole corpus");
+}
+
+// The canonicalisation that makes the above hold, stated directly so a
+// regression names itself.
+void test_integral_floats_encode_as_integers() {
+    check(encode_index_key(5) == encode_index_key(5.0),
+          "5 and 5.0 produce the same key - nlohmann considers them equal");
+    check(encode_index_key(-5) == encode_index_key(-5.0),
+          "negative integral floats canonicalise too");
+    check(encode_index_key(0) == encode_index_key(0.0), "zero canonicalises");
+    check(encode_index_key(5) != encode_index_key(5.5),
+          "a genuinely different number keeps a different key");
+}
+
+// Large integers must stay exact. Nanosecond timestamps exceed 2^53, so an
+// encoder that routed everything through double would collide neighbours - and
+// _created_at is stored in ns by default since v2.2.0.
+void test_large_integers_are_exact() {
+    const int64_t a = 1786263002080195076LL;
+    const int64_t b = 1786263002080195077LL;
+    check(static_cast<double>(a) == static_cast<double>(b),
+          "precondition: these two DO collide as doubles, which is the hazard");
+    check(encode_index_key(a) != encode_index_key(b),
+          "but their index keys differ - integers are not routed through double");
+}
+
+// memcmp order must match value order within a type, which is what makes a
+// cursor range walk possible later.
+void test_ordering_is_preserved() {
+    auto key = [](const nlohmann::json& j) { return *encode_index_key(j); };
+
+    check(key(-5) < key(-1), "negative integers ascend");
+    check(key(-1) < key(0), "negatives sort below zero");
+    check(key(0) < key(1), "zero below positives");
+    check(key(1) < key(5), "positive integers ascend");
+    check(key(std::numeric_limits<int64_t>::min()) < key(int64_t{0}),
+          "INT64_MIN is the smallest integer key");
+    check(key(int64_t{0}) < key(std::numeric_limits<int64_t>::max()),
+          "INT64_MAX is the largest");
+
+    check(key(-5.5) < key(-0.5), "negative reals ascend");
+    check(key(-0.5) < key(0.5), "negative reals sort below positive reals");
+    check(key(0.5) < key(5.5), "positive reals ascend");
+
+    check(key("a") < key("b"), "strings sort lexicographically");
+    check(key("") < key("a"), "the empty string sorts first");
+}
+
+void test_unindexable_values() {
+    check(!encode_index_key(nlohmann::json::object()).has_value(),
+          "an object is not a lookup key");
+    check(!encode_index_key(nlohmann::json::array({1, 2})).has_value(),
+          "an array is not a lookup key - CONTAINS needs a different shape");
+    check(!encode_index_key(std::numeric_limits<double>::quiet_NaN()).has_value(),
+          "NaN is never equal to anything, so it gets no entry");
+    check(encode_index_key(nullptr).has_value(),
+          "null IS indexable - it is a real value a filter can ask for");
+}
+
+void test_subdb_naming_roundtrip() {
+    const std::string n = index_subdb_name("executions", "workflowId");
+    check(is_index_subdb(n), "an index name is recognised as one");
+    check(n[0] == '_', "index sub-dbs start with '_' so they count as system "
+                       "and stay out of list_collections()");
+    auto parsed = parse_index_subdb(n);
+    check(parsed.has_value(), "the name parses back");
+    check(parsed && parsed->collection == "executions", "collection round-trips");
+    check(parsed && parsed->field == "workflowId", "field round-trips");
+
+    // A dotted path is a legal field name and must survive.
+    auto dotted = parse_index_subdb(index_subdb_name("c", "a.b.c"));
+    check(dotted && dotted->field == "a.b.c", "a nested field path round-trips");
+
+    check(!is_index_subdb("executions"), "an ordinary collection is not an index");
+    check(!is_index_subdb("_vectors_executions"), "nor is a vector sidecar");
+    check(!parse_index_subdb("_idx_noseparator").has_value(),
+          "a malformed index name does not parse");
+}
+
+}  // namespace
+
+int main() {
+    std::cout << "=== test_secondary_index_keys ===\n";
+    test_key_equality_matches_scan_equality();
+    test_integral_floats_encode_as_integers();
+    test_large_integers_are_exact();
+    test_ordering_is_preserved();
+    test_unindexable_values();
+    test_subdb_naming_roundtrip();
+    std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
+    return g_fail == 0 ? 0 : 1;
+}

+ 318 - 0
tests/test_subdb_identity.cpp

@@ -695,6 +695,320 @@ void test_existing_collection_readable_without_writing_first() {
     check(reader.count("nosuch") == 0, "and counts zero");
 }
 
+
+// v2.9.0 — a secondary index must stay exactly in step with the documents.
+//
+// The index is a second copy of a fact already stored in the row. Every way the
+// two can diverge is a silent-wrong-data bug: a stale entry returns a row that
+// no longer matches, a missing entry hides a row that does. So this walks the
+// full lifecycle - insert, update the indexed field, update something else,
+// delete, re-insert - and after every step asserts the index agrees with a
+// brute-force scan of the collection.
+void test_index_tracks_documents_through_every_write() {
+    TmpEnv t("idx-maint");
+    LmdbDocumentStore store(t.env);
+    store.set_indexed_fields("execs", {"workflowId"});
+
+    auto put = [&](const std::string& id, const std::string& wf, int n) {
+        Document d;
+        d.id = id;
+        d.collection = "execs";
+        d.set_data(nlohmann::json{{"workflowId", wf}, {"n", n}});
+        store.put("execs", id, d);
+    };
+
+    // What the index SHOULD say, computed by scanning every row - the oracle.
+    auto truth = [&](const std::string& wf) {
+        smartbotic::database::Query q;
+        q.limit = 10000;
+        std::vector<std::string> ids;
+        for (const auto& d : store.scan("execs", q).documents) {
+            if (d.data().value("workflowId", std::string{}) == wf) ids.push_back(d.id);
+        }
+        std::sort(ids.begin(), ids.end());
+        return ids;
+    };
+    auto indexed = [&](const std::string& wf) {
+        auto got = store.index_lookup_eq("execs", "workflowId", nlohmann::json(wf));
+        std::vector<std::string> ids = got.value_or(std::vector<std::string>{});
+        std::sort(ids.begin(), ids.end());
+        return ids;
+    };
+    auto agree = [&](const std::string& wf, const char* stage) {
+        const bool ok = indexed(wf) == truth(wf);
+        const std::string msg = "index agrees with a full scan for " + wf +
+                                " after " + stage;
+        check(ok, msg.c_str());
+    };
+
+    // Insert
+    put("e1", "wf-a", 1);
+    put("e2", "wf-a", 2);
+    put("e3", "wf-b", 3);
+    agree("wf-a", "inserts");
+    agree("wf-b", "inserts");
+    check(indexed("wf-a").size() == 2, "two rows under wf-a");
+
+    // Update the INDEXED field: the old posting must go, the new one appear.
+    put("e2", "wf-b", 2);
+    agree("wf-a", "moving e2 to wf-b");
+    agree("wf-b", "moving e2 to wf-b");
+    check(indexed("wf-a").size() == 1, "wf-a lost e2");
+    check(indexed("wf-b").size() == 2, "wf-b gained it");
+
+    // Update an UNindexed field: the index must be untouched, not duplicated.
+    put("e1", "wf-a", 99);
+    agree("wf-a", "updating an unindexed field");
+    check(indexed("wf-a").size() == 1,
+          "no duplicate posting from re-writing the same indexed value");
+
+    // Delete
+    check(store.del("execs", "e3"), "deleted e3");
+    agree("wf-b", "deleting e3");
+    check(indexed("wf-b").size() == 1, "e3 is gone from the index");
+
+    // Re-insert the same id
+    put("e3", "wf-b", 7);
+    agree("wf-b", "re-inserting e3");
+    check(indexed("wf-b").size() == 2, "e3 is back exactly once");
+
+    // A value with no rows is an empty result, NOT "no index".
+    auto none = store.index_lookup_eq("execs", "workflowId", nlohmann::json("wf-zzz"));
+    check(none.has_value() && none->empty(),
+          "an indexed field with no matching rows returns empty, not nullopt - "
+          "nullopt means 'no index' and would send the caller to a scan");
+
+    // An undeclared field has no index, and must say so rather than say 'none'.
+    auto unindexed = store.index_lookup_eq("execs", "n", nlohmann::json(1));
+    check(!unindexed.has_value(),
+          "an unindexed field returns nullopt so the caller falls back to a scan "
+          "instead of concluding there are no matches");
+}
+
+// A collection with no declared index must behave exactly as before, and pay
+// nothing. Also: declaring an index later must pick up the rows already there.
+void test_build_index_over_existing_rows() {
+    TmpEnv t("idx-build");
+    LmdbDocumentStore store(t.env);
+
+    // Write BEFORE declaring the index.
+    for (int i = 0; i < 20; ++i) {
+        Document d;
+        d.id = "d" + std::to_string(i);
+        d.collection = "c";
+        d.set_data(nlohmann::json{{"grp", i % 4 == 0 ? "hot" : "cold"}});
+        store.put("c", d.id, d);
+    }
+    check(!store.index_lookup_eq("c", "grp", nlohmann::json("hot")).has_value(),
+          "no index exists before it is declared");
+
+    const uint64_t built = store.build_index("c", "grp");
+    check(built == 20, "the backfill indexed every existing row");
+    store.set_indexed_fields("c", {"grp"});
+
+    auto hot = store.index_lookup_eq("c", "grp", nlohmann::json("hot"));
+    check(hot.has_value() && hot->size() == 5,
+          "the backfilled index finds the pre-existing rows (d0,d4,d8,d12,d16)");
+
+    // Idempotent: a second build must not double the postings.
+    store.build_index("c", "grp");
+    hot = store.index_lookup_eq("c", "grp", nlohmann::json("hot"));
+    check(hot.has_value() && hot->size() == 5,
+          "re-running the backfill does not duplicate postings");
+
+    // The count guard must see the same number without reading the ids.
+    auto n = store.index_count_eq("c", "grp", nlohmann::json("hot"));
+    check(n.has_value() && *n == 5, "index_count_eq agrees with the lookup");
+    auto cold = store.index_count_eq("c", "grp", nlohmann::json("cold"));
+    check(cold.has_value() && *cold == 15, "and counts the larger group");
+
+    check(store.drop_index("c", "grp"), "the index drops");
+    check(!store.index_lookup_eq("c", "grp", nlohmann::json("hot")).has_value(),
+          "after dropping, lookups report no index rather than no rows");
+}
+
+// Numbers are where an index most easily disagrees with a scan, because the scan
+// compares across numeric subtypes. A doc stored with 5 must be found by a
+// filter asking for 5.0 through EITHER path.
+void test_index_numeric_equality_matches_scan() {
+    TmpEnv t("idx-num");
+    LmdbDocumentStore store(t.env);
+    store.set_indexed_fields("m", {"code"});
+
+    Document a;
+    a.id = "a"; a.collection = "m";
+    a.set_data(nlohmann::json{{"code", 5}});          // integer
+    store.put("m", "a", a);
+
+    Document b;
+    b.id = "b"; b.collection = "m";
+    b.set_data(nlohmann::json{{"code", 5.0}});        // integral double
+    store.put("m", "b", b);
+
+    auto by_int = store.index_lookup_eq("m", "code", nlohmann::json(5));
+    auto by_dbl = store.index_lookup_eq("m", "code", nlohmann::json(5.0));
+    check(by_int.has_value() && by_int->size() == 2,
+          "asking for 5 finds BOTH the int and the integral-double row");
+    check(by_dbl == by_int,
+          "and asking for 5.0 returns exactly the same rows - the scan's EQ "
+          "compares numbers across subtypes, so the index must too");
+
+    // A big integer must not collide with its neighbour via double precision.
+    Document c;
+    c.id = "c"; c.collection = "m";
+    c.set_data(nlohmann::json{{"code", 1786263002080195076LL}});
+    store.put("m", "c", c);
+    auto near = store.index_lookup_eq("m", "code",
+                                       nlohmann::json(1786263002080195077LL));
+    check(near.has_value() && near->empty(),
+          "a neighbouring ns-scale integer does not collide - these two ARE "
+          "equal as doubles, so routing through double would false-match");
+}
+
+
+// v2.9.0 — an indexed plan must return EXACTLY what the unindexed plan returns.
+//
+// This is the whole safety argument for indexing. The index is an optimisation,
+// so any observable difference is a bug, and the interesting failures are silent:
+// a missing posting drops a row, a stale one adds a row that no longer matches,
+// and a different code path can disagree about ordering or total_matched.
+//
+// The test runs each query twice against the same data - once with the field
+// declared indexed, once not - and compares the complete result: ids in order,
+// total_matched, and has_more.
+void test_indexed_and_unindexed_plans_agree() {
+    TmpEnv t("idx-equiv");
+    LmdbDocumentStore store(t.env);
+
+    // 300 rows: `grp` is selective enough to use the index (10 groups of 30 =
+    // 10%), `bucket` deliberately is NOT (2 values, 50% each) so the guard has
+    // something to decline.
+    for (int i = 0; i < 300; ++i) {
+        Document d;
+        d.id = "r" + std::string(i < 10 ? "00" : (i < 100 ? "0" : "")) +
+               std::to_string(i);
+        d.collection = "c";
+        d.set_data(nlohmann::json{
+            {"grp", "g" + std::to_string(i % 10)},
+            {"bucket", (i % 2 == 0) ? "even" : "odd"},
+            {"n", i},
+            {"nest", {{"deep", "d" + std::to_string(i % 10)}}},
+        });
+        store.put("c", d.id, d);
+    }
+
+    using Op = smartbotic::database::FilterOp;
+    struct Case {
+        const char* name;
+        std::vector<smartbotic::database::Filter> filters;
+        std::optional<smartbotic::database::Sort> sort;
+        uint32_t limit;
+        uint32_t offset;
+    };
+    auto F = [](const char* f, Op op, const nlohmann::json& v) {
+        smartbotic::database::Filter x;
+        x.field = f; x.op = op; x.value = v;
+        return x;
+    };
+
+    const std::vector<Case> cases = {
+        {"eq indexed field",            {F("grp", Op::EQ, "g3")},                     std::nullopt, 100, 0},
+        {"eq + second predicate",       {F("grp", Op::EQ, "g3"), F("bucket", Op::EQ, "even")}, std::nullopt, 100, 0},
+        {"eq + range on another field", {F("grp", Op::EQ, "g3"), F("n", Op::GT, 100)}, std::nullopt, 100, 0},
+        {"eq with sort asc",            {F("grp", Op::EQ, "g3")},                     smartbotic::database::Sort{"n", false}, 100, 0},
+        {"eq with sort desc",           {F("grp", Op::EQ, "g3")},                     smartbotic::database::Sort{"n", true}, 100, 0},
+        {"eq paginated",                {F("grp", Op::EQ, "g3")},                     smartbotic::database::Sort{"n", false}, 7, 10},
+        {"eq offset past end",          {F("grp", Op::EQ, "g3")},                     std::nullopt, 10, 999},
+        {"eq matching nothing",         {F("grp", Op::EQ, "nope")},                   std::nullopt, 100, 0},
+        {"eq on unselective field",     {F("bucket", Op::EQ, "even")},                std::nullopt, 100, 0},
+        {"eq plus SEARCH",              {F("grp", Op::EQ, "g3"), F("", Op::SEARCH, "g3")}, std::nullopt, 100, 0},
+        {"ne on indexed field",         {F("grp", Op::NE, "g3")},                     std::nullopt, 100, 0},
+        {"eq on nested path",           {F("nest.deep", Op::EQ, "d4")},               std::nullopt, 100, 0},
+        {"limit zero",                  {F("grp", Op::EQ, "g3")},                     std::nullopt, 0, 0},
+    };
+
+    auto run = [&](const Case& c) {
+        smartbotic::database::Query q;
+        q.filters = c.filters;
+        q.sort = c.sort;
+        q.limit = c.limit;
+        q.offset = c.offset;
+        auto r = store.scan("c", q);
+        std::string sig = "total=" + std::to_string(r.total_matched) +
+                          " more=" + std::to_string(r.has_more ? 1 : 0) + " [";
+        for (const auto& d : r.documents) { sig += d.id; sig += ","; }
+        sig += "]";
+        return sig;
+    };
+
+    for (const auto& c : cases) {
+        store.set_indexed_fields("c", {});                 // no index
+        const std::string without = run(c);
+
+        store.set_indexed_fields("c", {"grp", "bucket", "nest.deep"});
+        store.build_index("c", "grp");
+        store.build_index("c", "bucket");
+        store.build_index("c", "nest.deep");
+        const std::string with = run(c);
+
+        const std::string msg = std::string("indexed and unindexed plans agree: ")
+                                + c.name;
+        if (with != without) {
+            std::cerr << "  without index: " << without << "\n"
+                      << "  with index:    " << with << "\n";
+        }
+        check(with == without, msg.c_str());
+    }
+
+    // The agreement above is only meaningful if the index plan was actually
+    // TAKEN for the selective cases. Otherwise the planner declined every time
+    // and the test compared the scan against itself.
+    store.set_indexed_fields("c", {"grp", "bucket", "nest.deep"});
+
+    {
+        store.reset_index_plan_stats();
+        smartbotic::database::Query q;
+        q.limit = 100;
+        q.filters.push_back(F("grp", Op::EQ, "g3"));
+        auto r = store.scan("c", q);
+        auto st = store.index_plan_stats();
+        check(r.total_matched == 30, "the selective query matches 30 of 300 rows");
+        check(st.indexed_scans == 1 && st.full_scans == 0,
+              "a selective EQ on an indexed field TAKES the index plan - without "
+              "this the equivalence cases above would prove nothing");
+    }
+
+    {
+        store.reset_index_plan_stats();
+        smartbotic::database::Query q;
+        q.limit = 100;
+        q.filters.push_back(F("bucket", Op::EQ, "even"));
+        auto r = store.scan("c", q);
+        auto st = store.index_plan_stats();
+        check(r.total_matched == 150, "the unselective query matches half the rows");
+        check(st.indexed_scans == 0 && st.declined_unselective == 1,
+              "and the guard DECLINES its index - 150 of 300 rows would cost more "
+              "through the index than a scan, so declaring an index must not be "
+              "able to pessimise a query");
+    }
+
+    {
+        // An index on a field the query does not filter on must not be consulted.
+        store.reset_index_plan_stats();
+        smartbotic::database::Query q;
+        q.limit = 100;
+        q.filters.push_back(F("n", Op::GT, 250));
+        store.scan("c", q);
+        auto st = store.index_plan_stats();
+        check(st.full_scans == 1 && st.indexed_scans == 0,
+              "a query whose predicates name no indexed field scans");
+    }
+
+    auto n = store.index_count_eq("c", "bucket", nlohmann::json("even"));
+    check(n.has_value() && *n == 150,
+          "the unselective index does exist and holds 150 of 300 rows");
+}
+
 }  // namespace
 
 int main() {
@@ -711,6 +1025,10 @@ int main() {
     test_filtered_scan_operator_matrix();
     test_concurrent_reads_do_not_rebind_cached_handles();
     test_existing_collection_readable_without_writing_first();
+    test_index_tracks_documents_through_every_write();
+    test_build_index_over_existing_rows();
+    test_index_numeric_equality_matches_scan();
+    test_indexed_and_unindexed_plans_agree();
 
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;