Forráskód Böngészése

feat(index): ranges, CONTAINS, two-list intersection, and CLI commands

Closes the limits v2.9.0 shipped with.

Ranges needed the numeric encoding unified. v2.9.0 gave integers and reals
separate type tags - correct for equality, but they ordered independently, so
every integer sorted below every real regardless of value and a range spanning
both could not walk the index. A number is now one tag with a 16-byte payload:
an order-preserving double, then an exact-int tiebreak. The double orders all
numbers; the tiebreak separates large integers whose doubles collide (2^60 and
2^60+1 share a double, and _created_at is nanoseconds by default). Equality
still holds - integral values take the integral branch either way, so 5 and 5.0
remain one key.

On zeus data (executions, 10,118 rows / 592 MB): startedAt > p98 (202 rows)
337ms -> 9ms, identical match count. The same predicate over every row stays
333ms with the index declined.

The on-disk key format changed, so the sub-db prefix now carries a format
version: _idx_ -> _idx2_. A stale index is never consulted - the new name does
not exist, the lookup reports no-index, the query scans - rather than misread,
which would return wrong rows. applyIndexDeclarations() self-heals a declared
field whose sub-db is absent, because otherwise writes would maintain the new
index while pre-upgrade rows had no postings, and an indexed query would
silently return only the newer ones.

Ranges are served only when the index holds one type and it matches the bound's.
Across types the scan compares by JSON type order and a walk compares by tag
byte; rather than assume those agree, mixed types scan. There is no cheap count
for a range, so the walk is the probe - and on exceeding the budget it returns
no-plan, never a truncated list, since a partial candidate list silently drops
matching rows.

CONTAINS: an array contributes one posting per element, which is what array
membership asks about. Maintenance became a set difference. EQ for a scalar on
an array field may now name rows whose array merely contains it - safe only
because every candidate is re-checked against every filter.

Intersection: two exact predicates each too broad alone are intersected and used
if the result fits the budget. Postings are ids, not documents, so reading two
lists costs nothing like the per-candidate decode the budget protects against.

CLI: indexes / index-create / index-drop.

ctest 20/20, index e2e 19+14, namespacing 44/44, views + policy + TLS/auth green.
fszontagh 1 hónapja
szülő
commit
c9bda17d32

A különbségek nem kerülnek megjelenítésre, a fájl túl nagy
+ 0 - 0
CLAUDE.md


+ 1 - 1
VERSION

@@ -1 +1 @@
-2.9.0
+2.9.1

+ 62 - 0
cli/main.cpp

@@ -72,6 +72,12 @@ void printUsage() {
               << "  " << C_CYAN << "status" << C_RESET << "                               Show read-only + recovery status\n"
               << C_BOLD << "  Access policy (v2.7.0+)" << C_RESET << "\n"
               << "  " << C_CYAN << "security" << C_RESET << " [project]                   Show whether a project enforces policy\n"
+              << "  " << C_CYAN << "indexes" << C_RESET << " <collection>"
+              << "                    list secondary indexes\n"
+              << "  " << C_CYAN << "index-create" << C_RESET << " <collection> <field>"
+              << "        declare an index and backfill it\n"
+              << "  " << C_CYAN << "index-drop" << C_RESET << " <collection> <field>"
+              << "          remove an index\n"
               << "  " << C_CYAN << "security-set" << C_RESET << " <project> <on|off> [enforce|audit]\n"
               << "  " << C_CYAN << "policies" << C_RESET << " [project]                   List principals with a policy\n"
               << "  " << C_CYAN << "policy" << C_RESET << " <project> <principal>       Show one policy\n"
@@ -551,6 +557,62 @@ bool execCommand(smartbotic::database::Client& client,
         }
 
 
+        // ===== v2.9.0 secondary indexes =====
+        //
+        // Declaration is explicit because an index costs write throughput and is
+        // not always a win: on a low-cardinality field the planner will decline to
+        // use it, since reading the index and then fetching most of the collection
+        // by id loses to scanning. `indexes` reports distinct values precisely so
+        // an operator can see that coming.
+        if (cmd == "indexes") {
+            if (params.empty()) { printError("usage: indexes <collection>"); return false; }
+            auto list = client.listIndexes(params[0]);
+            if (list.empty()) {
+                std::cout << "no indexes on " << params[0] << "\n";
+                return true;
+            }
+            std::cout << C_BOLD << "field                          distinct      entries"
+                      << C_RESET << "\n";
+            for (const auto& i : list) {
+                std::cout << "  " << i.field
+                          << std::string(i.field.size() < 29 ? 29 - i.field.size() : 1, ' ')
+                          << i.distinctValues
+                          << std::string(std::to_string(i.distinctValues).size() < 13
+                                             ? 13 - std::to_string(i.distinctValues).size()
+                                             : 1, ' ')
+                          << i.entries << "\n";
+            }
+            return true;
+        }
+
+        if (cmd == "index-create") {
+            if (params.size() < 2) {
+                printError("usage: index-create <collection> <field>");
+                return false;
+            }
+            uint64_t rows = 0;
+            if (!client.createIndex(params[0], params[1], rows)) {
+                printError("could not create the index (see the service log)");
+                return false;
+            }
+            std::cout << "indexed " << rows << " existing row(s) on "
+                      << params[0] << "#" << params[1] << "\n";
+            return true;
+        }
+
+        if (cmd == "index-drop") {
+            if (params.size() < 2) {
+                printError("usage: index-drop <collection> <field>");
+                return false;
+            }
+            if (!client.dropIndex(params[0], params[1])) {
+                printError("could not drop the index (see the service log)");
+                return false;
+            }
+            std::cout << "dropped " << params[0] << "#" << params[1] << "\n";
+            return true;
+        }
+
         // ===== v2.7.0 access policy =====
         //
         // Policy lives in the `_policies` collection and is managed through the

+ 20 - 19
docs/ROADMAP.md

@@ -1,6 +1,6 @@
 # Smartbotic Database - Status and Roadmap
 
-**Current version: 2.9.0** (see `VERSION`). Last reviewed: 2026-08-09.
+**Current version: 2.9.1** (see `VERSION`). Last reviewed: 2026-08-09.
 
 This is the single authoritative statement of what exists and what does not.
 If any other document in this repository disagrees with this one, this one is
@@ -53,7 +53,7 @@ Consequences of that substitution, which trip up readers:
 
 ## Shipped
 
-Every item below is in the installed product as of 2.9.0. `CLAUDE.md` has the
+Every item below is in the installed product as of 2.9.1. `CLAUDE.md` has the
 detail and the failure modes.
 
 - JSON document store: collections, version history, field-level encryption, TTL
@@ -67,8 +67,9 @@ detail and the failure modes.
 - Durable snapshots with tiered recovery and read-only lockout
 - Replication, events/subscribe, migrations, set operations
 - Paging fast path (no filter, no sort) and the two-pass filtered/sorted scan
-- Secondary indexes on declared fields, equality filters, with a measured
-  selectivity guard so declaring one cannot make a query slower
+- Secondary indexes on declared fields: equality, CONTAINS, ranges, and
+  two-list intersection, with a measured selectivity guard so declaring one
+  cannot make a query slower. CLI: `indexes`, `index-create`, `index-drop`
 
 ---
 
@@ -77,21 +78,21 @@ detail and the failure modes.
 Ordered by what a reader is most likely to need next. Nothing here has a
 committed date.
 
-### 1. Indexing - equality only (shipped v2.9.0)
-
-Equality filters are served from secondary indexes. What remains:
-
-- **Only `EQ`.** Ranges (`GT`/`GTE`/`LT`/`LTE`) still scan. The blocker is the key
-  encoding: integral and non-integral numbers sit under different type tags, which
-  is correct for equality but means they order independently, so a range spanning
-  both cannot walk the index in one pass. Unify the numeric encoding first.
-- **`CONTAINS` (array membership)** needs a different index shape - one posting per
-  element rather than per value.
-- **Multi-predicate intersection.** The planner picks the single most selective
-  indexed EQ and applies the rest as filters. Intersecting two posting lists would
-  help queries that are selective only in combination.
-- **No CLI surface.** `smartbotic-db-cli` has no index commands; management is via
-  the client API or RPC only.
+### 1. Indexing - remaining operators
+
+Equality, `CONTAINS` and numeric/string ranges are served from indexes, with
+two-list intersection when neither predicate is selective alone. What is left:
+
+- **`IN`** would be a union of posting lists - the mirror of the intersection
+  already implemented, and the most obviously worthwhile next one.
+- **`NE`, `REGEX`, `SEARCH`** are not index-servable in this shape. `NE` and
+  `REGEX` would need a full walk anyway; `SEARCH` reads whole documents.
+- **Ranges on a mixed-type field** decline deliberately: an index walk compares
+  by type tag while the scan compares by JSON type order, and those are not
+  assumed to agree. Serving them means defining one order and proving both paths
+  use it.
+- **Compound (multi-column) indexes.** Intersection covers much of the benefit;
+  a real compound index would beat it when the pair is queried constantly.
 
 ### 2. `encode_document` still serialises with nlohmann
 

+ 16 - 0
service/src/database_service.cpp

@@ -350,6 +350,22 @@ void DatabaseService::applyIndexDeclarations() {
             ++applied;
             spdlog::info("v2.9 index: {} field(s) active on {}",
                          cfg.indexedFields.size(), qualified);
+
+            // Self-heal a declared index whose sub-db is absent. That happens
+            // when the KEY FORMAT VERSION in the sub-db prefix changes (v2.9.1
+            // unified the numeric encoding, so `_idx_` became `_idx2_`), and it
+            // would otherwise leave the declaration pointing at nothing: writes
+            // would maintain the new index correctly, but rows written BEFORE the
+            // upgrade would have no postings, so an indexed query would silently
+            // return only the newer ones. Rebuilding is the only safe reading.
+            for (const auto& field : cfg.indexedFields) {
+                if (lmdb->index_stats(rc.collection, field).has_value()) continue;
+                const uint64_t rows = lmdb->build_index(rc.collection, field);
+                spdlog::warn("v2.9 index: rebuilt {}#{} over {} row(s) - the index "
+                             "was declared but its sub-db was absent (key format "
+                             "change or a manual removal)",
+                             qualified, field, rows);
+            }
         } catch (const std::exception& e) {
             // Advisory per collection: one unparseable name must not stop the
             // rest from being armed. Loud, because a missing declaration means

+ 197 - 47
service/src/storage/document_store_lmdb.cpp

@@ -534,58 +534,68 @@ void LmdbDocumentStore::maintainIndexes(
     // field at a time - never materialised into a Document. On a 505 MB
     // collection a full decode_document costs ~1ms per write; resolving just the
     // indexed fields off the yyjson tree is a fraction of that.
-    std::unordered_map<std::string, std::optional<std::string>> old_keys;
+    // A field contributes a SET of keys, not one: an array value contributes one
+    // key per element so CONTAINS can be served. So the update is a set
+    // difference - remove the keys the row no longer owns, add the ones it gained,
+    // leave the rest untouched.
+    std::unordered_map<std::string, std::vector<std::string>> old_keys;
     if (!old_payload.empty()) {
         yyjson_doc* d = yyjson_read(old_payload.data(), old_payload.size(), 0);
         if (d) {
             yyjson_val* root = yyjson_doc_get_root(d);
             for (const auto& f : fields) {
                 auto v = resolve_from_yyjson(root, f);
-                old_keys[f] = v ? encode_index_key(*v) : std::nullopt;
+                if (v) old_keys[f] = encode_index_keys(*v);
             }
             yyjson_doc_free(d);
         }
     }
 
     for (const auto& f : fields) {
-        std::optional<std::string> new_key;
+        std::vector<std::string> new_k;
         if (new_doc != nullptr) {
             // resolveFilterValue is the SAME resolver the scan uses, so an
             // indexed lookup and a scan cannot disagree about which field a
             // filter names.
-            auto v = filter_eval::resolveFilterValue(*new_doc, f);
-            if (v) new_key = encode_index_key(*v);
+            if (auto v = filter_eval::resolveFilterValue(*new_doc, f)) {
+                new_k = encode_index_keys(*v);
+            }
         }
-        std::optional<std::string> old_key;
-        if (auto it = old_keys.find(f); it != old_keys.end()) old_key = it->second;
-
-        // An unchanged value needs no index write at all - the common case for
-        // an update that touches other fields.
-        if (old_key == new_key) continue;
+        std::vector<std::string> old_k;
+        if (auto it = old_keys.find(f); it != old_keys.end()) old_k = it->second;
+
+        // Unchanged needs no index write at all - the common case for an update
+        // that touches other fields. Both sides are sorted and deduplicated by
+        // encode_index_keys, so this comparison is exact.
+        if (old_k == new_k) continue;
+
+        std::vector<std::string> to_remove;
+        std::vector<std::string> to_add;
+        std::set_difference(old_k.begin(), old_k.end(), new_k.begin(), new_k.end(),
+                            std::back_inserter(to_remove));
+        std::set_difference(new_k.begin(), new_k.end(), old_k.begin(), old_k.end(),
+                            std::back_inserter(to_add));
+        if (to_remove.empty() && to_add.empty()) continue;
 
         const std::string sub = index_subdb_name(collection, f);
         const unsigned int dbi = open_for_write(wtxn, sub, MDB_DUPSORT);
         to_cache.emplace_back(sub, dbi);
 
-        if (old_key) {
-            MDB_val k = to_val(*old_key);
+        for (const auto& key : to_remove) {
+            MDB_val k = to_val(key);
             MDB_val v = to_val(id);
             // DUPSORT: passing the data removes just THIS id from the value's
             // posting list, not every id under that value.
             const int rc = mdb_del(wtxn.raw(), dbi, &k, &v);
-            if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) {
-                throw_mdb(rc, "index del");
-            }
+            if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) throw_mdb(rc, "index del");
         }
-        if (new_key) {
-            MDB_val k = to_val(*new_key);
+        for (const auto& key : to_add) {
+            MDB_val k = to_val(key);
             MDB_val v = to_val(id);
             // MDB_NODUPDATA makes a repeat put a no-op rather than an error, so
             // re-indexing the same row twice is safe (build_index re-runs).
             const int rc = mdb_put(wtxn.raw(), dbi, &k, &v, MDB_NODUPDATA);
-            if (rc != MDB_SUCCESS && rc != MDB_KEYEXIST) {
-                throw_mdb(rc, "index put");
-            }
+            if (rc != MDB_SUCCESS && rc != MDB_KEYEXIST) throw_mdb(rc, "index put");
         }
     }
 }
@@ -671,6 +681,8 @@ LmdbDocumentStore::IndexPlanStats LmdbDocumentStore::index_plan_stats() const {
     st.full_scans = full_scans_.load(std::memory_order_relaxed);
     st.declined_unselective =
         declined_unselective_.load(std::memory_order_relaxed);
+    st.intersected_scans = intersected_scans_.load(std::memory_order_relaxed);
+    st.range_scans = range_scans_.load(std::memory_order_relaxed);
     return st;
 }
 
@@ -678,6 +690,8 @@ void LmdbDocumentStore::reset_index_plan_stats() {
     indexed_scans_.store(0, std::memory_order_relaxed);
     full_scans_.store(0, std::memory_order_relaxed);
     declined_unselective_.store(0, std::memory_order_relaxed);
+    intersected_scans_.store(0, std::memory_order_relaxed);
+    range_scans_.store(0, std::memory_order_relaxed);
 }
 
 std::optional<LmdbDocumentStore::IndexStats>
@@ -715,6 +729,86 @@ LmdbDocumentStore::index_stats(std::string_view collection,
     return out;
 }
 
+
+std::optional<std::vector<std::string>>
+LmdbDocumentStore::lookupIndexRangeTxn(ReadTxn& rtxn,
+                                        std::string_view collection,
+                                        const std::string& field,
+                                        smartbotic::database::FilterOp op,
+                                        const nlohmann::json& value,
+                                        uint64_t budget) {
+    using Op = smartbotic::database::FilterOp;
+    if (op != Op::GT && op != Op::GTE && op != Op::LT && op != Op::LTE) {
+        return std::nullopt;
+    }
+    const auto bound = encode_index_key(value);
+    if (!bound) return std::nullopt;
+
+    auto dbi_opt = try_open_for_read(rtxn, index_subdb_name(collection, field));
+    if (!dbi_opt) return std::nullopt;
+
+    MDB_cursor* cur = nullptr;
+    mdb_check(mdb_cursor_open(rtxn.raw(), *dbi_opt, &cur), "cursor_open (index range)");
+    struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+
+    MDB_val k{0, nullptr};
+    MDB_val v{0, nullptr};
+
+    // Type homogeneity. Comparing the first and last key's tag is enough: keys are
+    // ordered, so one tag at both ends means one tag throughout. The identity
+    // sentinel sorts before every tag (it starts with NUL), so skip it.
+    int rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST);
+    while (rc == MDB_SUCCESS && is_identity_key(to_sv(k))) {
+        rc = mdb_cursor_get(cur, &k, &v, MDB_NEXT_NODUP);
+    }
+    if (rc == MDB_NOTFOUND) return std::vector<std::string>{};  // empty index
+    if (rc != MDB_SUCCESS) throw_mdb(rc, "cursor first (index range)");
+    const unsigned char first_tag = index_key_tag(to_sv(k));
+
+    MDB_val lk{0, nullptr};
+    MDB_val lv{0, nullptr};
+    rc = mdb_cursor_get(cur, &lk, &lv, MDB_LAST);
+    if (rc != MDB_SUCCESS) throw_mdb(rc, "cursor last (index range)");
+    if (index_key_tag(to_sv(lk)) != first_tag) return std::nullopt;   // mixed types
+    if (first_tag != index_key_tag(*bound)) return std::nullopt;      // wrong type
+
+    const std::string_view bsv = *bound;
+    std::vector<std::string> ids;
+
+    const bool ascending = (op == Op::GT || op == Op::GTE);
+    if (ascending) {
+        MDB_val sk = to_val(bsv);
+        rc = mdb_cursor_get(cur, &sk, &v, MDB_SET_RANGE);   // first key >= bound
+        // For GT, step past the bound's own duplicates.
+        while (rc == MDB_SUCCESS && op == Op::GT && to_sv(sk) == bsv) {
+            rc = mdb_cursor_get(cur, &sk, &v, MDB_NEXT_NODUP);
+        }
+        k = sk;
+    } else {
+        rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST);
+    }
+
+    while (rc == MDB_SUCCESS) {
+        const auto key = to_sv(k);
+        if (!is_identity_key(key)) {
+            if (index_key_tag(key) != first_tag) break;    // left the type
+            if (!ascending) {
+                const int cmp = key.compare(bsv);
+                if (op == Op::LT ? cmp >= 0 : cmp > 0) break;
+            }
+            ids.emplace_back(to_sv(v));
+            // Bail out rather than truncate. A partial list would drop matching
+            // rows; declining sends the query to the scan, which is merely slower.
+            if (ids.size() > budget) return std::nullopt;
+        }
+        rc = mdb_cursor_get(cur, &k, &v, MDB_NEXT);
+    }
+    if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) {
+        throw_mdb(rc, "cursor next (index range)");
+    }
+    return ids;
+}
+
 uint64_t LmdbDocumentStore::build_index(std::string_view collection,
                                          const std::string& field) {
     // Walk the collection once and index every row. Idempotent thanks to
@@ -749,12 +843,16 @@ uint64_t LmdbDocumentStore::build_index(std::string_view collection,
             auto val = resolve_from_yyjson(yyjson_doc_get_root(d), field);
             yyjson_doc_free(d);
             if (!val) continue;                       // field absent on this row
-            const auto key = encode_index_key(*val);
-            if (!key) continue;                       // unindexable value
-            MDB_val kk = to_val(*key);
-            MDB_val vv = to_val(id);
-            const int rc = mdb_put(wtxn.raw(), dbi, &kk, &vv, MDB_NODUPDATA);
-            if (rc != MDB_SUCCESS && rc != MDB_KEYEXIST) throw_mdb(rc, "index build put");
+            const auto keys = encode_index_keys(*val);
+            if (keys.empty()) continue;               // unindexable value
+            for (const auto& key : keys) {
+                MDB_val kk = to_val(key);
+                MDB_val vv = to_val(id);
+                const int rc = mdb_put(wtxn.raw(), dbi, &kk, &vv, MDB_NODUPDATA);
+                if (rc != MDB_SUCCESS && rc != MDB_KEYEXIST) {
+                    throw_mdb(rc, "index build put");
+                }
+            }
             ++indexed;
         }
         wtxn.commit();
@@ -889,6 +987,7 @@ LmdbDocumentStore::planIndexCandidates(ReadTxn& rtxn,
                                         std::string_view collection,
                                         const smartbotic::database::Query& query,
                                         unsigned int coll_dbi) {
+    using Op = smartbotic::database::FilterOp;
     const auto fields = indexed_fields(collection);
     if (fields.empty()) return std::nullopt;
 
@@ -898,33 +997,84 @@ LmdbDocumentStore::planIndexCandidates(ReadTxn& rtxn,
     if (total == 0) return std::nullopt;
     const uint64_t budget = total / kIndexSelectivityDivisor;
 
-    const std::string* best_field = nullptr;
-    const nlohmann::json* best_value = nullptr;
-    uint64_t best_count = 0;
-    bool found = false;
+    const auto is_indexed = [&](const std::string& f) {
+        return std::find(fields.begin(), fields.end(), f) != fields.end();
+    };
 
+    // --- exact predicates: EQ, and CONTAINS (which asks about an array ELEMENT,
+    // and elements are exactly what an array contributes to the index) ---
+    struct Exact {
+        const smartbotic::database::Filter* f;
+        uint64_t count;
+    };
+    std::vector<Exact> exacts;
     for (const auto& f : query.filters) {
-        if (f.op != smartbotic::database::FilterOp::EQ) continue;
-        if (std::find(fields.begin(), fields.end(), f.field) == fields.end()) continue;
-        auto n = countIndexEqTxn(rtxn, collection, f.field, f.value);
-        if (!n) continue;                     // no index, or unindexable value
-        if (!found || *n < best_count) {
-            found = true;
-            best_count = *n;
-            best_field = &f.field;
-            best_value = &f.value;
+        if (f.op != Op::EQ && f.op != Op::CONTAINS) continue;
+        if (!is_indexed(f.field)) continue;
+        if (auto n = countIndexEqTxn(rtxn, collection, f.field, f.value)) {
+            exacts.push_back({&f, *n});
         }
     }
-    if (!found) return std::nullopt;
+    std::sort(exacts.begin(), exacts.end(),
+              [](const Exact& a, const Exact& b) { return a.count < b.count; });
+
+    // Cheapest plan: one selective exact predicate. Zero candidates is a
+    // legitimate, maximally selective answer - the query matches nothing and we
+    // never read a row.
+    if (!exacts.empty() && exacts.front().count <= budget) {
+        const auto* f = exacts.front().f;
+        return lookupIndexEqTxn(rtxn, collection, f->field, f->value);
+    }
 
-    // Zero candidates is a legitimate, extremely selective answer - the query
-    // matches nothing and we can say so without reading a single row.
-    if (best_count > budget) {
-        declined_unselective_.fetch_add(1, std::memory_order_relaxed);
-        return std::nullopt;
+    // --- intersection: two predicates each too broad alone, selective together.
+    //
+    // Worth trying because postings are ids, not documents: reading two lists
+    // costs nothing like the decode_document per candidate that the budget is
+    // protecting against. Only the INTERSECTION has to fit the budget.
+    //
+    // Correctness: each list is a superset of the rows matching its own
+    // predicate, so the intersection is a superset of the rows matching both -
+    // and every candidate is still checked against every filter afterwards.
+    if (exacts.size() >= 2) {
+        const uint64_t postings_cap = total;   // reading ids is cheap; decoding is not
+        if (exacts[0].count <= postings_cap && exacts[1].count <= postings_cap) {
+            auto a = lookupIndexEqTxn(rtxn, collection, exacts[0].f->field,
+                                       exacts[0].f->value);
+            auto b = lookupIndexEqTxn(rtxn, collection, exacts[1].f->field,
+                                       exacts[1].f->value);
+            if (a && b) {
+                std::sort(a->begin(), a->end());
+                std::sort(b->begin(), b->end());
+                std::vector<std::string> both;
+                std::set_intersection(a->begin(), a->end(), b->begin(), b->end(),
+                                      std::back_inserter(both));
+                if (both.size() <= budget) {
+                    intersected_scans_.fetch_add(1, std::memory_order_relaxed);
+                    return both;
+                }
+            }
+        }
+    }
+
+    // --- ranges. There is no cheap count for a range, so the walk itself is the
+    // probe: it collects ids and gives up the moment it passes the budget,
+    // returning no-plan rather than a truncated list.
+    for (const auto& f : query.filters) {
+        if (f.op != Op::GT && f.op != Op::GTE && f.op != Op::LT && f.op != Op::LTE) {
+            continue;
+        }
+        if (!is_indexed(f.field)) continue;
+        if (auto ids = lookupIndexRangeTxn(rtxn, collection, f.field, f.op,
+                                           f.value, budget)) {
+            range_scans_.fetch_add(1, std::memory_order_relaxed);
+            return ids;
+        }
     }
 
-    return lookupIndexEqTxn(rtxn, collection, *best_field, *best_value);
+    if (!exacts.empty()) {
+        declined_unselective_.fetch_add(1, std::memory_order_relaxed);
+    }
+    return std::nullopt;
 }
 
 ScanResult LmdbDocumentStore::scan(std::string_view collection,

+ 22 - 0
service/src/storage/document_store_lmdb.hpp

@@ -81,6 +81,8 @@ public:
         uint64_t indexed_scans = 0;     // an index served the query
         uint64_t full_scans = 0;        // walked the collection
         uint64_t declined_unselective = 0;  // an index existed but was too broad
+        uint64_t intersected_scans = 0;      // two posting lists intersected
+        uint64_t range_scans = 0;            // a range walked the index
     };
     IndexPlanStats index_plan_stats() const;
     void reset_index_plan_stats();
@@ -160,6 +162,8 @@ private:
     mutable std::atomic<uint64_t> indexed_scans_{0};
     mutable std::atomic<uint64_t> full_scans_{0};
     mutable std::atomic<uint64_t> declined_unselective_{0};
+    mutable std::atomic<uint64_t> intersected_scans_{0};
+    mutable std::atomic<uint64_t> range_scans_{0};
     std::unordered_map<std::string, unsigned int> dbi_cache_;
 
     // Resolve a collection name to an MDB_dbi handle.
@@ -191,6 +195,24 @@ private:
     countIndexEqTxn(class ReadTxn& rtxn, std::string_view collection,
                     const std::string& field, const nlohmann::json& value);
 
+    // v2.9.1 — candidate ids for a range predicate, or nullopt to decline.
+    //
+    // Declines when: there is no index, the bound is unindexable, the index holds
+    // more than one JSON type (a walk compares by tag byte while the scan compares
+    // by JSON type order - rather than assume those agree, mixed types scan), the
+    // index's type differs from the bound's, or the range would exceed `budget`
+    // candidates.
+    //
+    // ⚠ On exceeding the budget it returns nullopt, NEVER a truncated list. A
+    // partial candidate list would silently drop matching rows, which is the one
+    // failure mode an index must not have.
+    std::optional<std::vector<std::string>>
+    lookupIndexRangeTxn(class ReadTxn& rtxn, std::string_view collection,
+                        const std::string& field,
+                        smartbotic::database::FilterOp op,
+                        const nlohmann::json& value,
+                        uint64_t budget);
+
     // Choose an index for a query, or nullopt to walk the whole collection.
     std::optional<std::vector<std::string>>
     planIndexCandidates(class ReadTxn& rtxn, std::string_view collection,

+ 86 - 41
service/src/storage/secondary_index.cpp

@@ -2,6 +2,7 @@
 
 #include "storage/secondary_index.hpp"
 
+#include <algorithm>
 #include <cmath>
 #include <cstdint>
 #include <cstring>
@@ -16,8 +17,10 @@ constexpr char kSep = '#';
 constexpr unsigned char kTagNull = 0x10;
 constexpr unsigned char kTagFalse = 0x20;
 constexpr unsigned char kTagTrue = 0x21;
-constexpr unsigned char kTagInt = 0x30;
-constexpr unsigned char kTagReal = 0x31;
+// v2.9.1 — ONE tag for every number, so numeric ranges can walk the index in a
+// single pass. v2.9.0 had separate int and real tags: correct for equality, but
+// they ordered independently, so a range spanning both could not be served.
+constexpr unsigned char kTagNum = 0x30;
 constexpr unsigned char kTagStr = 0x40;
 
 // Append `v` big-endian so that memcmp order matches unsigned integer order.
@@ -27,35 +30,55 @@ void append_be64(std::string& out, uint64_t v) {
     }
 }
 
-std::string encode_int(int64_t v) {
-    std::string out;
-    out.reserve(9);
-    out.push_back(static_cast<char>(kTagInt));
-    // Flipping the sign bit maps the signed range onto unsigned while
-    // preserving order: INT64_MIN -> 0, -1 -> 0x7FFF..., 0 -> 0x8000..., etc.
-    const uint64_t biased =
-        static_cast<uint64_t>(v) ^ (uint64_t{1} << 63);
-    append_be64(out, biased);
-    return out;
+// Order-preserving transform for a signed integer: flipping the sign bit maps the
+// signed range onto unsigned while preserving order (INT64_MIN -> 0, -1 ->
+// 0x7FFF..., 0 -> 0x8000...).
+uint64_t order_bits_int(int64_t v) {
+    return static_cast<uint64_t>(v) ^ (uint64_t{1} << 63);
 }
 
-std::string encode_real(double d) {
-    std::string out;
-    out.reserve(9);
-    out.push_back(static_cast<char>(kTagReal));
+// The standard order-preserving transform for IEEE754: for negatives (sign bit
+// set) invert every bit, which reverses their descending magnitude order; for
+// non-negatives set the sign bit so they sort above all negatives.
+uint64_t order_bits_double(double d) {
     uint64_t bits = 0;
     static_assert(sizeof(double) == sizeof(uint64_t), "IEEE754 double expected");
     std::memcpy(&bits, &d, sizeof(bits));
-    // The standard order-preserving transform for IEEE754: for negatives
-    // (sign bit set) invert every bit, which reverses their descending
-    // magnitude order; for non-negatives just set the sign bit so they sort
-    // above all negatives.
-    if (bits & (uint64_t{1} << 63)) {
-        bits = ~bits;
-    } else {
-        bits |= (uint64_t{1} << 63);
-    }
-    append_be64(out, bits);
+    if (bits & (uint64_t{1} << 63)) return ~bits;
+    return bits | (uint64_t{1} << 63);
+}
+
+// True when `d` is exactly an int64, i.e. integral and inside the range.
+bool exact_int64(double d, int64_t& out) {
+    if (!std::isfinite(d) || d != std::floor(d)) return false;
+    // The int64 bounds are not exactly representable as double, so compare
+    // against the representable neighbours rather than the bounds themselves.
+    if (d < -9223372036854775808.0) return false;
+    if (d >= 9223372036854775808.0) return false;
+    out = static_cast<int64_t>(d);
+    return true;
+}
+
+// A number becomes a 17-byte key: tag, then an 8-byte order-preserving double,
+// then an 8-byte exactness tiebreak.
+//
+// The double sorts every number correctly against every other number, which is
+// what makes a single-pass range walk possible. But doubles cannot separate large
+// integers - 2^60 and 2^60+1 share a double - and nanosecond timestamps live up
+// there, so the second component carries the exact integer and breaks those ties.
+//
+// Two numbers with the same value therefore always produce the same key
+// (integral 5 and 5.0 both take the integral branch), and two different int64
+// values always produce different keys even when their doubles collide.
+std::string encode_number(double as_double, bool integral, int64_t exact) {
+    std::string out;
+    out.reserve(17);
+    out.push_back(static_cast<char>(kTagNum));
+    append_be64(out, order_bits_double(as_double));
+    // For a non-integral value the tiebreak is a constant. That cannot collide
+    // with an integral value's key, because sharing the primary would require the
+    // non-integral value to equal an integer.
+    append_be64(out, order_bits_int(integral ? exact : 0));
     return out;
 }
 
@@ -88,6 +111,10 @@ std::optional<IndexSubdb> parse_index_subdb(std::string_view name) {
     return out;
 }
 
+unsigned char index_key_tag(std::string_view key) {
+    return key.empty() ? 0 : static_cast<unsigned char>(key[0]);
+}
+
 std::optional<std::string> encode_index_key(const nlohmann::json& value) {
     switch (value.type()) {
         case nlohmann::json::value_t::null:
@@ -106,32 +133,32 @@ std::optional<std::string> encode_index_key(const nlohmann::json& value) {
             return out;
         }
 
-        case nlohmann::json::value_t::number_integer:
-            return encode_int(value.get<int64_t>());
+        case nlohmann::json::value_t::number_integer: {
+            const int64_t i = value.get<int64_t>();
+            return encode_number(static_cast<double>(i), true, i);
+        }
 
         case nlohmann::json::value_t::number_unsigned: {
             const uint64_t u = value.get<uint64_t>();
-            // Above INT64_MAX there is no int64 form, so such a value can only
-            // be compared against another value stored the same way. Encode it
-            // as a real, which is what nlohmann's own comparison degrades to.
+            // Above INT64_MAX there is no int64 form, so the tiebreak cannot
+            // represent it; the double alone orders it, which is also what
+            // nlohmann's own comparison degrades to up there.
             if (u > static_cast<uint64_t>(std::numeric_limits<int64_t>::max())) {
-                return encode_real(static_cast<double>(u));
+                return encode_number(static_cast<double>(u), false, 0);
             }
-            return encode_int(static_cast<int64_t>(u));
+            const int64_t i = static_cast<int64_t>(u);
+            return encode_number(static_cast<double>(i), true, i);
         }
 
         case nlohmann::json::value_t::number_float: {
             const double d = value.get<double>();
             if (std::isnan(d)) return std::nullopt;   // never equals anything
             // CANONICALISATION - the invariant in the header. nlohmann compares
-            // numbers across subtypes, so 5.0 must produce the same key as 5 or
-            // an indexed EQ would miss rows the scan finds.
-            if (std::isfinite(d) && d == std::floor(d) &&
-                d >= static_cast<double>(std::numeric_limits<int64_t>::min()) &&
-                d <= static_cast<double>(std::numeric_limits<int64_t>::max())) {
-                return encode_int(static_cast<int64_t>(d));
-            }
-            return encode_real(d);
+            // numbers across subtypes, so 5.0 must produce the same key as 5, or
+            // an indexed EQ would miss rows a scan finds.
+            int64_t exact = 0;
+            if (exact_int64(d, exact)) return encode_number(d, true, exact);
+            return encode_number(d, false, 0);
         }
 
         // A whole object or array is not a lookup key. CONTAINS asks about an
@@ -146,4 +173,22 @@ std::optional<std::string> encode_index_key(const nlohmann::json& value) {
     }
 }
 
+std::vector<std::string> encode_index_keys(const nlohmann::json& value) {
+    std::vector<std::string> out;
+    if (value.is_array()) {
+        out.reserve(value.size());
+        for (const auto& el : value) {
+            // Nested arrays and objects as elements are skipped rather than
+            // flattened: CONTAINS compares an element with ==, and a nested
+            // container needs the whole-value comparison the scan does.
+            if (auto k = encode_index_key(el)) out.push_back(std::move(*k));
+        }
+        std::sort(out.begin(), out.end());
+        out.erase(std::unique(out.begin(), out.end()), out.end());
+        return out;
+    }
+    if (auto k = encode_index_key(value)) out.push_back(std::move(*k));
+    return out;
+}
+
 }  // namespace smartbotic::db::storage

+ 42 - 9
service/src/storage/secondary_index.hpp

@@ -24,11 +24,16 @@
 // encode_index_key therefore CANONICALISES: any number whose value is integral
 // and fits int64 encodes to the integer form regardless of how it was stored.
 //
-// Known limitation, deliberate: integral and non-integral numbers land under
-// different type tags, so they order independently. That is fine for EQ (the
-// only operator wired to indexes today) but means a range query spanning both
-// cannot be served from an index yet. Ranges must keep using the scan until the
-// encoding is unified.
+// v2.9.1: every number now shares ONE tag and one ordering, so a numeric range
+// walks the index in a single pass. The payload is a 8-byte order-preserving
+// double followed by an 8-byte exact-integer tiebreak - the double orders all
+// numbers, and the tiebreak separates large integers whose doubles collide
+// (2^60 and 2^60+1 share a double, and nanosecond timestamps live up there).
+//
+// Ranges are only served when an index holds ONE type, checked by comparing the
+// tag of its first and last key. Across types the scan compares by JSON type
+// order, and an index walk compares by tag byte; rather than assume those agree,
+// a mixed-type index declines the range and the query scans.
 
 #pragma once
 
@@ -43,7 +48,16 @@ namespace smartbotic::db::storage {
 
 // Prefix marking a sub-db as an index. Leading '_' means is_system_subdb()
 // already treats these as internal, so they stay out of list_collections().
-inline constexpr std::string_view kIndexSubdbPrefix = "_idx_";
+// The prefix carries the KEY FORMAT VERSION. v2.9.1 changed the numeric encoding
+// (separate int/real tags -> one ordered numeric space), so keys written by
+// v2.9.0 are unreadable under the new rules. Bumping the prefix means a stale
+// index is never consulted: the new name simply does not exist yet, the lookup
+// reports "no index", and the query scans until the index is rebuilt. A silent
+// misread would have been the alternative, and that returns wrong rows.
+//
+// v2.9.0's `_idx_` sub-dbs are left behind as harmless litter; `drop_index`
+// removes only the current format's.
+inline constexpr std::string_view kIndexSubdbPrefix = "_idx2_";
 
 // Sub-db name for one indexed field.
 //
@@ -78,10 +92,29 @@ std::optional<IndexSubdb> parse_index_subdb(std::string_view name);
 //   0x10  null      (no payload)
 //   0x20  false     (no payload)
 //   0x21  true      (no payload)
-//   0x30  integer   8 bytes, big-endian int64 with the sign bit flipped
-//   0x31  real      8 bytes, big-endian IEEE754 with the standard
-//                   order-preserving transform
+//   0x30  number    16 bytes: order-preserving double, then exact-int tiebreak
 //   0x40  string    raw UTF-8 bytes
 std::optional<std::string> encode_index_key(const nlohmann::json& value);
 
+// Leading type tag of an encoded key, 0 for an empty key. A range walk uses it to
+// stay inside one type.
+unsigned char index_key_tag(std::string_view key);
+
+// Every index key a single field value contributes.
+//
+// A scalar contributes one key. An ARRAY contributes one key PER ELEMENT, which
+// is what lets CONTAINS - array membership, compared element-wise with
+// nlohmann's == - be answered from the index.
+//
+// Note what this means for EQ on an array-valued field: EQ compares the whole
+// array, which has no key here, so encode_index_key returns nullopt and the
+// planner declines. And if EQ asks for a SCALAR on an array-valued field, the
+// index may return rows whose array merely contains it - a superset. That is
+// safe only because every candidate is re-checked against every filter; the
+// index narrows work, it never decides the answer.
+//
+// Duplicate elements collapse: DUPSORT stores a value's ids as a set, so one row
+// cannot appear twice under one key.
+std::vector<std::string> encode_index_keys(const nlohmann::json& value);
+
 }  // namespace smartbotic::db::storage

+ 36 - 0
tests/load_test/test_indexes.cpp

@@ -68,6 +68,9 @@ void setup(Client& c) {
         nlohmann::json d{
             {"wf", "wf-" + std::to_string(i % 30)},      // 30 groups of 10 = 3.3%
             {"state", (i % 2 == 0) ? "done" : "pending"},  // 50%, unselective
+            {"seq", i},                                    // for range queries
+            {"labels", nlohmann::json::array(
+                 {(i % 2 == 0) ? "even" : "odd", "all"})},  // for CONTAINS
         };
         c.upsert(kColl, d, "r" + std::to_string(1000 + i));
     }
@@ -153,6 +156,39 @@ void verify(Client& c) {
 
     check(c.dropIndex(kColl, "nosuch"),
           "dropping an index that does not exist is idempotent success");
+
+    // v2.9.1 — ranges and CONTAINS through the real client. These go through the
+    // same scan path the unit tests cover, but only an e2e proves the wire types
+    // survive the round trip.
+    uint64_t n2 = 0;
+    check(c.createIndex(kColl, "seq", n2), "index a numeric field");
+    check(c.createIndex(kColl, "labels"), "index an array field");
+
+    {
+        Client::QueryOptions o;
+        o.limit = 1000;
+        o.filters.push_back(Client::Filter::gt("seq", 280));
+        const size_t indexed = c.find(kColl, o).size();
+        check(c.dropIndex(kColl, "seq"), "drop the numeric index");
+        const size_t scanned = c.find(kColl, o).size();
+        check(indexed == scanned,
+              "a range returns the same rows with and without an index (" +
+              std::to_string(indexed) + ")");
+        check(indexed > 0 && indexed < 60, "and the range is genuinely selective");
+    }
+
+    {
+        Client::QueryOptions o;
+        o.limit = 1000;
+        o.filters.push_back(Client::Filter::contains("labels", "even"));
+        const size_t indexed = c.find(kColl, o).size();
+        check(c.dropIndex(kColl, "labels"), "drop the array index");
+        const size_t scanned = c.find(kColl, o).size();
+        check(indexed == scanned,
+              "CONTAINS returns the same rows with and without an index (" +
+              std::to_string(indexed) + ")");
+        check(indexed > 0, "and it matched something");
+    }
 }
 
 }  // namespace

+ 28 - 0
tests/test_secondary_index_keys.cpp

@@ -117,6 +117,34 @@ void test_ordering_is_preserved() {
 
     check(key("a") < key("b"), "strings sort lexicographically");
     check(key("") < key("a"), "the empty string sorts first");
+
+    // v2.9.1 — THE point of unifying the numeric encoding. Integers and reals
+    // must interleave in one ordered space, otherwise a range spanning both
+    // cannot walk the index in a single pass. v2.9.0 kept them under separate
+    // tags, where every integer sorted below every real regardless of value.
+    check(key(5) < key(5.5), "an integer sorts below a larger real");
+    check(key(5.5) < key(6), "and a real below a larger integer");
+    check(key(-6) < key(-5.5), "the same holds through negatives");
+    check(key(-5.5) < key(-5), "negative real between two negative integers");
+    check(key(-1) < key(-0.5), "just below zero");
+    check(key(-0.5) < key(0), "and up to zero");
+    check(key(0) < key(0.5), "and past it");
+
+    // A full mixed sequence must come out in numeric order.
+    const std::vector<nlohmann::json> ascending = {
+        -1000, -10.5, -10, -1, -0.25, 0, 0.25, 1, 2.5, 3, 1e6, 1e300,
+    };
+    bool ordered = true;
+    for (size_t i = 1; i < ascending.size(); ++i) {
+        if (!(key(ascending[i - 1]) < key(ascending[i]))) ordered = false;
+    }
+    check(ordered, "a mixed integer/real sequence encodes in numeric order");
+
+    // Large integers keep their exact ordering even though their doubles collide.
+    const int64_t big = 1786263002080195076LL;
+    check(key(big) < key(big + 1),
+          "neighbouring ns timestamps order correctly - the tiebreak, since their "
+          "doubles are equal");
 }
 
 void test_unindexable_values() {

+ 135 - 0
tests/test_subdb_identity.cpp

@@ -1043,6 +1043,140 @@ void test_empty_id_reads_as_absent() {
     check(store.count("c") == 1, "and nothing was disturbed");
 }
 
+
+// v2.9.1 — ranges, CONTAINS and intersection must agree with the scan too, and
+// must actually be USED. Each is a new way for the index to disagree with a
+// brute-force answer, and each disagreement would be silent.
+void test_range_contains_and_intersection() {
+    TmpEnv t("idx-v291");
+    LmdbDocumentStore store(t.env);
+
+    // n mixes integers and reals on purpose: v2.9.0 encoded those under separate
+    // type tags, so a range spanning both could not be served at all.
+    for (int i = 0; i < 400; ++i) {
+        Document d;
+        d.id = "r" + std::string(i < 10 ? "00" : (i < 100 ? "0" : "")) + std::to_string(i);
+        d.collection = "c";
+        nlohmann::json data{
+            {"n", (i % 2 == 0) ? nlohmann::json(i) : nlohmann::json(i + 0.5)},
+            {"grp", "g" + std::to_string(i % 20)},
+            {"other", "o" + std::to_string(i % 20)},
+            // Deliberately COARSE: 4 and 5 values, so each matches 25% and 20% of
+            // 400 rows - both above the 10% budget alone - while their pair
+            // narrows to i%20, i.e. 20 rows (5%). That is the only shape where
+            // intersecting two posting lists earns its keep.
+            {"q4", i % 4},
+            {"q5", i % 5},
+            {"tags", nlohmann::json::array({"t" + std::to_string(i % 25), "all"})},
+        };
+        d.set_data(data);
+        store.put("c", d.id, d);
+    }
+
+    using Op = smartbotic::database::FilterOp;
+    auto F = [](const char* f, Op op, const nlohmann::json& v) {
+        smartbotic::database::Filter x;
+        x.field = f; x.op = op; x.value = v;
+        return x;
+    };
+    auto sig = [&](const std::vector<smartbotic::database::Filter>& fs,
+                   std::optional<smartbotic::database::Sort> so = std::nullopt) {
+        smartbotic::database::Query q;
+        q.filters = fs;
+        q.sort = so;
+        q.limit = 1000;
+        auto r = store.scan("c", q);
+        std::string out = "total=" + std::to_string(r.total_matched) + " [";
+        std::vector<std::string> ids;
+        for (const auto& d : r.documents) ids.push_back(d.id);
+        std::sort(ids.begin(), ids.end());
+        for (const auto& i : ids) { out += i; out += ","; }
+        return out + "]";
+    };
+
+    const std::vector<std::pair<const char*, std::vector<smartbotic::database::Filter>>> cases = {
+        {"GT on a mixed int/real field",   {F("n", Op::GT, 380)}},
+        {"GTE on a mixed field",           {F("n", Op::GTE, 380)}},
+        {"LT on a mixed field",            {F("n", Op::LT, 12)}},
+        {"LTE on a mixed field",           {F("n", Op::LTE, 12)}},
+        {"GT with a real bound",           {F("n", Op::GT, 380.5)}},
+        {"range that matches nothing",     {F("n", Op::GT, 100000)}},
+        {"range that matches everything",  {F("n", Op::GT, -1)}},
+        {"CONTAINS an array element",      {F("tags", Op::CONTAINS, "t3")}},
+        {"CONTAINS a common element",      {F("tags", Op::CONTAINS, "all")}},
+        {"CONTAINS a missing element",     {F("tags", Op::CONTAINS, "nope")}},
+        {"two broad EQs, narrow together", {F("q4", Op::EQ, 1), F("q5", Op::EQ, 2)}},
+        {"two broad EQs, disjoint result",  {F("q4", Op::EQ, 1), F("q5", Op::EQ, 2), F("grp", Op::EQ, "g19")}},
+        {"range plus EQ",                  {F("n", Op::GT, 300), F("grp", Op::EQ, "g3")}},
+        {"EQ on an array field (whole)",   {F("tags", Op::EQ, nlohmann::json::array({"t3", "all"}))}},
+    };
+
+    for (const auto& [name, filters] : cases) {
+        store.set_indexed_fields("c", {});
+        const std::string without = sig(filters);
+
+        store.set_indexed_fields("c", {"n", "grp", "other", "tags", "q4", "q5"});
+        for (const char* f : {"n", "grp", "other", "tags", "q4", "q5"}) {
+            store.build_index("c", f);
+        }
+        const std::string with = sig(filters);
+
+        if (with != without) {
+            std::cerr << "    without: " << without.substr(0, 200) << "\n"
+                      << "    with:    " << with.substr(0, 200) << "\n";
+        }
+        const std::string msg = std::string("indexed == unindexed: ") + name;
+        check(with == without, msg.c_str());
+    }
+
+    // Now prove each new plan is actually taken.
+    store.set_indexed_fields("c", {"n", "grp", "other", "tags", "q4", "q5"});
+
+    {
+        store.reset_index_plan_stats();
+        (void)sig({F("n", Op::GT, 380)});
+        auto st = store.index_plan_stats();
+        check(st.range_scans == 1 && st.indexed_scans == 1,
+              "a selective range WALKS the index - impossible before the numeric "
+              "encoding was unified");
+    }
+    {
+        store.reset_index_plan_stats();
+        (void)sig({F("n", Op::GT, -1)});          // matches everything
+        auto st = store.index_plan_stats();
+        check(st.range_scans == 0 && st.full_scans == 1,
+              "an unselective range gives up and scans - and must return no-plan "
+              "rather than a truncated candidate list");
+    }
+    {
+        store.reset_index_plan_stats();
+        (void)sig({F("tags", Op::CONTAINS, "t3")});
+        auto st = store.index_plan_stats();
+        check(st.indexed_scans == 1,
+              "CONTAINS is served from the per-element postings an array writes");
+    }
+    {
+        store.reset_index_plan_stats();
+        (void)sig({F("q4", Op::EQ, 1), F("q5", Op::EQ, 2)});
+        auto st = store.index_plan_stats();
+        check(st.intersected_scans == 1 && st.indexed_scans == 1,
+              "q4 matches 25% and q5 20% - each too broad alone - so their posting "
+              "lists are INTERSECTED down to 5% instead of scanning");
+    }
+
+    // An array-valued field: EQ on a scalar may over-return from the index, which
+    // is safe only because every candidate is re-filtered. Pin that.
+    {
+        const std::string indexed = sig({F("tags", Op::EQ, "t3")});
+        store.set_indexed_fields("c", {});
+        const std::string scanned = sig({F("tags", Op::EQ, "t3")});
+        store.set_indexed_fields("c", {"n", "grp", "other", "tags", "q4", "q5"});
+        check(indexed == scanned,
+              "EQ for a scalar on an array field agrees - the index may name rows "
+              "whose array merely CONTAINS it, and re-filtering drops them");
+    }
+}
+
 }  // namespace
 
 int main() {
@@ -1064,6 +1198,7 @@ int main() {
     test_index_numeric_equality_matches_scan();
     test_indexed_and_unindexed_plans_agree();
     test_empty_id_reads_as_absent();
+    test_range_contains_and_intersection();
 
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;

Nem az összes módosított fájl került megjelenítésre, mert túl sok fájl változott