ソースを参照

feat(index): serve totals, range-ordering, and distinct values from the index

The remaining full scans were never about finding rows - they were about
COUNTING them. total_matched is part of the contract, and counting matches meant
visiting every one. That, not lookup, is why status="completed" (66% of rows)
was declined.

Counted plan: for a single unsorted EQ or IN on a non-multivalued index the total
comes from mdb_cursor_count and only the page's rows are read, so no selectivity
budget applies and breadth is free. On zeus executions (10,118 rows / 592 MB):
status=completed limit=10 338ms -> 4ms, offset=6000 354ms -> 12ms, total 6,689
exact both ways. Page order needs no work - DUPSORT stores a key's ids ascending
and document keys are ids.

A count is only the match count while no row owns more than one posting, so an
index now records a multivalued marker the first time an array value is indexed,
and the counted plan declines once it is set: under an array field a key's
duplicate count counts rows CONTAINING the element, while EQ compares the whole
array.

Range + sort on the same field: the walk's order already IS the sort order, so no
sorting pass and only the page is decoded. Reversing an ascending walk reproduces
sort_documents' reverse-id tie-break. Rows lacking the field need no guard here,
unlike pure ordering - a missing value fails a range anyway.

Distinct values + min/max via GetIndexValues: one step per distinct value rather
than per row. The value is read from a row holding it, never decoded from the
index key, so there is no decoder to drift from the encoder and the value is
returned as stored. An array field reports nothing, its keys being elements.

Covering reads investigated and closed as not viable: the response always retains
six metadata fields and only _id exists in a posting.

Unique constraints built and deliberately NOT shipped. The check is correct but
cannot be enforced: it throws from put(), inside applyDualWriteMirror, which
catches every exception, bumps drift and flips mirror_healthy_ - so a rejection
would be swallowed AND would send every read to MemoryStore, the v2.8.1 fault.
MemoryStore also mutates before the mirror runs, so rejecting cleanly needs a
rollback. Nothing exposes it, so no constraint is advertised that does not hold.

Also fixed: is_index_meta_key excludes both reserved keys from every index walk
and count, and rowCountTxn excludes the identity sentinel from a collection's row
count - a raw mdb_stat includes it, which silently disabled the v2.9.2 ordered
plan until a test asserted the plan had actually run.

test_subdb_identity 236, index e2e 33, ctest 20/20, all suites green.
fszontagh 1 ヶ月 前
親
コミット
408da82c75

ファイルの差分が大きいため隠しています
+ 0 - 0
CLAUDE.md


+ 1 - 1
VERSION

@@ -1 +1 @@
-2.9.2
+2.10.0

+ 24 - 0
cli/main.cpp

@@ -78,6 +78,8 @@ void printUsage() {
               << "        declare an index and backfill it\n"
               << "  " << C_CYAN << "index-drop" << C_RESET << " <collection> <field>"
               << "          remove an index\n"
+              << "  " << C_CYAN << "index-values" << C_RESET << " <coll> <field> [n] [asc|desc]"
+              << "  distinct values + counts\n"
               << "  " << C_CYAN << "security-set" << C_RESET << " <project> <on|off> [enforce|audit]\n"
               << "  " << C_CYAN << "policies" << C_RESET << " [project]                   List principals with a policy\n"
               << "  " << C_CYAN << "policy" << C_RESET << " <project> <principal>       Show one policy\n"
@@ -585,6 +587,28 @@ bool execCommand(smartbotic::database::Client& client,
             return true;
         }
 
+        if (cmd == "index-values") {
+            if (params.size() < 2) {
+                printError("usage: index-values <collection> <field> [limit] [asc|desc]");
+                return false;
+            }
+            const uint32_t limit = params.size() > 2 ? std::stoul(params[2]) : 20;
+            const bool asc = params.size() > 3 ? (params[3] != "desc") : true;
+            auto vals = client.indexValues(params[0], params[1], limit, asc);
+            if (vals.empty()) {
+                std::cout << "no values (is " << params[1] << " indexed?)\n";
+                return true;
+            }
+            std::cout << C_BOLD << "rows       value" << C_RESET << "\n";
+            for (const auto& v : vals) {
+                const std::string c = std::to_string(v.count);
+                std::cout << "  " << c
+                          << std::string(c.size() < 9 ? 9 - c.size() : 1, ' ')
+                          << v.value.dump() << "\n";
+            }
+            return true;
+        }
+
         if (cmd == "index-create") {
             if (params.size() < 2) {
                 printError("usage: index-create <collection> <field>");

+ 25 - 0
client/include/smartbotic/database/client.hpp

@@ -666,6 +666,31 @@ public:
     };
     [[nodiscard]] std::vector<IndexDefinition> listIndexes(const std::string& collection);
 
+    /**
+     * Distinct values an indexed field holds, with how many rows hold each. v2.10.0.
+     *
+     * Answered from the index with one step per distinct VALUE rather than per row,
+     * so "which statuses exist" over 10k rows costs four steps. Useful for filter
+     * dropdowns, and for judging whether a filter on this field will be selective
+     * enough for the planner to use its index at all.
+     *
+     * `ascending=false` walks from the highest value down, so limit=1 returns the
+     * MAXIMUM; ascending with limit=1 returns the minimum.
+     *
+     * Bounded on purpose: each distinct value costs one document read, because the
+     * value is returned as stored rather than reconstructed from the index key.
+     * Returns empty for an unindexed field, and for an array-valued field, whose
+     * keys are elements rather than values.
+     */
+    struct IndexValue {
+        nlohmann::json value;
+        uint64_t count = 0;
+    };
+    [[nodiscard]] std::vector<IndexValue> indexValues(const std::string& collection,
+                                                      const std::string& field,
+                                                      uint32_t limit = 100,
+                                                      bool ascending = true);
+
     /**
      * List all views.
      */

+ 38 - 0
client/src/client.cpp

@@ -1232,6 +1232,38 @@ public:
         return out;
     }
 
+    std::vector<Client::IndexValue> indexValues(const std::string& collection,
+                                                const std::string& field,
+                                                uint32_t limit, bool ascending) {
+        smartbotic::databasepb::GetIndexValuesRequest request;
+        request.set_collection(qualify(collection));
+        request.set_field(field);
+        request.set_limit(limit);
+        request.set_ascending(ascending);
+        smartbotic::databasepb::GetIndexValuesResponse response;
+        grpc::ClientContext context;
+        setDeadline(context);
+
+        std::vector<Client::IndexValue> out;
+        auto status = stub_->GetIndexValues(&context, request, &response);
+        if (!status.ok()) {
+            spdlog::error("Client::indexValues failed: {}", status.error_message());
+            return out;
+        }
+        out.reserve(response.values_size());
+        for (const auto& v : response.values()) {
+            Client::IndexValue iv;
+            try {
+                iv.value = nlohmann::json::parse(v.value());
+            } catch (const nlohmann::json::exception&) {
+                continue;
+            }
+            iv.count = v.count();
+            out.push_back(std::move(iv));
+        }
+        return out;
+    }
+
     bool dropView(const std::string& name) {
         smartbotic::databasepb::DropViewRequest request;
         request.set_name(qualify(name));
@@ -2159,6 +2191,12 @@ std::vector<Client::IndexDefinition> Client::listIndexes(const std::string& coll
     return impl_->listIndexes(collection);
 }
 
+std::vector<Client::IndexValue> Client::indexValues(const std::string& collection,
+                                                    const std::string& field,
+                                                    uint32_t limit, bool ascending) {
+    return impl_->indexValues(collection, field, limit, ascending);
+}
+
 std::vector<Client::ViewDefinition> Client::listViews() {
     return impl_->listViews();
 }

+ 29 - 19
docs/ROADMAP.md

@@ -1,6 +1,6 @@
 # Smartbotic Database - Status and Roadmap
 
-**Current version: 2.9.2** (see `VERSION`). Last reviewed: 2026-08-09.
+**Current version: 2.10.0** (see `VERSION`). Last reviewed: 2026-08-09.
 
 This is the single authoritative statement of what exists and what does not.
 If any other document in this repository disagrees with this one, this one is
@@ -53,7 +53,7 @@ Consequences of that substitution, which trip up readers:
 
 ## Shipped
 
-Every item below is in the installed product as of 2.9.2. `CLAUDE.md` has the
+Every item below is in the installed product as of 2.10.0. `CLAUDE.md` has the
 detail and the failure modes.
 
 - JSON document store: collections, version history, field-level encryption, TTL
@@ -67,9 +67,11 @@ detail and the failure modes.
 - Durable snapshots with tiered recovery and read-only lockout
 - Replication, events/subscribe, migrations, set operations
 - Paging fast path (no filter, no sort) and the two-pass filtered/sorted scan
-- Secondary indexes on declared fields: equality, CONTAINS, ranges, and
-  two-list intersection, with a measured selectivity guard so declaring one
-  cannot make a query slower. CLI: `indexes`, `index-create`, `index-drop`
+- Secondary indexes on declared fields: equality, IN, CONTAINS, EXISTS,
+  ranges, intersection, result ordering, filtered totals, and distinct
+  values / min / max. A measured selectivity guard means declaring an index
+  cannot make a query slower. CLI: `indexes`, `index-create`, `index-drop`,
+  `index-values`
 
 ---
 
@@ -78,21 +80,29 @@ detail and the failure modes.
 Ordered by what a reader is most likely to need next. Nothing here has a
 committed date.
 
-### 1. Indexing - remaining operators
-
-Equality, `CONTAINS` and numeric/string ranges are served from indexes, with
-two-list intersection when neither predicate is selective alone. What is left:
-
-- **`IN`** would be a union of posting lists - the mirror of the intersection
-  already implemented, and the most obviously worthwhile next one.
-- **`NE`, `REGEX`, `SEARCH`** are not index-servable in this shape. `NE` and
-  `REGEX` would need a full walk anyway; `SEARCH` reads whole documents.
-- **Ranges on a mixed-type field** decline deliberately: an index walk compares
-  by type tag while the scan compares by JSON type order, and those are not
-  assumed to agree. Serving them means defining one order and proving both paths
-  use it.
+### 1. Indexing - what is left
+
+Served from an index: equality, `IN`, `CONTAINS`, `EXISTS=true`, numeric and
+string ranges, two-list intersection, result ORDERING, filtered TOTALS, and
+distinct values / min / max. Remaining:
+
+- **Unique constraints.** Built and deliberately not shipped - see the v2.10.0
+  entry in `CLAUDE.md`. The check is correct; it cannot be enforced while
+  `applyDualWriteMirror` swallows every exception and MemoryStore mutates before
+  the mirror runs. Needs the write path restructured, which wants its own change.
+- **`NE` and `REGEX`** would need a full walk either way. A prefix-anchored
+  `REGEX` (`^abc`) could become a range over the string tag - the one real
+  opportunity here.
+- **`SEARCH`** reads whole documents by definition.
+- **Sort + unselective filter** stays a full pass, and this one is structural:
+  an exact `total_matched` requires visiting every match, which is precisely what
+  stopping early avoids. Changing it means making the total approximate - a
+  contract decision, not an optimisation.
+- **Covering reads** are closed as not viable: the response always retains six
+  metadata fields and only `_id` exists in a posting. Would need an ids-only
+  response mode.
 - **Compound (multi-column) indexes.** Intersection covers much of the benefit;
-  a real compound index would beat it when the pair is queried constantly.
+  a real compound index would beat it for a pair queried constantly.
 
 ### 2. `encode_document` still serialises with nlohmann
 

+ 24 - 0
proto/database.proto

@@ -80,6 +80,9 @@ service DatabaseService {
     rpc CreateIndex(CreateIndexRequest) returns (CreateIndexResponse);
     rpc DropIndex(DropIndexRequest) returns (DropIndexResponse);
     rpc ListIndexes(ListIndexesRequest) returns (ListIndexesResponse);
+    // v2.10.0 — the distinct values an indexed field holds, with row counts.
+    // Answered from the index, one step per distinct value rather than per row.
+    rpc GetIndexValues(GetIndexValuesRequest) returns (GetIndexValuesResponse);
     rpc GetCollectionConfig(GetCollectionConfigRequest) returns (GetCollectionConfigResponse);
     rpc MigrateCollectionTimestamps(MigrateCollectionTimestampsRequest) returns (MigrateCollectionTimestampsResponse);
 
@@ -969,6 +972,27 @@ message ListIndexesResponse {
     repeated IndexInfo indexes = 3;
 }
 
+message GetIndexValuesRequest {
+    string collection = 1;
+    string field = 2;
+    // Bounded on purpose: reading the values costs one document read each.
+    uint32 limit = 3;
+    // false walks from the highest value down, so limit=1 gives the maximum.
+    bool ascending = 4;
+}
+
+message IndexValueCount {
+    // The value as stored, JSON-encoded.
+    string value = 1;
+    uint64 count = 2;
+}
+
+message GetIndexValuesResponse {
+    bool success = 1;
+    string error = 2;
+    repeated IndexValueCount values = 3;
+}
+
 message GetCollectionConfigRequest {
     string collection = 1;
 }

+ 44 - 0
service/src/database_grpc_impl.cpp

@@ -2908,6 +2908,50 @@ grpc::Status DatabaseGrpcImpl::ListIndexes(
     }
 }
 
+
+grpc::Status DatabaseGrpcImpl::GetIndexValues(
+    grpc::ServerContext* context,
+    const pb::GetIndexValuesRequest* request,
+    pb::GetIndexValuesResponse* response
+) {
+    smartbotic::database::Decision dec;
+    if (auto st = gate(context, request->collection(),
+                       smartbotic::database::Access::Read, dec); !st.ok()) {
+        return st;
+    }
+    // A masked field's values are precisely what the mask hides, so listing them
+    // would be the inference channel the mask exists to close.
+    if (!dec.mask.empty() &&
+        std::find(dec.mask.begin(), dec.mask.end(), request->field()) != dec.mask.end()) {
+        response->set_success(false);
+        response->set_error("field is masked for this principal");
+        return grpc::Status::OK;
+    }
+    try {
+        const auto rc = smartbotic::database::resolveCollection(request->collection());
+        auto* lmdb = dynamic_cast<smartbotic::db::storage::LmdbDocumentStore*>(
+            service_.docStore(rc.project));
+        if (lmdb == nullptr) {
+            response->set_success(false);
+            response->set_error("secondary indexes require the LMDB substrate");
+            return grpc::Status::OK;
+        }
+        const size_t limit = request->limit() == 0 ? 100 : request->limit();
+        for (const auto& v : lmdb->index_values(rc.collection, request->field(),
+                                                limit, request->ascending())) {
+            auto* out = response->add_values();
+            out->set_value(v.value.dump());
+            out->set_count(v.count);
+        }
+        response->set_success(true);
+        return grpc::Status::OK;
+    } catch (const std::invalid_argument& e) {
+        return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT, e.what());
+    } catch (const std::exception& e) {
+        return grpc::Status(grpc::StatusCode::INTERNAL, e.what());
+    }
+}
+
 grpc::Status DatabaseGrpcImpl::ConfigureCollection(
     grpc::ServerContext* context,
     const pb::ConfigureCollectionRequest* request,

+ 6 - 0
service/src/database_grpc_impl.hpp

@@ -283,6 +283,12 @@ public:
         pb::ListIndexesResponse* response
     ) override;
 
+    grpc::Status GetIndexValues(
+        grpc::ServerContext* context,
+        const pb::GetIndexValuesRequest* request,
+        pb::GetIndexValuesResponse* response
+    ) override;
+
     grpc::Status ConfigureCollection(
         grpc::ServerContext* context,
         const pb::ConfigureCollectionRequest* request,

+ 372 - 13
service/src/storage/document_store_lmdb.cpp

@@ -37,6 +37,7 @@
 #include <cctype>
 #include <cstring>
 #include <functional>
+#include <limits>
 #include <regex>
 #include <sstream>
 #include <stdexcept>
@@ -529,6 +530,7 @@ void LmdbDocumentStore::maintainIndexes(
 
     const auto fields = indexed_fields(collection);
     if (fields.empty()) return;   // unindexed collections pay nothing
+    const auto uniques = unique_fields(collection);
 
     // Old keys come from the STORED bytes, parsed with yyjson and resolved one
     // field at a time - never materialised into a Document. On a 505 MB
@@ -553,12 +555,14 @@ void LmdbDocumentStore::maintainIndexes(
 
     for (const auto& f : fields) {
         std::vector<std::string> new_k;
+        bool new_is_array = false;
         if (new_doc != nullptr) {
             // resolveFilterValue is the SAME resolver the scan uses, so an
             // indexed lookup and a scan cannot disagree about which field a
             // filter names.
             if (auto v = filter_eval::resolveFilterValue(*new_doc, f)) {
                 new_k = encode_index_keys(*v);
+                new_is_array = v->is_array();
             }
         }
         std::vector<std::string> old_k;
@@ -581,6 +585,11 @@ void LmdbDocumentStore::maintainIndexes(
         const unsigned int dbi = open_for_write(wtxn, sub, MDB_DUPSORT);
         to_cache.emplace_back(sub, dbi);
 
+        // An array value means one row owns several postings, which breaks the
+        // "duplicate count == match count" identity a count depends on. Record it
+        // once, in the same transaction, so a count can check cheaply.
+        if (new_is_array) markIndexMultiValued(wtxn, dbi);
+
         for (const auto& key : to_remove) {
             MDB_val k = to_val(key);
             MDB_val v = to_val(id);
@@ -589,6 +598,35 @@ void LmdbDocumentStore::maintainIndexes(
             const int rc = mdb_del(wtxn.raw(), dbi, &k, &v);
             if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) throw_mdb(rc, "index del");
         }
+        // UNIQUE enforcement, before any posting is written. Checked here rather
+        // than at a handler because this is the only place that knows the value's
+        // encoded key, and because it must share the document's transaction: a
+        // rejection must leave neither the row nor an index entry behind.
+        if (std::find(uniques.begin(), uniques.end(), f) != uniques.end()) {
+            for (const auto& key : to_add) {
+                MDB_cursor* cur = nullptr;
+                mdb_check(mdb_cursor_open(wtxn.raw(), dbi, &cur),
+                          "cursor_open (unique check)");
+                struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+                MDB_val ck = to_val(key);
+                MDB_val cv{0, nullptr};
+                int crc = mdb_cursor_get(cur, &ck, &cv, MDB_SET);
+                while (crc == MDB_SUCCESS) {
+                    const auto holder = to_sv(cv);
+                    // Its own posting is not a conflict - that is this row being
+                    // rewritten with the value it already had.
+                    if (holder != id) {
+                        throw UniqueViolation(std::string(collection), f,
+                                              std::string(holder));
+                    }
+                    crc = mdb_cursor_get(cur, &ck, &cv, MDB_NEXT_DUP);
+                }
+                if (crc != MDB_SUCCESS && crc != MDB_NOTFOUND) {
+                    throw_mdb(crc, "cursor (unique check)");
+                }
+            }
+        }
+
         for (const auto& key : to_add) {
             MDB_val k = to_val(key);
             MDB_val v = to_val(id);
@@ -686,6 +724,8 @@ LmdbDocumentStore::IndexPlanStats LmdbDocumentStore::index_plan_stats() const {
     st.ordered_scans = ordered_scans_.load(std::memory_order_relaxed);
     st.union_scans = union_scans_.load(std::memory_order_relaxed);
     st.exists_scans = exists_scans_.load(std::memory_order_relaxed);
+    st.counted_scans = counted_scans_.load(std::memory_order_relaxed);
+    st.range_ordered_scans = range_ordered_scans_.load(std::memory_order_relaxed);
     return st;
 }
 
@@ -698,6 +738,70 @@ void LmdbDocumentStore::reset_index_plan_stats() {
     ordered_scans_.store(0, std::memory_order_relaxed);
     union_scans_.store(0, std::memory_order_relaxed);
     exists_scans_.store(0, std::memory_order_relaxed);
+    counted_scans_.store(0, std::memory_order_relaxed);
+    range_ordered_scans_.store(0, std::memory_order_relaxed);
+}
+
+
+std::vector<LmdbDocumentStore::IndexValue>
+LmdbDocumentStore::index_values(std::string_view collection,
+                                 const std::string& field,
+                                 size_t limit,
+                                 bool ascending) {
+    std::vector<IndexValue> out;
+    if (limit == 0) return out;
+
+    ReadTxn rtxn(env_);
+    auto idbi = try_open_for_read(rtxn, index_subdb_name(collection, field));
+    if (!idbi) return out;
+    auto cdbi = try_open_for_read(rtxn, collection);
+    if (!cdbi) return out;
+
+    MDB_cursor* cur = nullptr;
+    mdb_check(mdb_cursor_open(rtxn.raw(), *idbi, &cur), "cursor_open (index values)");
+    struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+
+    MDB_val k{0, nullptr};
+    MDB_val v{0, nullptr};
+    int rc = mdb_cursor_get(cur, &k, &v, ascending ? MDB_FIRST : MDB_LAST);
+    while (rc == MDB_SUCCESS && out.size() < limit) {
+        if (!is_index_meta_key(to_sv(k))) {
+            size_t n = 0;
+            mdb_check(mdb_cursor_count(cur, &n), "cursor_count (index values)");
+
+            // Read the value from a row that holds it. Walking backwards lands on
+            // a key's LAST duplicate, which is just as good - any holder has the
+            // same value, that being what the key means.
+            MDB_val hk = to_val(to_sv(v));
+            MDB_val hv{0, nullptr};
+            if (mdb_get(rtxn.raw(), *cdbi, &hk, &hv) == MDB_SUCCESS) {
+                const auto bytes = to_sv(hv);
+                yyjson_doc* d = yyjson_read(bytes.data(), bytes.size(), 0);
+                if (d) {
+                    auto val = resolve_from_yyjson(yyjson_doc_get_root(d), field);
+                    yyjson_doc_free(d);
+                    if (val) {
+                        IndexValue iv;
+                        // An array field contributes one posting per element, so
+                        // the row's value is the whole array while this key names
+                        // one element. Reporting the array would be wrong, so a
+                        // multivalued index reports nothing rather than something
+                        // misleading.
+                        if (!val->is_array()) {
+                            iv.value = std::move(*val);
+                            iv.count = static_cast<uint64_t>(n);
+                            out.push_back(std::move(iv));
+                        }
+                    }
+                }
+            }
+        }
+        rc = mdb_cursor_get(cur, &k, &v, ascending ? MDB_NEXT_NODUP : MDB_PREV_NODUP);
+    }
+    if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) {
+        throw_mdb(rc, "cursor step (index values)");
+    }
+    return out;
 }
 
 std::optional<LmdbDocumentStore::IndexStats>
@@ -715,9 +819,6 @@ LmdbDocumentStore::index_stats(std::string_view collection,
     // Reporting it would overstate every index by exactly one, the same
     // off-by-one count() already subtracts for document sub-dbs.
     out.entries = st.ms_entries;
-    if (!read_subdb_identity(rtxn, *dbi_opt).empty() && out.entries > 0) {
-        --out.entries;
-    }
 
     // Distinct keys need a walk: MDB_NEXT_NODUP skips a key's remaining
     // duplicates, so this costs one step per distinct value, not per posting.
@@ -726,12 +827,21 @@ LmdbDocumentStore::index_stats(std::string_view collection,
     struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
     MDB_val k{0, nullptr};
     MDB_val v{0, nullptr};
+    // Subtract the reserved keys from `entries` by counting them as we walk,
+    // rather than probing for each: they are the only NUL-prefixed keys and sort
+    // before every encoded value, so this sees them first.
+    uint64_t meta = 0;
     int rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST);
     while (rc == MDB_SUCCESS) {
-        if (!is_identity_key(to_sv(k))) ++out.distinct_values;
+        if (is_index_meta_key(to_sv(k))) {
+            ++meta;
+        } else {
+            ++out.distinct_values;
+        }
         rc = mdb_cursor_get(cur, &k, &v, MDB_NEXT_NODUP);
     }
     if (rc != MDB_NOTFOUND) throw_mdb(rc, "cursor next_nodup (index stat)");
+    out.entries = out.entries >= meta ? out.entries - meta : 0;
     return out;
 }
 
@@ -764,7 +874,7 @@ LmdbDocumentStore::lookupIndexRangeTxn(ReadTxn& rtxn,
     // ordered, so one tag at both ends means one tag throughout. The identity
     // sentinel sorts before every tag (it starts with NUL), so skip it.
     int rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST);
-    while (rc == MDB_SUCCESS && is_identity_key(to_sv(k))) {
+    while (rc == MDB_SUCCESS && is_index_meta_key(to_sv(k))) {
         rc = mdb_cursor_get(cur, &k, &v, MDB_NEXT_NODUP);
     }
     if (rc == MDB_NOTFOUND) return std::vector<std::string>{};  // empty index
@@ -796,7 +906,7 @@ LmdbDocumentStore::lookupIndexRangeTxn(ReadTxn& rtxn,
 
     while (rc == MDB_SUCCESS) {
         const auto key = to_sv(k);
-        if (!is_identity_key(key)) {
+        if (!is_index_meta_key(key)) {
             if (index_key_tag(key) != first_tag) break;    // left the type
             if (!ascending) {
                 const int cmp = key.compare(bsv);
@@ -851,6 +961,7 @@ uint64_t LmdbDocumentStore::build_index(std::string_view collection,
             if (!val) continue;                       // field absent on this row
             const auto keys = encode_index_keys(*val);
             if (keys.empty()) continue;               // unindexable value
+            if (val->is_array()) markIndexMultiValued(wtxn, dbi);
             for (const auto& key : keys) {
                 MDB_val kk = to_val(key);
                 MDB_val vv = to_val(id);
@@ -1001,6 +1112,245 @@ uint64_t LmdbDocumentStore::rowCountTxn(ReadTxn& rtxn, unsigned int dbi) {
     return n;
 }
 
+
+
+void LmdbDocumentStore::set_unique_fields(std::string_view collection,
+                                           std::vector<std::string> fields) {
+    std::lock_guard<std::mutex> lock(index_mutex_);
+    if (fields.empty()) {
+        unique_fields_.erase(std::string(collection));
+    } else {
+        unique_fields_[std::string(collection)] = std::move(fields);
+    }
+}
+
+std::vector<std::string>
+LmdbDocumentStore::unique_fields(std::string_view collection) {
+    std::lock_guard<std::mutex> lock(index_mutex_);
+    auto it = unique_fields_.find(std::string(collection));
+    if (it == unique_fields_.end()) return {};
+    return it->second;
+}
+
+std::vector<LmdbDocumentStore::DuplicateValue>
+LmdbDocumentStore::find_duplicate_values(std::string_view collection,
+                                          const std::string& field,
+                                          size_t limit) {
+    std::vector<DuplicateValue> out;
+    ReadTxn rtxn(env_);
+    auto dbi_opt = try_open_for_read(rtxn, index_subdb_name(collection, field));
+    if (!dbi_opt) return out;
+
+    MDB_cursor* cur = nullptr;
+    mdb_check(mdb_cursor_open(rtxn.raw(), *dbi_opt, &cur), "cursor_open (dupes)");
+    struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+
+    MDB_val k{0, nullptr};
+    MDB_val v{0, nullptr};
+    int rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST);
+    while (rc == MDB_SUCCESS && out.size() < limit) {
+        if (!is_index_meta_key(to_sv(k))) {
+            size_t n = 0;
+            mdb_check(mdb_cursor_count(cur, &n), "cursor_count (dupes)");
+            if (n > 1) {
+                DuplicateValue d;
+                d.count = static_cast<uint64_t>(n);
+                d.sample_id = std::string(to_sv(v));
+                out.push_back(std::move(d));
+            }
+        }
+        rc = mdb_cursor_get(cur, &k, &v, MDB_NEXT_NODUP);
+    }
+    if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) throw_mdb(rc, "cursor (dupes)");
+    return out;
+}
+
+void LmdbDocumentStore::markIndexMultiValued(WriteTxn& wtxn, unsigned int dbi) {
+    MDB_val k{kIndexMultiValuedKey.size(),
+              const_cast<char*>(kIndexMultiValuedKey.data())};
+    MDB_val v{1, const_cast<char*>("1")};
+    // MDB_NODUPDATA so re-marking is a no-op rather than an error.
+    const int rc = mdb_put(wtxn.raw(), dbi, &k, &v, MDB_NODUPDATA);
+    if (rc != MDB_SUCCESS && rc != MDB_KEYEXIST) throw_mdb(rc, "mark multivalued");
+}
+
+bool LmdbDocumentStore::indexIsMultiValuedTxn(ReadTxn& rtxn,
+                                               std::string_view collection,
+                                               const std::string& field) {
+    auto dbi_opt = try_open_for_read(rtxn, index_subdb_name(collection, field));
+    if (!dbi_opt) return true;      // unknown: assume the worst
+    MDB_val k{kIndexMultiValuedKey.size(),
+              const_cast<char*>(kIndexMultiValuedKey.data())};
+    MDB_val v{0, nullptr};
+    return mdb_get(rtxn.raw(), *dbi_opt, &k, &v) == MDB_SUCCESS;
+}
+
+
+
+bool LmdbDocumentStore::rangeOrderedFromIndex(
+    ReadTxn& rtxn,
+    std::string_view collection,
+    const smartbotic::database::Query& query,
+    unsigned int coll_dbi,
+    ScanResult& result) {
+
+    using Op = smartbotic::database::FilterOp;
+    if (query.filters.size() != 1) return false;
+    if (!query.sort || query.sort->field.empty()) return false;
+
+    const auto& f = query.filters.front();
+    if (f.op != Op::GT && f.op != Op::GTE && f.op != Op::LT && f.op != Op::LTE) {
+        return false;
+    }
+    if (f.field != query.sort->field) return false;   // the walk's order IS the sort
+
+    const auto fields = indexed_fields(collection);
+    if (std::find(fields.begin(), fields.end(), f.field) == fields.end()) return false;
+    // One posting per row, or the count is not the match count.
+    if (indexIsMultiValuedTxn(rtxn, collection, f.field)) return false;
+
+    // No budget: the walk collects ids only, and the total needs all of them.
+    // Passing UINT64_MAX asks for the whole range rather than declining on size.
+    auto ids = lookupIndexRangeTxn(rtxn, collection, f.field, f.op, f.value,
+                                   std::numeric_limits<uint64_t>::max());
+    if (!ids) return false;   // no index, unindexable bound, or mixed types
+
+    // lookupIndexRangeTxn walks ascending. Reversing gives descending keys AND
+    // descending ids within a key, which is exactly how sort_documents breaks a
+    // tie when descending.
+    if (query.sort->descending) std::reverse(ids->begin(), ids->end());
+
+    const uint64_t total = ids->size();
+    const uint64_t start = std::min<uint64_t>(query.offset, total);
+    const uint64_t end = std::min<uint64_t>(
+        static_cast<uint64_t>(query.offset) + query.limit, total);
+
+    result.documents.clear();
+    result.documents.reserve(end > start ? end - start : 0);
+    for (uint64_t i = start; i < end; ++i) {
+        MDB_val hk = to_val((*ids)[i]);
+        MDB_val hv{0, nullptr};
+        if (mdb_get(rtxn.raw(), coll_dbi, &hk, &hv) != MDB_SUCCESS) continue;
+        result.documents.push_back(decode_document(to_sv(hv)));
+    }
+    result.total_matched = total;
+    result.has_more = end < total;
+
+    if (!query.projection.empty()) {
+        for (auto& d : result.documents) {
+            nlohmann::json data = d.data();
+            apply_projection_inplace(data, query.projection);
+            d.set_data(data);
+        }
+    }
+    range_ordered_scans_.fetch_add(1, std::memory_order_relaxed);
+    return true;
+}
+
+bool LmdbDocumentStore::countAndPageFromIndex(
+    ReadTxn& rtxn,
+    std::string_view collection,
+    const smartbotic::database::Query& query,
+    unsigned int coll_dbi,
+    ScanResult& result) {
+
+    using Op = smartbotic::database::FilterOp;
+    // Exactly one predicate: with two, the total is the size of an intersection
+    // that no single index knows.
+    if (query.filters.size() != 1) return false;
+    if (query.sort && !query.sort->field.empty()) return false;   // order differs
+
+    const auto& f = query.filters.front();
+    if (f.op != Op::EQ && f.op != Op::IN) return false;
+
+    const auto fields = indexed_fields(collection);
+    if (std::find(fields.begin(), fields.end(), f.field) == fields.end()) return false;
+
+    // The count is only the match count while every row owns at most one posting.
+    // Once an array has been indexed, postings under a key can name rows whose
+    // value merely CONTAINS it, and EQ compares the whole value.
+    if (indexIsMultiValuedTxn(rtxn, collection, f.field)) return false;
+
+    const uint64_t want = static_cast<uint64_t>(query.offset) + query.limit;
+    std::vector<std::string> page_ids;
+    uint64_t total = 0;
+
+    if (f.op == Op::EQ) {
+        auto n = countIndexEqTxn(rtxn, collection, f.field, f.value);
+        if (!n) return false;                     // no index, or unindexable value
+        total = *n;
+        // Walk only as far as the page needs. The ids under one key are ascending,
+        // which is the order a collection cursor would have produced.
+        if (total > 0 && want > 0) {
+            auto dbi_opt = try_open_for_read(rtxn,
+                                             index_subdb_name(collection, f.field));
+            if (!dbi_opt) return false;
+            const auto key = encode_index_key(f.value);
+            if (!key) return false;
+            MDB_cursor* cur = nullptr;
+            mdb_check(mdb_cursor_open(rtxn.raw(), *dbi_opt, &cur),
+                      "cursor_open (index page)");
+            struct G { MDB_cursor* c; ~G() { if (c) mdb_cursor_close(c); } } g{cur};
+            MDB_val k = to_val(*key);
+            MDB_val v{0, nullptr};
+            int rc = mdb_cursor_get(cur, &k, &v, MDB_SET);
+            if (rc == MDB_SUCCESS) rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST_DUP);
+            uint64_t seen = 0;
+            while (rc == MDB_SUCCESS && page_ids.size() < query.limit) {
+                if (seen >= query.offset) page_ids.emplace_back(to_sv(v));
+                ++seen;
+                rc = mdb_cursor_get(cur, &k, &v, MDB_NEXT_DUP);
+            }
+            if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) {
+                throw_mdb(rc, "cursor next_dup (index page)");
+            }
+        }
+    } else {
+        // IN. A non-multivalued field gives each row one value, so a row appears
+        // under exactly one of the requested keys - the per-value counts are
+        // disjoint and their sum is the exact total.
+        if (!f.value.is_array()) return false;
+        std::vector<std::string> all;
+        for (const auto& el : f.value) {
+            auto n = countIndexEqTxn(rtxn, collection, f.field, el);
+            if (!n) return false;
+            total += *n;
+            if (want == 0) continue;
+            auto part = lookupIndexEqTxn(rtxn, collection, f.field, el);
+            if (!part) return false;
+            all.insert(all.end(), part->begin(), part->end());
+        }
+        // Merging is needed because the page must be in id order across values.
+        // Only ids are read, never documents, so this stays cheap.
+        std::sort(all.begin(), all.end());
+        const uint64_t start = std::min<uint64_t>(query.offset, all.size());
+        const uint64_t end = std::min<uint64_t>(want, all.size());
+        page_ids.assign(all.begin() + start, all.begin() + end);
+    }
+
+    // Materialise ONLY the page.
+    result.documents.clear();
+    result.documents.reserve(page_ids.size());
+    for (const auto& id : page_ids) {
+        MDB_val hk = to_val(id);
+        MDB_val hv{0, nullptr};
+        if (mdb_get(rtxn.raw(), coll_dbi, &hk, &hv) != MDB_SUCCESS) continue;
+        result.documents.push_back(decode_document(to_sv(hv)));
+    }
+    result.total_matched = total;
+    result.has_more = std::min<uint64_t>(want, total) < total;
+
+    if (!query.projection.empty()) {
+        for (auto& d : result.documents) {
+            nlohmann::json data = d.data();
+            apply_projection_inplace(data, query.projection);
+            d.set_data(data);
+        }
+    }
+    counted_scans_.fetch_add(1, std::memory_order_relaxed);
+    return true;
+}
+
 bool LmdbDocumentStore::orderedHitsFromIndex(
     ReadTxn& rtxn,
     std::string_view collection,
@@ -1026,11 +1376,8 @@ bool LmdbDocumentStore::orderedHitsFromIndex(
     // Total order or nothing. entries != rows means some row has no posting (the
     // field is absent, or its value is unindexable) or several (an array), and in
     // either case the index cannot state where that row sorts.
-    MDB_stat ist{};
-    if (mdb_stat(rtxn.raw(), *dbi_opt, &ist) != MDB_SUCCESS) return false;
-    uint64_t entries = ist.ms_entries;
-    if (!read_subdb_identity(rtxn, *dbi_opt).empty() && entries > 0) --entries;
-    if (entries != rows) return false;
+    const auto istats = index_stats(collection, query.sort->field);
+    if (!istats || istats->entries != rows) return false;
 
     MDB_cursor* cur = nullptr;
     mdb_check(mdb_cursor_open(rtxn.raw(), *dbi_opt, &cur), "cursor_open (index order)");
@@ -1049,7 +1396,7 @@ bool LmdbDocumentStore::orderedHitsFromIndex(
     int rc = mdb_cursor_get(cur, &k, &v, desc ? MDB_LAST : MDB_FIRST);
     std::vector<std::string> ids;
     while (rc == MDB_SUCCESS && ids.size() < want) {
-        if (!is_identity_key(to_sv(k))) ids.emplace_back(to_sv(v));
+        if (!is_index_meta_key(to_sv(k))) ids.emplace_back(to_sv(v));
         rc = mdb_cursor_get(cur, &k, &v, desc ? MDB_PREV : MDB_NEXT);
     }
     if (rc != MDB_SUCCESS && rc != MDB_NOTFOUND) {
@@ -1102,7 +1449,7 @@ LmdbDocumentStore::lookupIndexAllTxn(ReadTxn& rtxn,
     MDB_val v{0, nullptr};
     int rc = mdb_cursor_get(cur, &k, &v, MDB_FIRST);
     while (rc == MDB_SUCCESS) {
-        if (!is_identity_key(to_sv(k))) {
+        if (!is_index_meta_key(to_sv(k))) {
             ids.emplace_back(to_sv(v));
             // Bail out rather than truncate, as everywhere else: a partial
             // candidate list drops matching rows.
@@ -1401,6 +1748,18 @@ ScanResult LmdbDocumentStore::scan(std::string_view collection,
             yyjson_doc_free(ydoc);
         };
 
+        // v2.10.0 — count AND page from the index. Cheapest of all when it
+        // applies: nothing but the page is read, and no selectivity guard is
+        // needed because the total never requires visiting a match.
+        if (countAndPageFromIndex(rtxn, collection, query, *dbi_opt, result)) {
+            return result;
+        }
+        // A range on the field being sorted: the walk's order already IS the sort
+        // order, so no sorting pass and only the page is decoded.
+        if (rangeOrderedFromIndex(rtxn, collection, query, *dbi_opt, result)) {
+            return result;
+        }
+
         // v2.9.2 — ORDERING from the index. Try this first: when it applies it
         // reads the page and nothing else, where every other plan still visits
         // every matching row.

+ 121 - 0
service/src/storage/document_store_lmdb.hpp

@@ -15,6 +15,7 @@
 #include <atomic>
 #include <mutex>
 #include <optional>
+#include <stdexcept>
 #include <string>
 #include <string_view>
 #include <unordered_map>
@@ -28,6 +29,23 @@ namespace smartbotic::db::storage {
 
 class LmdbEnv;
 
+// v2.10.0 — thrown when a write would put two rows under one value of a UNIQUE
+// index. A distinct type so the gRPC layer can answer ALREADY_EXISTS instead of
+// INTERNAL: a duplicate is the caller's business, not a server fault.
+class UniqueViolation : public std::runtime_error {
+public:
+    UniqueViolation(std::string collection, std::string field, std::string existing_id)
+        : std::runtime_error("unique constraint violated on " + collection + "#" +
+                             field + ": value already held by document '" +
+                             existing_id + "'"),
+          collection(std::move(collection)),
+          field(std::move(field)),
+          existing_id(std::move(existing_id)) {}
+    std::string collection;
+    std::string field;
+    std::string existing_id;
+};
+
 class LmdbDocumentStore : public DocumentStore {
 public:
     explicit LmdbDocumentStore(LmdbEnv& env);
@@ -58,6 +76,40 @@ public:
                             std::vector<std::string> fields);
     std::vector<std::string> indexed_fields(std::string_view collection);
 
+    // ⚠ v2.10.0 — UNREACHABLE BY DESIGN, pending write-path work. Nothing calls
+    // set_unique_fields, so the enforcement below never runs, and no RPC or config
+    // key exposes it. Kept because the mechanism is correct in isolation and is
+    // the base for finishing the feature.
+    //
+    // Why it cannot be turned on yet: the check throws from put(), which runs
+    // inside applyDualWriteMirror - and that catches every exception, logs it,
+    // bumps mirror drift and flips mirror_healthy_. So a rejection would be
+    // swallowed (the row still lands in MemoryStore, unenforced) AND would send
+    // every read in the process to MemoryStore, which is the v2.8.1 fault. Worse,
+    // MemoryStore mutates BEFORE the mirror runs, so a clean rejection needs the
+    // in-memory write rolled back.
+    //
+    // Finishing it means either propagating UniqueViolation through the mirror
+    // without touching health, plus rollback, or moving enforcement ahead of the
+    // MemoryStore mutation under the same collection lock. Both touch the write
+    // path that produced the v2.4.3, v2.4.4 and v2.8.0 incidents, so it wants its
+    // own change. See docs/ROADMAP.md.
+    void set_unique_fields(std::string_view collection,
+                          std::vector<std::string> fields);
+    std::vector<std::string> unique_fields(std::string_view collection);
+
+    // Existing duplicate values for a field, up to `limit` examples. Empty means
+    // the field can carry a unique constraint. Used to REFUSE declaring
+    // uniqueness over data that already violates it, rather than accepting a
+    // constraint that is already false.
+    struct DuplicateValue {
+        std::string sample_id;      // one document holding the value
+        uint64_t count = 0;         // how many hold it
+    };
+    std::vector<DuplicateValue> find_duplicate_values(std::string_view collection,
+                                                       const std::string& field,
+                                                       size_t limit = 5);
+
     // Doc ids whose `field` equals `value`, straight from the index. nullopt
     // means "no index for this field" - the caller must fall back to a scan
     // rather than treat it as "no matches", which would be silently wrong.
@@ -86,10 +138,34 @@ public:
         uint64_t ordered_scans = 0;          // the index supplied the ORDER
         uint64_t union_scans = 0;            // IN unioned posting lists
         uint64_t exists_scans = 0;           // EXISTS served from all postings
+        uint64_t counted_scans = 0;          // total AND page from the index alone
+        uint64_t range_ordered_scans = 0;    // a range walk supplied the order too
     };
     IndexPlanStats index_plan_stats() const;
     void reset_index_plan_stats();
 
+    // v2.10.0 — the DISTINCT values a field holds, in ascending order, with how
+    // many rows hold each. Walks one step per distinct value (MDB_NEXT_NODUP), not
+    // one per row, so a low-cardinality field is answered almost for free -
+    // "which statuses exist" over 10k rows costs four steps.
+    //
+    // The value itself is read from one row holding it rather than decoded from the
+    // index key. That keeps the encoding write-only (no decoder to drift from the
+    // encoder) and is exact by construction: it returns the value as stored, not a
+    // reconstruction of it. The cost is one document read per distinct value, which
+    // is why `limit` is not optional in spirit - ask for a bounded number.
+    //
+    // `ascending=false` walks from the end, so limit=1 gives the MAXIMUM and
+    // ascending limit=1 gives the MINIMUM.
+    struct IndexValue {
+        nlohmann::json value;
+        uint64_t count = 0;
+    };
+    std::vector<IndexValue> index_values(std::string_view collection,
+                                          const std::string& field,
+                                          size_t limit,
+                                          bool ascending = true);
+
     // Size of one index, read from the index itself rather than by scanning.
     // distinct_values lets an operator judge selectivity: postings concentrated
     // in few values mean the planner will usually decline the index.
@@ -162,6 +238,7 @@ private:
     // consults it on every put and must not contend with dbi cache lookups.
     std::mutex index_mutex_;
     std::unordered_map<std::string, std::vector<std::string>> indexed_fields_;
+    std::unordered_map<std::string, std::vector<std::string>> unique_fields_;
     mutable std::atomic<uint64_t> indexed_scans_{0};
     mutable std::atomic<uint64_t> full_scans_{0};
     mutable std::atomic<uint64_t> declined_unselective_{0};
@@ -170,6 +247,8 @@ private:
     mutable std::atomic<uint64_t> ordered_scans_{0};
     mutable std::atomic<uint64_t> union_scans_{0};
     mutable std::atomic<uint64_t> exists_scans_{0};
+    mutable std::atomic<uint64_t> counted_scans_{0};
+    mutable std::atomic<uint64_t> range_ordered_scans_{0};
     std::unordered_map<std::string, unsigned int> dbi_cache_;
 
     // Resolve a collection name to an MDB_dbi handle.
@@ -219,6 +298,12 @@ private:
                         const nlohmann::json& value,
                         uint64_t budget);
 
+    // v2.9.3 — record that an index has taken an array value, and read it back.
+    // A count can only be trusted while this is false; see kIndexMultiValuedKey.
+    void markIndexMultiValued(class WriteTxn& wtxn, unsigned int dbi);
+    bool indexIsMultiValuedTxn(class ReadTxn& rtxn, std::string_view collection,
+                               const std::string& field);
+
     // Rows in a collection sub-db, excluding the v2.4.4 identity sentinel.
     uint64_t rowCountTxn(class ReadTxn& rtxn, unsigned int dbi);
 
@@ -248,6 +333,42 @@ private:
                               std::vector<std::string>& out,
                               uint64_t& exact_total);
 
+    // v2.10.0 — answer a query's TOTAL and its page from the index alone.
+    //
+    // The remaining full scans were not about finding rows, they were about
+    // COUNTING them: total_matched is part of the contract, and counting matches
+    // meant visiting every one. That is why status="completed" (66% of rows) was
+    // declined - not because the rows were hard to find, but because 6,689 of them
+    // had to be counted.
+    //
+    // For a single EQ or IN predicate on a non-multivalued index the count is
+    // already in the index (mdb_cursor_count), so the total costs nothing and only
+    // the page's rows are read. No selectivity budget applies: an unselective
+    // predicate is now as cheap as a selective one.
+    //
+    // Page ORDER matches the general path without extra work: DUPSORT stores a
+    // key's ids ascending and document keys are ids, so postings come out in the
+    // same order a collection cursor would produce. Only valid with no sort.
+    //
+    // Fills `result` and returns true, or returns false to decline.
+    bool countAndPageFromIndex(class ReadTxn& rtxn, std::string_view collection,
+                               const smartbotic::database::Query& query,
+                               unsigned int coll_dbi, ScanResult& result);
+
+    // v2.10.0 — a RANGE on the same field the query sorts by.
+    //
+    // A range walk already emerges in key order, so it IS the sort order: no
+    // separate sorting pass, and only the page is materialised. Total is exact
+    // because on a non-multivalued index each posting in the range is one distinct
+    // matching row.
+    //
+    // Rows lacking the field need no special guard here, unlike pure ordering: a
+    // missing value fails a range predicate anyway, so excluding them is correct
+    // rather than a silent omission.
+    bool rangeOrderedFromIndex(class ReadTxn& rtxn, std::string_view collection,
+                               const smartbotic::database::Query& query,
+                               unsigned int coll_dbi, ScanResult& result);
+
     // Every id the index holds for `field`, regardless of value. Serves
     // EXISTS=true, which is selective exactly when the field is sparse.
     std::optional<std::vector<std::string>>

+ 9 - 0
service/src/storage/secondary_index.cpp

@@ -2,6 +2,8 @@
 
 #include "storage/secondary_index.hpp"
 
+#include "storage/subdb_identity.hpp"
+
 #include <algorithm>
 #include <cmath>
 #include <cstdint>
@@ -111,6 +113,13 @@ std::optional<IndexSubdb> parse_index_subdb(std::string_view name) {
     return out;
 }
 
+bool is_index_meta_key(std::string_view key) {
+    // Both reserved keys start with NUL, which no encoded value can, so the cheap
+    // test comes first.
+    if (key.empty() || key[0] != '\0') return false;
+    return key == kIndexMultiValuedKey || is_identity_key(key);
+}
+
 unsigned char index_key_tag(std::string_view key) {
     return key.empty() ? 0 : static_cast<unsigned char>(key[0]);
 }

+ 22 - 0
service/src/storage/secondary_index.hpp

@@ -96,6 +96,28 @@ std::optional<IndexSubdb> parse_index_subdb(std::string_view name);
 //   0x40  string    raw UTF-8 bytes
 std::optional<std::string> encode_index_key(const nlohmann::json& value);
 
+// v2.9.3 — reserved key marking an index as MULTIVALUED: at least one row has
+// contributed more than one posting, because its value was an array.
+//
+// This exists so a count can be trusted. For a scalar-only field each matching
+// row owns exactly one posting under its value, so the index's duplicate count
+// for a key IS the number of rows with that value - which is the match count for
+// EQ, available without reading a single row. Once any array has been indexed
+// that identity breaks: the postings under one key can name rows whose value
+// merely CONTAINS it, and EQ compares the whole array.
+//
+// Set on write, never cleared: it is a conservative "has ever been", because
+// clearing it would need proof that no array remains anywhere in the collection.
+// Leading NUL so it cannot collide with an encoded value, which always begins
+// with a type tag.
+inline constexpr std::string_view kIndexMultiValuedKey =
+    std::string_view("\0__multivalued__", 17);
+
+// True for any reserved key in an index sub-db - the v2.4.4 identity sentinel or
+// the multivalued marker. Every walk over an index must skip these, and every
+// count must exclude them.
+bool is_index_meta_key(std::string_view key);
+
 // Leading type tag of an encoded key, 0 for an empty key. A range walk uses it to
 // stay inside one type.
 unsigned char index_key_tag(std::string_view key);

+ 319 - 12
tests/test_subdb_identity.cpp

@@ -974,9 +974,9 @@ void test_indexed_and_unindexed_plans_agree() {
         auto r = store.scan("c", q);
         auto st = store.index_plan_stats();
         check(r.total_matched == 30, "the selective query matches 30 of 300 rows");
-        check(st.indexed_scans == 1 && st.full_scans == 0,
-              "a selective EQ on an indexed field TAKES the index plan - without "
-              "this the equivalence cases above would prove nothing");
+        check(st.counted_scans == 1 && st.full_scans == 0,
+              "a single unsorted EQ is answered by the COUNTED plan - total from "
+              "the index, only the page's rows read");
     }
 
     {
@@ -987,10 +987,10 @@ void test_indexed_and_unindexed_plans_agree() {
         auto r = store.scan("c", q);
         auto st = store.index_plan_stats();
         check(r.total_matched == 150, "the unselective query matches half the rows");
-        check(st.indexed_scans == 0 && st.declined_unselective == 1,
-              "and the guard DECLINES its index - 150 of 300 rows would cost more "
-              "through the index than a scan, so declaring an index must not be "
-              "able to pessimise a query");
+        check(st.counted_scans == 1 && st.full_scans == 0,
+              "an UNSELECTIVE EQ is now cheap too. The old selectivity guard "
+              "declined it, but only because counting 150 matches meant visiting "
+              "them; with the count read from the index, breadth costs nothing");
     }
 
     {
@@ -1001,10 +1001,26 @@ void test_indexed_and_unindexed_plans_agree() {
         q.filters.push_back(F("n", Op::GT, 250));
         store.scan("c", q);
         auto st = store.index_plan_stats();
-        check(st.full_scans == 1 && st.indexed_scans == 0,
+        check(st.full_scans == 1 && st.indexed_scans == 0 && st.counted_scans == 0,
               "a query whose predicates name no indexed field scans");
     }
 
+    {
+        // The selectivity guard still governs the shapes the counted plan cannot
+        // serve. A SORT rules it out (postings come in id order, not sort order),
+        // so an unselective EQ plus a sort must still decline to the scan.
+        store.reset_index_plan_stats();
+        smartbotic::database::Query q;
+        q.limit = 10;
+        q.filters.push_back(F("bucket", Op::EQ, "even"));
+        q.sort = smartbotic::database::Sort{"n", true};
+        store.scan("c", q);
+        auto st = store.index_plan_stats();
+        check(st.counted_scans == 0 && st.declined_unselective == 1,
+              "unselective EQ + a sort still declines - the counted plan cannot "
+              "order, and the guard is what stops the index pessimising it");
+    }
+
     auto n = store.index_count_eq("c", "bucket", nlohmann::json("even"));
     check(n.has_value() && *n == 150,
           "the unselective index does exist and holds 150 of 300 rows");
@@ -1378,16 +1394,33 @@ void test_index_in_and_exists() {
     {
         store.reset_index_plan_stats();
         (void)sig({F("grp", Op::IN, nlohmann::json::array({"g1","g2","g3"}))});
-        check(store.index_plan_stats().union_scans == 1,
-              "IN over a narrow set UNIONS posting lists - 30 of 400 rows");
+        auto st = store.index_plan_stats();
+        check(st.counted_scans == 1,
+              "a single unsorted IN is answered by the COUNTED plan: the per-value "
+              "counts are disjoint on a non-multivalued field, so their sum is the "
+              "exact total");
     }
     {
         store.reset_index_plan_stats();
         (void)sig({F("grp", Op::IN, nlohmann::json::array(
               {"g0","g1","g2","g3","g4","g5","g6","g7","g8","g9","g10","g11"}))});
         auto st = store.index_plan_stats();
-        check(st.union_scans == 0 && st.full_scans == 1,
-              "a broad IN is rejected on the summed counts, before any list is read");
+        check(st.counted_scans == 1 && st.full_scans == 0,
+              "even a BROAD IN is served now - the total is a sum of counts and "
+              "only ids are merged, so no document outside the page is touched. "
+              "The union plan remains for the sorted case");
+    }
+    {
+        // The union plan still exists for the shape the counted plan declines.
+        store.reset_index_plan_stats();
+        smartbotic::database::Query q;
+        q.limit = 1000;
+        q.filters.push_back(F("grp", Op::IN, nlohmann::json::array({"g1","g2","g3"})));
+        q.sort = smartbotic::database::Sort{"grp", false};
+        store.scan("c", q);
+        auto st = store.index_plan_stats();
+        check(st.union_scans == 1,
+              "a SORTED IN still unions posting lists as filter candidates");
     }
     {
         store.reset_index_plan_stats();
@@ -1412,6 +1445,277 @@ void test_index_in_and_exists() {
     }
 }
 
+
+// v2.10.0 — the total and the page come from the index, so nothing outside the
+// page is read. The risk moves from "are the rows right" to "is the TOTAL right",
+// and a wrong total silently breaks paging rather than erroring.
+void test_counted_plan() {
+    using Op = smartbotic::database::FilterOp;
+    TmpEnv t("counted");
+    LmdbDocumentStore store(t.env);
+
+    // 200 rows, 100 of them state="done" - deliberately unselective, the case the
+    // old guard declined.
+    for (int i = 0; i < 200; ++i) {
+        Document d;
+        d.id = "r" + std::string(i < 10 ? "00" : (i < 100 ? "0" : "")) + std::to_string(i);
+        d.collection = "c";
+        d.set_data(nlohmann::json{{"state", (i % 2 == 0) ? "done" : "todo"},
+                                  {"grp", "g" + std::to_string(i % 5)}});
+        store.put("c", d.id, d);
+    }
+
+    auto F = [](const char* f, Op op, const nlohmann::json& v) {
+        smartbotic::database::Filter x;
+        x.field = f; x.op = op; x.value = v;
+        return x;
+    };
+    auto sig = [&](const smartbotic::database::Query& q) {
+        auto r = store.scan("c", q);
+        std::string out = "total=" + std::to_string(r.total_matched) +
+                          " more=" + std::to_string(r.has_more ? 1 : 0) + " [";
+        for (const auto& d : r.documents) { out += d.id; out += ","; }
+        return out + "]";
+    };
+    auto Q = [&](std::vector<smartbotic::database::Filter> fs, uint32_t limit,
+                 uint32_t offset) {
+        smartbotic::database::Query q;
+        q.filters = std::move(fs);
+        q.limit = limit;
+        q.offset = offset;
+        return q;
+    };
+
+    const std::vector<std::pair<const char*, smartbotic::database::Query>> cases = {
+        {"EQ first page",            Q({F("state", Op::EQ, "done")}, 10, 0)},
+        {"EQ deep offset",           Q({F("state", Op::EQ, "done")}, 10, 80)},
+        {"EQ last partial page",     Q({F("state", Op::EQ, "done")}, 10, 95)},
+        {"EQ offset past the end",   Q({F("state", Op::EQ, "done")}, 10, 500)},
+        {"EQ limit=0",               Q({F("state", Op::EQ, "done")}, 0, 0)},
+        {"EQ whole match set",       Q({F("state", Op::EQ, "done")}, 1000, 0)},
+        {"EQ matching nothing",      Q({F("state", Op::EQ, "nope")}, 10, 0)},
+        {"IN two values",            Q({F("grp", Op::IN, nlohmann::json::array({"g1","g3"}))}, 10, 0)},
+        {"IN paged",                 Q({F("grp", Op::IN, nlohmann::json::array({"g1","g3"}))}, 7, 55)},
+        {"IN one missing value",     Q({F("grp", Op::IN, nlohmann::json::array({"g1","zz"}))}, 10, 0)},
+    };
+
+    uint64_t counted_used = 0;
+    for (const auto& [name, q] : cases) {
+        store.set_indexed_fields("c", {});
+        const std::string without = sig(q);
+        store.set_indexed_fields("c", {"state", "grp"});
+        store.build_index("c", "state");
+        store.build_index("c", "grp");
+        store.reset_index_plan_stats();
+        const std::string with = sig(q);
+        counted_used += store.index_plan_stats().counted_scans;
+        if (with != without) {
+            std::cerr << "    scan:    " << without.substr(0, 200) << "\n"
+                      << "    counted: " << with.substr(0, 200) << "\n";
+        }
+        const std::string msg = std::string("counted plan == scan: ") + name;
+        check(with == without, msg.c_str());
+    }
+    check(counted_used == cases.size(),
+          ("the counted plan ran in all " + std::to_string(cases.size()) +
+           " cases (got " + std::to_string(counted_used) + ")").c_str());
+
+    // ---- an ARRAY-valued field must NOT be counted from the index ----
+    //
+    // EQ compares the whole array, but the index holds one posting per element, so
+    // a key's duplicate count is the number of rows CONTAINING that element - not
+    // the number whose value equals it. Counting from the index there would report
+    // a total for rows that do not match.
+    {
+        TmpEnv t2("counted-multi");
+        LmdbDocumentStore s2(t2.env);
+        for (int i = 0; i < 60; ++i) {
+            Document d;
+            d.id = "m" + std::to_string(i);
+            d.collection = "c";
+            d.set_data(nlohmann::json{{"tags", nlohmann::json::array(
+                 {"t" + std::to_string(i % 3), "all"})}});
+            s2.put("c", d.id, d);
+        }
+        auto sig2 = [&](const smartbotic::database::Query& q) {
+            auto r = s2.scan("c", q);
+            return "total=" + std::to_string(r.total_matched) +
+                   " rows=" + std::to_string(r.documents.size());
+        };
+        smartbotic::database::Query q;
+        q.limit = 100;
+        q.filters.push_back(F("tags", Op::EQ, "t1"));
+
+        s2.set_indexed_fields("c", {});
+        const std::string without = sig2(q);
+        s2.set_indexed_fields("c", {"tags"});
+        s2.build_index("c", "tags");
+        s2.reset_index_plan_stats();
+        const std::string with = sig2(q);
+
+        check(with == without,
+              "EQ on an array-valued field agrees with the scan (both find 0 - EQ "
+              "compares the WHOLE array)");
+        check(s2.index_plan_stats().counted_scans == 0,
+              "and the counted plan DECLINES, because the index is marked "
+              "multivalued: a key's duplicate count counts rows CONTAINING the "
+              "element, which is not the EQ match count");
+
+        // CONTAINS on the same data still works, via candidates.
+        smartbotic::database::Query qc;
+        qc.limit = 100;
+        qc.filters.push_back(F("tags", Op::CONTAINS, "t1"));
+        s2.set_indexed_fields("c", {});
+        const std::string cwithout = sig2(qc);
+        s2.set_indexed_fields("c", {"tags"});
+        const std::string cwith = sig2(qc);
+        check(cwith == cwithout, "CONTAINS on the same array field still agrees");
+    }
+}
+
+
+// v2.10.0 — a range on the field being sorted. The walk's order already IS the
+// sort order, so no sorting pass runs and only the page is decoded. The subtle
+// part is descending ties: reversing an ascending walk must reproduce
+// sort_documents' reverse-id tie-break exactly.
+void test_range_ordered_plan() {
+    using Op = smartbotic::database::FilterOp;
+    TmpEnv t("range-ordered");
+    LmdbDocumentStore store(t.env);
+
+    for (int i = 0; i < 250; ++i) {
+        Document d;
+        d.id = "r" + std::string(i < 10 ? "00" : (i < 100 ? "0" : "")) + std::to_string(i);
+        d.collection = "c";
+        // `dup` deliberately ties three rows per value so the tie-break matters.
+        store.put("c", d.id, [&]{
+            d.set_data(nlohmann::json{{"n", i}, {"dup", i / 3}});
+            return d;
+        }());
+    }
+
+    auto F = [](const char* f, Op op, const nlohmann::json& v) {
+        smartbotic::database::Filter x;
+        x.field = f; x.op = op; x.value = v;
+        return x;
+    };
+    auto sig = [&](const smartbotic::database::Query& q) {
+        auto r = store.scan("c", q);
+        std::string out = "total=" + std::to_string(r.total_matched) +
+                          " more=" + std::to_string(r.has_more ? 1 : 0) + " [";
+        for (const auto& d : r.documents) { out += d.id; out += ","; }
+        return out + "]";
+    };
+    auto Q = [&](smartbotic::database::Filter f, const char* sortField, bool desc,
+                 uint32_t limit, uint32_t offset) {
+        smartbotic::database::Query q;
+        q.filters.push_back(std::move(f));
+        q.sort = smartbotic::database::Sort{sortField, desc};
+        q.limit = limit;
+        q.offset = offset;
+        return q;
+    };
+
+    const std::vector<std::pair<const char*, smartbotic::database::Query>> cases = {
+        {"GT sorted desc on the same field",  Q(F("n", Op::GT, 200), "n", true, 10, 0)},
+        {"GT sorted asc on the same field",   Q(F("n", Op::GT, 200), "n", false, 10, 0)},
+        {"GTE sorted desc",                   Q(F("n", Op::GTE, 200), "n", true, 10, 0)},
+        {"LT sorted asc",                     Q(F("n", Op::LT, 40), "n", false, 10, 0)},
+        {"LTE sorted desc",                   Q(F("n", Op::LTE, 40), "n", true, 10, 0)},
+        {"paged inside the range",            Q(F("n", Op::GT, 100), "n", true, 7, 20)},
+        {"offset past the range",             Q(F("n", Op::GT, 240), "n", true, 10, 99)},
+        {"limit=0 over a range",              Q(F("n", Op::GT, 100), "n", true, 0, 0)},
+        {"range matching nothing",            Q(F("n", Op::GT, 9999), "n", true, 10, 0)},
+        {"range over the whole collection",   Q(F("n", Op::GTE, 0), "n", true, 5, 0)},
+        {"TIED keys, desc",                   Q(F("dup", Op::GT, 40), "dup", true, 12, 0)},
+        {"TIED keys, asc",                    Q(F("dup", Op::GT, 40), "dup", false, 12, 0)},
+        {"TIED keys, desc, paged",            Q(F("dup", Op::GTE, 20), "dup", true, 5, 13)},
+        // A range on one field but sorted by ANOTHER must NOT take this plan.
+        {"range sorted by a different field", Q(F("n", Op::GT, 200), "dup", true, 10, 0)},
+    };
+
+    uint64_t used = 0;
+    for (const auto& [name, q] : cases) {
+        store.set_indexed_fields("c", {});
+        const std::string without = sig(q);
+        store.set_indexed_fields("c", {"n", "dup"});
+        store.build_index("c", "n");
+        store.build_index("c", "dup");
+        store.reset_index_plan_stats();
+        const std::string with = sig(q);
+        used += store.index_plan_stats().range_ordered_scans;
+        if (with != without) {
+            std::cerr << "    scan:     " << without.substr(0, 200) << "\n"
+                      << "    rangeord: " << with.substr(0, 200) << "\n";
+        }
+        const std::string msg = std::string("range-ordered == scan: ") + name;
+        check(with == without, msg.c_str());
+    }
+    // All but the last case (sorted by a different field) should use the plan.
+    check(used == cases.size() - 1,
+          ("the range-ordered plan ran in " + std::to_string(used) + " of " +
+           std::to_string(cases.size() - 1) + " applicable cases").c_str());
+
+    {
+        store.set_indexed_fields("c", {"n", "dup"});
+        store.reset_index_plan_stats();
+        (void)sig(Q(F("n", Op::GT, 200), "dup", true, 10, 0));
+        auto st = store.index_plan_stats();
+        check(st.range_ordered_scans == 0,
+              "a range on one field sorted by another does NOT take the plan - the "
+              "walk's order is not the requested order");
+    }
+}
+
+
+// v2.10.0 — distinct values and min/max, answered from the index.
+void test_index_values() {
+    TmpEnv t("idx-values");
+    LmdbDocumentStore store(t.env);
+    for (int i = 0; i < 120; ++i) {
+        Document d;
+        d.id = "r" + std::to_string(100 + i);
+        d.collection = "c";
+        d.set_data(nlohmann::json{
+            {"state", (i % 3 == 0) ? "done" : ((i % 3 == 1) ? "todo" : "wip")},
+            {"n", i},
+            {"tags", nlohmann::json::array({"t" + std::to_string(i % 2)})},
+        });
+        store.put("c", d.id, d);
+    }
+    store.build_index("c", "state");
+    store.build_index("c", "n");
+    store.build_index("c", "tags");
+
+    auto vals = store.index_values("c", "state", 20, true);
+    check(vals.size() == 3, "three distinct states");
+    check(!vals.empty() && vals[0].value == "done", "ascending: 'done' sorts first");
+    uint64_t sum = 0;
+    for (const auto& v : vals) sum += v.count;
+    check(sum == 120, "the counts add up to every row - none double-counted");
+
+    auto desc = store.index_values("c", "state", 20, false);
+    check(desc.size() == 3 && desc[0].value == "wip",
+          "descending walks from the highest value");
+
+    // limit=1 in each direction is min and max.
+    auto mn = store.index_values("c", "n", 1, true);
+    auto mx = store.index_values("c", "n", 1, false);
+    check(mn.size() == 1 && mn[0].value == 0, "ascending limit=1 is the MINIMUM");
+    check(mx.size() == 1 && mx[0].value == 119, "descending limit=1 is the MAXIMUM");
+
+    check(store.index_values("c", "n", 5, true).size() == 5, "limit is honoured");
+    check(store.index_values("c", "n", 0, true).empty(), "limit=0 returns nothing");
+    check(store.index_values("c", "nosuch", 5, true).empty(),
+          "an unindexed field returns nothing rather than erroring");
+
+    // An array field's keys are ELEMENTS, so reporting the row's value (the whole
+    // array) would be misleading. It reports nothing instead.
+    check(store.index_values("c", "tags", 5, true).empty(),
+          "an array-valued field reports no values - its keys are elements, not "
+          "values, and returning the array would misstate what the key means");
+}
+
 }  // namespace
 
 int main() {
@@ -1436,6 +1740,9 @@ int main() {
     test_range_contains_and_intersection();
     test_index_supplies_ordering();
     test_index_in_and_exists();
+    test_counted_plan();
+    test_range_ordered_plan();
+    test_index_values();
 
     std::cout << "passed: " << g_pass << ", failed: " << g_fail << "\n";
     return g_fail == 0 ? 0 : 1;

この差分においてかなりの量のファイルが変更されているため、一部のファイルを表示していません