Răsfoiți Sursa

feat(vectors): apply metadata filters on vector search

POST /search accepted a `filters` array, documented it in both openapi.json and
llms.txt, and then ignored it: the handler only ever read query_text/
query_vector, top_k and min_score. A filter that matched nothing returned
everything, and a malformed filter raised no error, which is what made it look
like a request-shape problem from the outside.

The database evaluates the filters, exactly as it already does for
/documents: the same parseFilter() grammar and the same eleven operators, so
the two endpoints can never diverge in semantics. find() does not return stored
vectors, so scoring still comes from similaritySearch and the two result sets
are intersected - a document therefore scores identically with and without a
filter, and min_score keeps one meaning.

The ranking is scanned in descending score order and widened until top_k
matches are collected or the collection is exhausted. Stopping at top_k is
exact, not approximate: anything not yet seen scores lower than everything
already kept. The scan is bounded at 10000 candidates.

Malformed specs and unknown operators now return 400 bad_filter instead of
being silently dropped.

Docs realigned to what the code does, after auditing every route and body
parameter against the handlers:
- the operator list was missing in/exists/regex/search, and did not say that
  contains is array membership while search is the substring op
- top_k was documented as required in both files; it defaults to 5
- /vectors still claimed the collection's embedding_model chooses the model,
  which stopped being true when per-project overrides landed in 0.3.0
- GET /docs and POST /ui/logout were undocumented
fszontagh 1 săptămână în urmă
părinte
comite
0a8bb07e36
6 a modificat fișierele cu 235 adăugiri și 10 ștergeri
  1. 1 1
      CLAUDE.md
  2. 19 3
      api/llms.txt
  3. 6 5
      api/openapi.json
  4. 67 1
      src/handlers/vectors.cpp
  5. 129 0
      tests/test_api_integration.cpp
  6. 13 0
      tests/test_embeddings.cpp

+ 1 - 1
CLAUDE.md

@@ -34,7 +34,7 @@ cd webui && npm install && npm run build   # → webui/dist ; `npm run dev` for
 ## Conventions
 
 - C++20, `nlohmann/json`, `spdlog`, cpp-httplib (FetchContent, OpenSSL on). Handlers `throw svapi::ApiError`; a global exception handler maps to `{error:{code,message}}` + HTTP status (`errors.hpp`: 400/401/403/404/422/503/500).
-- Filters: `field:op:value` (`op` ∈ eq/ne/gt/gte/lt/lte/in/contains/exists/regex/search).
+- Filters: `field:op:value` (`op` ∈ eq/ne/gt/gte/lt/lte/in/contains/exists/regex/search). Used by `/documents` (`?filter=`, repeatable) **and** vector `/search` (`filters` body array) - both parse with `parseFilter()` and are evaluated by the DB, so semantics never diverge. `contains` is array membership; `search` is the substring op. Vector search intersects the DB's filter result with the similarity ranking (scores stay identical to an unfiltered query), scanning at most 10000 candidates.
 - **Embedding config precedence:** collection `embedding_model` pin → project override → global default; endpoint + API key are the project override when set, else global. Resolved by the pure `resolveEmbedding()` (`project_settings.hpp`) — unit-test precedence there, not through the handlers.
 - **Every outbound request to an external provider goes through `applyOutboundDefaults()`** (`outbound_http.hpp`): `User-Agent: smartbotic-vectorapi/<version>` and nothing else, TLS verification on, redirects off. Never let cpp-httplib's default UA reach a provider.
 - TS: no `any`, no TS `enum`, no constructor parameter-property shorthand (tsconfig `erasableSyntaxOnly`). Iterating parsed JSON: bind to a named var first (range-for over `nlohmann::json::parse(...)[...]` dangles).

+ 19 - 3
api/llms.txt

@@ -73,7 +73,19 @@ Arbitrary JSON CRUD under a `kind=json` collection.
 - `GET   /api/v1/projects/{project}/collections/{name}/documents` - Find documents.
 - `GET   /api/v1/projects/{project}/collections/{name}/documents/{id}/delete-impact` - Preview what deleting this document would affect, without deleting it: `{would_be_blocked, impacts:[{relation, child_collection, child_field, on_delete, child_count, sample_child_ids, blocks}]}`. Read-only; follows the collection's read grant (a scoped key with a read rule may call it).
 
-**Filter grammar**: `?filter=field:op:value` (repeatable; ANDed). Supported ops: `eq`, `ne`, `lt`, `lte`, `gt`, `gte`, `contains`. Example: `?filter=age:gte:30&filter=name:eq:Alice`.
+**Filter grammar**: `field:op:value`, repeatable and ANDed. On documents it is the `?filter=` query param; on vector search it is the `filters` body array. The same grammar, operators and semantics apply to both, because the database evaluates them.
+
+Supported ops: `eq`, `ne`, `lt`, `lte`, `gt`, `gte`, `in`, `contains`, `exists`, `regex`, `search`.
+
+- `in` - value is a JSON array, e.g. `colour:in:["red","yellow"]`.
+- `contains` - **array membership**, not substring: it matches when the field is an array holding the value. `colour:contains:yellow` does NOT match the string `"yellow"`.
+- `search` - substring match on a string field, e.g. `colour:search:ell` matches `"yellow"`.
+- `regex` - regular-expression match.
+- `exists` - `field:exists:true` / `field:exists:false`.
+
+The value is parsed as JSON when it parses (numbers, booleans, arrays), otherwise it is taken as a string. Everything after the second `:` is the value, so a value may itself contain `:`. A malformed spec or unknown operator is rejected with **400** `bad_filter`.
+
+Example: `?filter=age:gte:30&filter=name:eq:Alice`.
 
 Additional query params: `limit` (default 20), `offset` (default 0), `sort` (field name), `desc` (bool).
 
@@ -100,8 +112,10 @@ Relation names are per-project and **unqualified**: requests and responses alway
 
 Requires a `kind=vector` collection. Vectors are stored with cosine-similarity indexing.
 
-- `POST /api/v1/projects/{project}/collections/{name}/vectors` `{id?, text?, vector?, metadata?}` - Store a vector. Supply either `text` (the server embeds it using the collection's `embedding_model`) or a pre-computed `vector` array. `vector` dimension must match the collection's `vector_dimension`. `metadata` is an arbitrary JSON object stored alongside the vector.
-- `POST /api/v1/projects/{project}/collections/{name}/search` `{query_text?, query_vector?, top_k, min_score?, filters?}` - Cosine-similarity search. Supply either `query_text` or `query_vector`. `top_k` (required) limits results. `min_score` (0-1) filters low-confidence matches. `filters` apply the same `field:op:value` grammar to vector metadata.
+- `POST /api/v1/projects/{project}/collections/{name}/vectors` `{id?, text?, vector?, metadata?}` - Store a vector. Supply either `text` (the server embeds it, resolving the model as: the collection's `embedding_model` pin, else the project's embedding override, else the global default) or a pre-computed `vector` array. `vector` dimension must match the collection's `vector_dimension`. `metadata` is an arbitrary JSON object stored alongside the vector.
+- `POST /api/v1/projects/{project}/collections/{name}/search` `{query_text?, query_vector?, top_k?, min_score?, filters?}` - Cosine-similarity search. Supply either `query_text` or `query_vector`; neither is **422**. `top_k` caps the results, **default 5**. `min_score` (0-1) drops low-confidence matches, default 0. `filters` is an array of `field:op:value` strings applied to the vector's metadata, using the grammar and operators above - metadata is stored at the top level of the document, so a vector written with `metadata:{"colour":"red"}` is filtered as `colour:eq:red`.
+
+  Filtering is evaluated by the database, then intersected with the similarity ranking, so scores are identical with and without a filter and `min_score` means the same thing either way. The scan is bounded at 10000 candidates: within a collection of that size a filtered search returns exactly the true top-k, and beyond it a filter matching only very low-ranked vectors may return fewer.
 
   **Latency tip:** `query_text` triggers a synchronous server-side embedding call to the configured provider, which usually dominates request time (seconds for large models). Cosine search itself is sub-millisecond. If your client issues many searches (e.g. an LLM/n8n agent doing RAG), embed the query once on your side and pass `query_vector` to skip server-side embedding entirely.
 
@@ -142,8 +156,10 @@ Public (no auth required):
 - `GET /readyz`  - Readiness probe; 200 when the database is reachable, 503 otherwise.
 - `GET /openapi.json` - OpenAPI 3.1 machine-readable spec.
 - `GET /llms.txt`     - This file.
+- `GET /docs`         - Rendered API reference page (RapiDoc over the spec above).
 
 Session login (used by the web UI):
 
 - `POST /ui/login` `{key}` - Exchange an API key for an HttpOnly session cookie (`svapi_session`). The cookie may be used in place of the `Authorization: Bearer` header for all `/api/v1/*` endpoints.
+- `POST /ui/logout` - Invalidate the session and clear the cookie.
 - `GET  /ui/session` - Return `{ok, admin, projects}` for the caller's current session (cookie or bearer), or **401** when there is none. Sessions are held in memory, so a server restart invalidates every cookie; a client must call this before trusting any locally stored "logged in" state.

+ 6 - 5
api/openapi.json

@@ -1212,19 +1212,20 @@
                   "top_k": {
                     "type": "integer",
                     "minimum": 1,
-                    "description": "Maximum number of results to return"
+                    "default": 5,
+                    "description": "Maximum number of results to return (default 5)"
                   },
                   "min_score": {
                     "type": "number",
-                    "description": "Minimum cosine similarity score (0-1)"
+                    "default": 0,
+                    "description": "Minimum cosine similarity score (0-1), default 0"
                   },
                   "filters": {
                     "type": "array",
                     "items": { "type": "string" },
-                    "description": "Metadata filters in field:op:value grammar (e.g. src:eq:n8n)"
+                    "description": "Metadata filters in field:op:value grammar, ANDed (e.g. src:eq:n8n). Same operators as the documents endpoint (eq, ne, lt, lte, gt, gte, in, contains, exists, regex, search) and evaluated by the database, so a document scores the same with or without a filter. Metadata is stored at the document top level: metadata {\"colour\":\"red\"} is filtered as colour:eq:red. Note that contains is array membership - use search for substrings. A malformed spec or unknown operator returns 400 bad_filter. The filtered scan is bounded at 10000 candidates."
                   }
-                },
-                "required": ["top_k"]
+                }
               }
             }
           }

+ 67 - 1
src/handlers/vectors.cpp

@@ -2,11 +2,36 @@
 #include "embedding_cache.hpp"
 #include "embeddings.hpp"
 #include "errors.hpp"
+#include "filters.hpp"
 #include "project_settings.hpp"
 #include "json_http.hpp"
 #include "server.hpp"
+#include <algorithm>
+#include <unordered_set>
 namespace svapi {
 namespace {
+
+// Upper bound on how deep a filtered search will scan the similarity ranking,
+// and on how many filter-matching ids we fetch. A filtered search stays exact
+// while the collection is within this bound.
+constexpr uint32_t kFilterScanCap = 10000;
+
+// Parse the optional `filters` body field into DB filters. Same grammar and
+// operator set as the documents endpoint, because the DB evaluates them.
+std::vector<smartbotic::database::Client::Filter> parseSearchFilters(const nlohmann::json& body) {
+    std::vector<smartbotic::database::Client::Filter> out;
+    if (!body.contains("filters") || body["filters"].is_null()) return out;
+    if (!body["filters"].is_array())
+        throw ApiError(ErrCode::Unprocessable, "validation",
+                       "filters must be an array of \"field:op:value\" strings");
+    for (const auto& f : body["filters"]) {
+        if (!f.is_string())
+            throw ApiError(ErrCode::Unprocessable, "validation",
+                           "each filter must be a \"field:op:value\" string");
+        out.push_back(parseFilter(f.get<std::string>()));
+    }
+    return out;
+}
 CollectionMeta requireVectorCollection(ServerDeps* d, const httplib::Request& req,
                                         std::string& projectOut, std::string& nameOut, KeyOp op) {
     ApiKey k = requireKey(*d, req);
@@ -82,7 +107,48 @@ void registerVectorRoutes(ApiServer& s) {
         std::vector<float> q = resolveVector(d, project, meta, body, "query_vector", "query_text", wantsNoStore(req));
         uint32_t topK = body.value("top_k", 5u);
         float minScore = body.value("min_score", 0.0f);
-        auto results = d->db.client().similaritySearch(qualify(project, name), q, topK, minScore);
+        auto filters = parseSearchFilters(body);
+        const std::string coll = qualify(project, name);
+
+        std::vector<smartbotic::database::Client::SimilarityResult> results;
+        if (filters.empty()) {
+            results = d->db.client().similaritySearch(coll, q, topK, minScore);
+        } else {
+            // The DB evaluates the filters (same engine as /documents), so every
+            // operator behaves identically on both endpoints. It will not return
+            // the stored vectors, so scoring still has to come from
+            // similaritySearch - we intersect the two.
+            smartbotic::database::Client::QueryOptions opts;
+            opts.filters = filters;
+            opts.limit   = kFilterScanCap;
+            std::unordered_set<std::string> allowed;
+            for (const auto& doc : d->db.client().find(coll, opts)) {
+                const std::string id = doc.value("_id", "");
+                if (!id.empty()) allowed.insert(id);
+            }
+            if (!allowed.empty()) {
+                // Walk the ranking from the top, widening until topK matches are
+                // found or the collection is exhausted. Because the scan is in
+                // descending score order, stopping as soon as topK matches are
+                // collected still yields the true topK: anything unseen scores
+                // lower than everything already kept.
+                uint32_t scan = std::max<uint32_t>(topK * 4, 64);
+                while (true) {
+                    auto page = d->db.client().similaritySearch(coll, q, scan, minScore);
+                    results.clear();
+                    for (const auto& r : page) {
+                        if (!allowed.count(r.id)) continue;
+                        results.push_back(r);
+                        if (results.size() >= topK) break;
+                    }
+                    if (results.size() >= topK) break;      // enough matches
+                    if (page.size() < scan) break;          // ranking exhausted
+                    if (scan >= kFilterScanCap) break;      // bounded work
+                    scan = std::min(scan * 4, kFilterScanCap);
+                }
+            }
+        }
+
         nlohmann::json arr = nlohmann::json::array();
         for (const auto& r : results) arr.push_back({{"id", r.id}, {"score", r.score}, {"data", r.data}});
         sendJson(res, 200, {{"results", arr}});

+ 129 - 0
tests/test_api_integration.cpp

@@ -1477,3 +1477,132 @@ TEST_F(ApiFixture, SessionEndpointAcceptsABearerKeyToo) {
     ASSERT_TRUE(r); ASSERT_EQ(r->status, 200);
     EXPECT_TRUE(nlohmann::json::parse(r->body)["admin"].get<bool>());
 }
+
+// --- metadata filters on vector search ----------------------------------------
+
+namespace {
+// Two orthogonal-ish unit vectors so the ranking is unambiguous, plus a query
+// that is closest to "red".
+const nlohmann::json RED_VEC    = {1.0, 0.0, 0.0, 0.0};
+const nlohmann::json YELLOW_VEC = {0.0, 1.0, 0.0, 0.0};
+const nlohmann::json QUERY_VEC  = {0.9, 0.4, 0.0, 0.0};
+
+std::string seedColourCollection(ApiFixture& f, httplib::Client& c, const std::string& project,
+                                 const std::string& coll, int dim) {
+    std::string base = "/api/v1/projects/" + project + "/collections";
+    EXPECT_EQ(c.Post(base.c_str(),
+        nlohmann::json{{"name", coll}, {"kind", "vector"}, {"vector_dimension", dim}}.dump(),
+        "application/json")->status, 201);
+    EXPECT_EQ(c.Post((base + "/" + coll + "/vectors").c_str(),
+        nlohmann::json{{"id","a"},{"vector",RED_VEC},{"metadata",{{"colour","red"},{"n",1}}}}.dump(),
+        "application/json")->status, 201);
+    EXPECT_EQ(c.Post((base + "/" + coll + "/vectors").c_str(),
+        nlohmann::json{{"id","b"},{"vector",YELLOW_VEC},{"metadata",{{"colour","yellow"},{"n",2}}}}.dump(),
+        "application/json")->status, 201);
+    return base + "/" + coll + "/search";
+}
+}  // namespace
+
+TEST_F(ApiFixture, VectorSearchFiltersRestrictResults) {
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "colours", 4);
+
+    auto run = [&](nlohmann::json filters) {
+        nlohmann::json body{{"query_vector", QUERY_VEC}, {"top_k", 10}};
+        if (!filters.is_null()) body["filters"] = filters;
+        auto r = c.Post(search.c_str(), body.dump(), "application/json");
+        EXPECT_EQ(r->status, 200);
+        return nlohmann::json::parse(r->body)["results"];
+    };
+
+    EXPECT_EQ(run(nullptr).size(), 2u) << "no filter: both vectors";
+    auto red = run(nlohmann::json::array({"colour:eq:red"}));
+    ASSERT_EQ(red.size(), 1u) << "colour:eq:red must return only the red vector";
+    EXPECT_EQ(red[0]["id"], "a");
+    EXPECT_EQ(run(nlohmann::json::array({"colour:eq:blue"})).size(), 0u) << "no match must return nothing";
+}
+
+TEST_F(ApiFixture, VectorSearchFiltersSupportTheDocumentOperators) {
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "ops", 4);
+    auto count = [&](const std::string& spec) {
+        auto r = c.Post(search.c_str(),
+            nlohmann::json{{"query_vector", QUERY_VEC}, {"top_k", 10},
+                           {"filters", nlohmann::json::array({spec})}}.dump(), "application/json");
+        EXPECT_EQ(r->status, 200) << spec << " -> " << r->body;
+        return nlohmann::json::parse(r->body)["results"].size();
+    };
+    EXPECT_EQ(count("colour:ne:red"), 1u);
+    EXPECT_EQ(count("n:gt:1"), 1u);
+    EXPECT_EQ(count("n:gte:1"), 2u);
+    EXPECT_EQ(count("n:lt:2"), 1u);
+    EXPECT_EQ(count("colour:in:[\"red\",\"yellow\"]"), 2u);
+    // `search` is the substring operator; `contains` is array-membership, and
+    // both behave exactly as they do on /documents because the DB evaluates them.
+    EXPECT_EQ(count("colour:search:ell"), 1u);
+    EXPECT_EQ(count("colour:contains:yellow"), 0u);
+    EXPECT_EQ(count("colour:exists:true"), 2u);
+    EXPECT_EQ(count("missing_field:exists:true"), 0u);
+}
+
+TEST_F(ApiFixture, VectorSearchFiltersCombineWithTopKAndMinScore) {
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "combo", 4);
+    // top_k still caps a filtered result set.
+    auto r = c.Post(search.c_str(),
+        nlohmann::json{{"query_vector", QUERY_VEC}, {"top_k", 1},
+                       {"filters", nlohmann::json::array({"colour:exists:true"})}}.dump(),
+        "application/json");
+    ASSERT_EQ(r->status, 200);
+    auto res = nlohmann::json::parse(r->body)["results"];
+    ASSERT_EQ(res.size(), 1u);
+    EXPECT_EQ(res[0]["id"], "a") << "the nearer vector must win, not simply the first match";
+
+    // min_score still applies on top of the filter.
+    auto r2 = c.Post(search.c_str(),
+        nlohmann::json{{"query_vector", QUERY_VEC}, {"top_k", 10}, {"min_score", 0.99},
+                       {"filters", nlohmann::json::array({"colour:exists:true"})}}.dump(),
+        "application/json");
+    ASSERT_EQ(r2->status, 200);
+    EXPECT_EQ(nlohmann::json::parse(r2->body)["results"].size(), 0u);
+}
+
+TEST_F(ApiFixture, VectorSearchRejectsMalformedFilters) {
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "bad", 4);
+    // A filter the server cannot parse must fail loudly rather than be ignored.
+    for (const char* spec : {"colour", "colour:nosuchop:red"}) {
+        auto r = c.Post(search.c_str(),
+            nlohmann::json{{"query_vector", QUERY_VEC}, {"top_k", 5},
+                           {"filters", nlohmann::json::array({spec})}}.dump(), "application/json");
+        EXPECT_GE(r->status, 400) << spec << " should be rejected";
+        EXPECT_LT(r->status, 500) << spec << " should be a client error";
+    }
+    // Wrong type for the field entirely.
+    auto r = c.Post(search.c_str(),
+        nlohmann::json{{"query_vector", QUERY_VEC}, {"filters", "colour:eq:red"}}.dump(),
+        "application/json");
+    EXPECT_GE(r->status, 400);
+}
+
+TEST_F(ApiFixture, VectorSearchFilteredScoresMatchUnfilteredScores) {
+    // A filtered search must report the same score for a document as an
+    // unfiltered one, otherwise min_score means two different things.
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "scores", 4);
+    auto scoreOf = [&](nlohmann::json body, const std::string& id) -> double {
+        auto r = c.Post(search.c_str(), body.dump(), "application/json");
+        EXPECT_EQ(r->status, 200);
+        // Bind before iterating: a range-for over json::parse(...)[...] walks a
+        // destroyed temporary (see CLAUDE.md).
+        const nlohmann::json parsed = nlohmann::json::parse(r->body);
+        for (const auto& h : parsed["results"])
+            if (h["id"] == id) return h["score"].get<double>();
+        return -1.0;
+    };
+    double plain    = scoreOf({{"query_vector", QUERY_VEC}, {"top_k", 10}}, "a");
+    double filtered = scoreOf({{"query_vector", QUERY_VEC}, {"top_k", 10},
+                               {"filters", nlohmann::json::array({"colour:eq:red"})}}, "a");
+    ASSERT_GT(plain, 0.0);
+    EXPECT_NEAR(plain, filtered, 1e-6);
+}

+ 13 - 0
tests/test_embeddings.cpp

@@ -293,3 +293,16 @@ TEST(EmbeddingCache, ReconfigureEvictsToTighterBounds) {
     EXPECT_FALSE(cache.get(kA).has_value());  // evicted
     EXPECT_TRUE(cache.get(kC).has_value());   // most-recently used survives
 }
+
+// A base URL that already ends in /v1 is the common OpenRouter mistake: the
+// client appends "/v1/embeddings" to whatever path the base carries, so the
+// request lands on /api/v1/v1/embeddings, which OpenRouter answers with 404.
+TEST(Embeddings, BaseEndingInV1DuplicatesTheVersionSegment) {
+    auto wrong = splitEmbeddingBase("https://openrouter.ai/api/v1");
+    EXPECT_EQ(wrong.origin, "https://openrouter.ai");
+    EXPECT_EQ(wrong.path, "/api/v1");
+    EXPECT_EQ(wrong.path + "/v1/embeddings", "/api/v1/v1/embeddings");
+
+    auto right = splitEmbeddingBase("https://openrouter.ai/api");
+    EXPECT_EQ(right.path + "/v1/embeddings", "/api/v1/embeddings");
+}