Ver código fonte

Merge: vector search metadata filters + doc alignment

fszontagh 1 semana atrás
pai
commit
77dad03590
6 arquivos alterados com 235 adições e 10 exclusões
  1. 1 1
      CLAUDE.md
  2. 19 3
      api/llms.txt
  3. 6 5
      api/openapi.json
  4. 67 1
      src/handlers/vectors.cpp
  5. 129 0
      tests/test_api_integration.cpp
  6. 13 0
      tests/test_embeddings.cpp

+ 1 - 1
CLAUDE.md

@@ -34,7 +34,7 @@ cd webui && npm install && npm run build   # → webui/dist ; `npm run dev` for
 ## Conventions
 
 - C++20, `nlohmann/json`, `spdlog`, cpp-httplib (FetchContent, OpenSSL on). Handlers `throw svapi::ApiError`; a global exception handler maps to `{error:{code,message}}` + HTTP status (`errors.hpp`: 400/401/403/404/422/503/500).
-- Filters: `field:op:value` (`op` ∈ eq/ne/gt/gte/lt/lte/in/contains/exists/regex/search).
+- Filters: `field:op:value` (`op` ∈ eq/ne/gt/gte/lt/lte/in/contains/exists/regex/search). Used by `/documents` (`?filter=`, repeatable) **and** vector `/search` (`filters` body array) - both parse with `parseFilter()` and are evaluated by the DB, so semantics never diverge. `contains` is array membership; `search` is the substring op. Vector search intersects the DB's filter result with the similarity ranking (scores stay identical to an unfiltered query), scanning at most 10000 candidates.
 - **Embedding config precedence:** collection `embedding_model` pin → project override → global default; endpoint + API key are the project override when set, else global. Resolved by the pure `resolveEmbedding()` (`project_settings.hpp`) — unit-test precedence there, not through the handlers.
 - **Every outbound request to an external provider goes through `applyOutboundDefaults()`** (`outbound_http.hpp`): `User-Agent: smartbotic-vectorapi/<version>` and nothing else, TLS verification on, redirects off. Never let cpp-httplib's default UA reach a provider.
 - TS: no `any`, no TS `enum`, no constructor parameter-property shorthand (tsconfig `erasableSyntaxOnly`). Iterating parsed JSON: bind to a named var first (range-for over `nlohmann::json::parse(...)[...]` dangles).

+ 19 - 3
api/llms.txt

@@ -73,7 +73,19 @@ Arbitrary JSON CRUD under a `kind=json` collection.
 - `GET   /api/v1/projects/{project}/collections/{name}/documents` - Find documents.
 - `GET   /api/v1/projects/{project}/collections/{name}/documents/{id}/delete-impact` - Preview what deleting this document would affect, without deleting it: `{would_be_blocked, impacts:[{relation, child_collection, child_field, on_delete, child_count, sample_child_ids, blocks}]}`. Read-only; follows the collection's read grant (a scoped key with a read rule may call it).
 
-**Filter grammar**: `?filter=field:op:value` (repeatable; ANDed). Supported ops: `eq`, `ne`, `lt`, `lte`, `gt`, `gte`, `contains`. Example: `?filter=age:gte:30&filter=name:eq:Alice`.
+**Filter grammar**: `field:op:value`, repeatable and ANDed. On documents it is the `?filter=` query param; on vector search it is the `filters` body array. The same grammar, operators and semantics apply to both, because the database evaluates them.
+
+Supported ops: `eq`, `ne`, `lt`, `lte`, `gt`, `gte`, `in`, `contains`, `exists`, `regex`, `search`.
+
+- `in` - value is a JSON array, e.g. `colour:in:["red","yellow"]`.
+- `contains` - **array membership**, not substring: it matches when the field is an array holding the value. `colour:contains:yellow` does NOT match the string `"yellow"`.
+- `search` - substring match on a string field, e.g. `colour:search:ell` matches `"yellow"`.
+- `regex` - regular-expression match.
+- `exists` - `field:exists:true` / `field:exists:false`.
+
+The value is parsed as JSON when it parses (numbers, booleans, arrays), otherwise it is taken as a string. Everything after the second `:` is the value, so a value may itself contain `:`. A malformed spec or unknown operator is rejected with **400** `bad_filter`.
+
+Example: `?filter=age:gte:30&filter=name:eq:Alice`.
 
 Additional query params: `limit` (default 20), `offset` (default 0), `sort` (field name), `desc` (bool).
 
@@ -100,8 +112,10 @@ Relation names are per-project and **unqualified**: requests and responses alway
 
 Requires a `kind=vector` collection. Vectors are stored with cosine-similarity indexing.
 
-- `POST /api/v1/projects/{project}/collections/{name}/vectors` `{id?, text?, vector?, metadata?}` - Store a vector. Supply either `text` (the server embeds it using the collection's `embedding_model`) or a pre-computed `vector` array. `vector` dimension must match the collection's `vector_dimension`. `metadata` is an arbitrary JSON object stored alongside the vector.
-- `POST /api/v1/projects/{project}/collections/{name}/search` `{query_text?, query_vector?, top_k, min_score?, filters?}` - Cosine-similarity search. Supply either `query_text` or `query_vector`. `top_k` (required) limits results. `min_score` (0-1) filters low-confidence matches. `filters` apply the same `field:op:value` grammar to vector metadata.
+- `POST /api/v1/projects/{project}/collections/{name}/vectors` `{id?, text?, vector?, metadata?}` - Store a vector. Supply either `text` (the server embeds it, resolving the model as: the collection's `embedding_model` pin, else the project's embedding override, else the global default) or a pre-computed `vector` array. `vector` dimension must match the collection's `vector_dimension`. `metadata` is an arbitrary JSON object stored alongside the vector.
+- `POST /api/v1/projects/{project}/collections/{name}/search` `{query_text?, query_vector?, top_k?, min_score?, filters?}` - Cosine-similarity search. Supply either `query_text` or `query_vector`; neither is **422**. `top_k` caps the results, **default 5**. `min_score` (0-1) drops low-confidence matches, default 0. `filters` is an array of `field:op:value` strings applied to the vector's metadata, using the grammar and operators above - metadata is stored at the top level of the document, so a vector written with `metadata:{"colour":"red"}` is filtered as `colour:eq:red`.
+
+  Filtering is evaluated by the database, then intersected with the similarity ranking, so scores are identical with and without a filter and `min_score` means the same thing either way. The scan is bounded at 10000 candidates: within a collection of that size a filtered search returns exactly the true top-k, and beyond it a filter matching only very low-ranked vectors may return fewer.
 
   **Latency tip:** `query_text` triggers a synchronous server-side embedding call to the configured provider, which usually dominates request time (seconds for large models). Cosine search itself is sub-millisecond. If your client issues many searches (e.g. an LLM/n8n agent doing RAG), embed the query once on your side and pass `query_vector` to skip server-side embedding entirely.
 
@@ -142,8 +156,10 @@ Public (no auth required):
 - `GET /readyz`  - Readiness probe; 200 when the database is reachable, 503 otherwise.
 - `GET /openapi.json` - OpenAPI 3.1 machine-readable spec.
 - `GET /llms.txt`     - This file.
+- `GET /docs`         - Rendered API reference page (RapiDoc over the spec above).
 
 Session login (used by the web UI):
 
 - `POST /ui/login` `{key}` - Exchange an API key for an HttpOnly session cookie (`svapi_session`). The cookie may be used in place of the `Authorization: Bearer` header for all `/api/v1/*` endpoints.
+- `POST /ui/logout` - Invalidate the session and clear the cookie.
 - `GET  /ui/session` - Return `{ok, admin, projects}` for the caller's current session (cookie or bearer), or **401** when there is none. Sessions are held in memory, so a server restart invalidates every cookie; a client must call this before trusting any locally stored "logged in" state.

+ 6 - 5
api/openapi.json

@@ -1212,19 +1212,20 @@
                   "top_k": {
                     "type": "integer",
                     "minimum": 1,
-                    "description": "Maximum number of results to return"
+                    "default": 5,
+                    "description": "Maximum number of results to return (default 5)"
                   },
                   "min_score": {
                     "type": "number",
-                    "description": "Minimum cosine similarity score (0-1)"
+                    "default": 0,
+                    "description": "Minimum cosine similarity score (0-1), default 0"
                   },
                   "filters": {
                     "type": "array",
                     "items": { "type": "string" },
-                    "description": "Metadata filters in field:op:value grammar (e.g. src:eq:n8n)"
+                    "description": "Metadata filters in field:op:value grammar, ANDed (e.g. src:eq:n8n). Same operators as the documents endpoint (eq, ne, lt, lte, gt, gte, in, contains, exists, regex, search) and evaluated by the database, so a document scores the same with or without a filter. Metadata is stored at the document top level: metadata {\"colour\":\"red\"} is filtered as colour:eq:red. Note that contains is array membership - use search for substrings. A malformed spec or unknown operator returns 400 bad_filter. The filtered scan is bounded at 10000 candidates."
                   }
-                },
-                "required": ["top_k"]
+                }
               }
             }
           }

+ 67 - 1
src/handlers/vectors.cpp

@@ -2,11 +2,36 @@
 #include "embedding_cache.hpp"
 #include "embeddings.hpp"
 #include "errors.hpp"
+#include "filters.hpp"
 #include "project_settings.hpp"
 #include "json_http.hpp"
 #include "server.hpp"
+#include <algorithm>
+#include <unordered_set>
 namespace svapi {
 namespace {
+
+// Upper bound on how deep a filtered search will scan the similarity ranking,
+// and on how many filter-matching ids we fetch. A filtered search stays exact
+// while the collection is within this bound.
+constexpr uint32_t kFilterScanCap = 10000;
+
+// Parse the optional `filters` body field into DB filters. Same grammar and
+// operator set as the documents endpoint, because the DB evaluates them.
+std::vector<smartbotic::database::Client::Filter> parseSearchFilters(const nlohmann::json& body) {
+    std::vector<smartbotic::database::Client::Filter> out;
+    if (!body.contains("filters") || body["filters"].is_null()) return out;
+    if (!body["filters"].is_array())
+        throw ApiError(ErrCode::Unprocessable, "validation",
+                       "filters must be an array of \"field:op:value\" strings");
+    for (const auto& f : body["filters"]) {
+        if (!f.is_string())
+            throw ApiError(ErrCode::Unprocessable, "validation",
+                           "each filter must be a \"field:op:value\" string");
+        out.push_back(parseFilter(f.get<std::string>()));
+    }
+    return out;
+}
 CollectionMeta requireVectorCollection(ServerDeps* d, const httplib::Request& req,
                                         std::string& projectOut, std::string& nameOut, KeyOp op) {
     ApiKey k = requireKey(*d, req);
@@ -82,7 +107,48 @@ void registerVectorRoutes(ApiServer& s) {
         std::vector<float> q = resolveVector(d, project, meta, body, "query_vector", "query_text", wantsNoStore(req));
         uint32_t topK = body.value("top_k", 5u);
         float minScore = body.value("min_score", 0.0f);
-        auto results = d->db.client().similaritySearch(qualify(project, name), q, topK, minScore);
+        auto filters = parseSearchFilters(body);
+        const std::string coll = qualify(project, name);
+
+        std::vector<smartbotic::database::Client::SimilarityResult> results;
+        if (filters.empty()) {
+            results = d->db.client().similaritySearch(coll, q, topK, minScore);
+        } else {
+            // The DB evaluates the filters (same engine as /documents), so every
+            // operator behaves identically on both endpoints. It will not return
+            // the stored vectors, so scoring still has to come from
+            // similaritySearch - we intersect the two.
+            smartbotic::database::Client::QueryOptions opts;
+            opts.filters = filters;
+            opts.limit   = kFilterScanCap;
+            std::unordered_set<std::string> allowed;
+            for (const auto& doc : d->db.client().find(coll, opts)) {
+                const std::string id = doc.value("_id", "");
+                if (!id.empty()) allowed.insert(id);
+            }
+            if (!allowed.empty()) {
+                // Walk the ranking from the top, widening until topK matches are
+                // found or the collection is exhausted. Because the scan is in
+                // descending score order, stopping as soon as topK matches are
+                // collected still yields the true topK: anything unseen scores
+                // lower than everything already kept.
+                uint32_t scan = std::max<uint32_t>(topK * 4, 64);
+                while (true) {
+                    auto page = d->db.client().similaritySearch(coll, q, scan, minScore);
+                    results.clear();
+                    for (const auto& r : page) {
+                        if (!allowed.count(r.id)) continue;
+                        results.push_back(r);
+                        if (results.size() >= topK) break;
+                    }
+                    if (results.size() >= topK) break;      // enough matches
+                    if (page.size() < scan) break;          // ranking exhausted
+                    if (scan >= kFilterScanCap) break;      // bounded work
+                    scan = std::min(scan * 4, kFilterScanCap);
+                }
+            }
+        }
+
         nlohmann::json arr = nlohmann::json::array();
         for (const auto& r : results) arr.push_back({{"id", r.id}, {"score", r.score}, {"data", r.data}});
         sendJson(res, 200, {{"results", arr}});

+ 129 - 0
tests/test_api_integration.cpp

@@ -1477,3 +1477,132 @@ TEST_F(ApiFixture, SessionEndpointAcceptsABearerKeyToo) {
     ASSERT_TRUE(r); ASSERT_EQ(r->status, 200);
     EXPECT_TRUE(nlohmann::json::parse(r->body)["admin"].get<bool>());
 }
+
+// --- metadata filters on vector search ----------------------------------------
+
+namespace {
+// Two orthogonal-ish unit vectors so the ranking is unambiguous, plus a query
+// that is closest to "red".
+const nlohmann::json RED_VEC    = {1.0, 0.0, 0.0, 0.0};
+const nlohmann::json YELLOW_VEC = {0.0, 1.0, 0.0, 0.0};
+const nlohmann::json QUERY_VEC  = {0.9, 0.4, 0.0, 0.0};
+
+std::string seedColourCollection(ApiFixture& f, httplib::Client& c, const std::string& project,
+                                 const std::string& coll, int dim) {
+    std::string base = "/api/v1/projects/" + project + "/collections";
+    EXPECT_EQ(c.Post(base.c_str(),
+        nlohmann::json{{"name", coll}, {"kind", "vector"}, {"vector_dimension", dim}}.dump(),
+        "application/json")->status, 201);
+    EXPECT_EQ(c.Post((base + "/" + coll + "/vectors").c_str(),
+        nlohmann::json{{"id","a"},{"vector",RED_VEC},{"metadata",{{"colour","red"},{"n",1}}}}.dump(),
+        "application/json")->status, 201);
+    EXPECT_EQ(c.Post((base + "/" + coll + "/vectors").c_str(),
+        nlohmann::json{{"id","b"},{"vector",YELLOW_VEC},{"metadata",{{"colour","yellow"},{"n",2}}}}.dump(),
+        "application/json")->status, 201);
+    return base + "/" + coll + "/search";
+}
+}  // namespace
+
+TEST_F(ApiFixture, VectorSearchFiltersRestrictResults) {
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "colours", 4);
+
+    auto run = [&](nlohmann::json filters) {
+        nlohmann::json body{{"query_vector", QUERY_VEC}, {"top_k", 10}};
+        if (!filters.is_null()) body["filters"] = filters;
+        auto r = c.Post(search.c_str(), body.dump(), "application/json");
+        EXPECT_EQ(r->status, 200);
+        return nlohmann::json::parse(r->body)["results"];
+    };
+
+    EXPECT_EQ(run(nullptr).size(), 2u) << "no filter: both vectors";
+    auto red = run(nlohmann::json::array({"colour:eq:red"}));
+    ASSERT_EQ(red.size(), 1u) << "colour:eq:red must return only the red vector";
+    EXPECT_EQ(red[0]["id"], "a");
+    EXPECT_EQ(run(nlohmann::json::array({"colour:eq:blue"})).size(), 0u) << "no match must return nothing";
+}
+
+TEST_F(ApiFixture, VectorSearchFiltersSupportTheDocumentOperators) {
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "ops", 4);
+    auto count = [&](const std::string& spec) {
+        auto r = c.Post(search.c_str(),
+            nlohmann::json{{"query_vector", QUERY_VEC}, {"top_k", 10},
+                           {"filters", nlohmann::json::array({spec})}}.dump(), "application/json");
+        EXPECT_EQ(r->status, 200) << spec << " -> " << r->body;
+        return nlohmann::json::parse(r->body)["results"].size();
+    };
+    EXPECT_EQ(count("colour:ne:red"), 1u);
+    EXPECT_EQ(count("n:gt:1"), 1u);
+    EXPECT_EQ(count("n:gte:1"), 2u);
+    EXPECT_EQ(count("n:lt:2"), 1u);
+    EXPECT_EQ(count("colour:in:[\"red\",\"yellow\"]"), 2u);
+    // `search` is the substring operator; `contains` is array-membership, and
+    // both behave exactly as they do on /documents because the DB evaluates them.
+    EXPECT_EQ(count("colour:search:ell"), 1u);
+    EXPECT_EQ(count("colour:contains:yellow"), 0u);
+    EXPECT_EQ(count("colour:exists:true"), 2u);
+    EXPECT_EQ(count("missing_field:exists:true"), 0u);
+}
+
+TEST_F(ApiFixture, VectorSearchFiltersCombineWithTopKAndMinScore) {
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "combo", 4);
+    // top_k still caps a filtered result set.
+    auto r = c.Post(search.c_str(),
+        nlohmann::json{{"query_vector", QUERY_VEC}, {"top_k", 1},
+                       {"filters", nlohmann::json::array({"colour:exists:true"})}}.dump(),
+        "application/json");
+    ASSERT_EQ(r->status, 200);
+    auto res = nlohmann::json::parse(r->body)["results"];
+    ASSERT_EQ(res.size(), 1u);
+    EXPECT_EQ(res[0]["id"], "a") << "the nearer vector must win, not simply the first match";
+
+    // min_score still applies on top of the filter.
+    auto r2 = c.Post(search.c_str(),
+        nlohmann::json{{"query_vector", QUERY_VEC}, {"top_k", 10}, {"min_score", 0.99},
+                       {"filters", nlohmann::json::array({"colour:exists:true"})}}.dump(),
+        "application/json");
+    ASSERT_EQ(r2->status, 200);
+    EXPECT_EQ(nlohmann::json::parse(r2->body)["results"].size(), 0u);
+}
+
+TEST_F(ApiFixture, VectorSearchRejectsMalformedFilters) {
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "bad", 4);
+    // A filter the server cannot parse must fail loudly rather than be ignored.
+    for (const char* spec : {"colour", "colour:nosuchop:red"}) {
+        auto r = c.Post(search.c_str(),
+            nlohmann::json{{"query_vector", QUERY_VEC}, {"top_k", 5},
+                           {"filters", nlohmann::json::array({spec})}}.dump(), "application/json");
+        EXPECT_GE(r->status, 400) << spec << " should be rejected";
+        EXPECT_LT(r->status, 500) << spec << " should be a client error";
+    }
+    // Wrong type for the field entirely.
+    auto r = c.Post(search.c_str(),
+        nlohmann::json{{"query_vector", QUERY_VEC}, {"filters", "colour:eq:red"}}.dump(),
+        "application/json");
+    EXPECT_GE(r->status, 400);
+}
+
+TEST_F(ApiFixture, VectorSearchFilteredScoresMatchUnfilteredScores) {
+    // A filtered search must report the same score for a document as an
+    // unfiltered one, otherwise min_score means two different things.
+    auto c = admin();
+    std::string search = seedColourCollection(*this, c, project_, "scores", 4);
+    auto scoreOf = [&](nlohmann::json body, const std::string& id) -> double {
+        auto r = c.Post(search.c_str(), body.dump(), "application/json");
+        EXPECT_EQ(r->status, 200);
+        // Bind before iterating: a range-for over json::parse(...)[...] walks a
+        // destroyed temporary (see CLAUDE.md).
+        const nlohmann::json parsed = nlohmann::json::parse(r->body);
+        for (const auto& h : parsed["results"])
+            if (h["id"] == id) return h["score"].get<double>();
+        return -1.0;
+    };
+    double plain    = scoreOf({{"query_vector", QUERY_VEC}, {"top_k", 10}}, "a");
+    double filtered = scoreOf({{"query_vector", QUERY_VEC}, {"top_k", 10},
+                               {"filters", nlohmann::json::array({"colour:eq:red"})}}, "a");
+    ASSERT_GT(plain, 0.0);
+    EXPECT_NEAR(plain, filtered, 1e-6);
+}

+ 13 - 0
tests/test_embeddings.cpp

@@ -293,3 +293,16 @@ TEST(EmbeddingCache, ReconfigureEvictsToTighterBounds) {
     EXPECT_FALSE(cache.get(kA).has_value());  // evicted
     EXPECT_TRUE(cache.get(kC).has_value());   // most-recently used survives
 }
+
+// A base URL that already ends in /v1 is the common OpenRouter mistake: the
+// client appends "/v1/embeddings" to whatever path the base carries, so the
+// request lands on /api/v1/v1/embeddings, which OpenRouter answers with 404.
+TEST(Embeddings, BaseEndingInV1DuplicatesTheVersionSegment) {
+    auto wrong = splitEmbeddingBase("https://openrouter.ai/api/v1");
+    EXPECT_EQ(wrong.origin, "https://openrouter.ai");
+    EXPECT_EQ(wrong.path, "/api/v1");
+    EXPECT_EQ(wrong.path + "/v1/embeddings", "/api/v1/v1/embeddings");
+
+    auto right = splitEmbeddingBase("https://openrouter.ai/api");
+    EXPECT_EQ(right.path + "/v1/embeddings", "/api/v1/embeddings");
+}