Эх сурвалжийг харах

fix(storage): apply query field projection in the adapter

The executions list was returning 32 MB for 20 rows, so the page took a long
time to render. ExecutionController already asked for summary fields only,
precisely to avoid this, but the projection stopped working when we moved to
the upstream database client: it has no server-side projection, so the adapter
logged a warning and returned whole documents. Nothing filtered in its place.

Execution documents carry nodeExecutions, which now hold base64 images from
the fetch, vision and generation nodes. In the largest row nodeExecutions was
7,690,987 of 7,709,959 bytes, 99.7% of the document, none of which the list
view reads.

The adapter now applies `fields` to the returned documents, always keeping
metadata keys so callers still get the id and version. Every caller that
passes `fields` benefits, rather than each having to filter for itself.

Measured on the same request: 32,461,603 bytes to 7,433. The single-execution
endpoint is untouched and still returns the full document.
fszontagh 1 сар өмнө
parent
commit
1be9dda784

+ 24 - 3
lib/storage/storage_client.cpp

@@ -155,9 +155,10 @@ Result<QueryResult> StorageClient::query(const std::string& collection, const Qu
         up.sortField = opts.sorts.front().first;
         up.sortDescending = !opts.sorts.front().second;
     }
-    if (!opts.fields.empty()) {
-        LOG_WARN("Storage: field projection not supported by upstream client; returning all fields");
-    }
+    // Upstream has no server-side projection, so it is applied below once the
+    // documents are back. That still saves callers from serialising fields they
+    // never asked for, which is the difference between a summary listing and one
+    // carrying every node output.
     up.limit = static_cast<uint32_t>(opts.page_size);
     up.offset = static_cast<uint32_t>(std::max(0, (opts.page - 1) * opts.page_size));
 
@@ -166,6 +167,26 @@ Result<QueryResult> StorageClient::query(const std::string& collection, const Qu
     out.documents = std::move(fr.documents);
     out.total_count = static_cast<int64_t>(fr.totalCount);
     out.has_more = fr.hasMore;
+
+    if (!opts.fields.empty()) {
+        for (auto& doc : out.documents) {
+            if (!doc.is_object()) {
+                continue;
+            }
+            nlohmann::json projected = nlohmann::json::object();
+            for (auto it = doc.begin(); it != doc.end(); ++it) {
+                // Metadata keys are always kept: callers need the id and version
+                // regardless of which payload fields they asked for.
+                const bool is_metadata = !it.key().empty() && it.key().front() == '_';
+                if (is_metadata ||
+                    std::find(opts.fields.begin(), opts.fields.end(), it.key()) != opts.fields.end()) {
+                    projected[it.key()] = it.value();
+                }
+            }
+            doc = std::move(projected);
+        }
+    }
+
     return out;
 }
 

+ 4 - 4
lib/storage/storage_client.hpp

@@ -26,10 +26,10 @@ struct StorageClientConfig {
     int max_message_size_mb = 64;            // kept for back-compat; ignored by adapter
 };
 
-// Query options - field projection and multi-sort kept for source compatibility,
-// but the upstream client supports only a single sort field. The adapter will
-// log a warning and use only the first entry of `sorts`. Field projection is
-// silently ignored by upstream; callers must filter in-process if needed.
+// Query options. The upstream client supports only a single sort field, so the
+// adapter logs a warning and uses the first entry of `sorts`. Upstream has no
+// server-side projection either, so `fields` is applied by the adapter after the
+// query returns; metadata keys are always retained.
 struct QueryOptions {
     std::vector<std::pair<std::string, nlohmann::json>> filters;
     std::vector<std::pair<std::string, bool>> sorts;  // field, ascending