ソースを参照

feat: keep file bytes in the file store, and stop guessing at document metadata

Downloads and generated images were written into the collection as base64
inside the document. That made anime_images 413 MB across 112 documents
and sdcpp_outputs 333 MB across 78 - and nothing ever read those bytes
back. No workflow reads .data from a stored document; they were written
once and carried for ever, through every query that touched the
collection.

Both nodes now put the bytes in the file store and keep in the document
what somebody would search or read: the name, the type, the size, where it
came from, and the id to fetch by. A 10 KB download that made a 14 KB
document now makes one of 913 bytes.

Two consequences handled rather than left to be discovered:

- Storing a file needs write access to "files" as well as to the
  collection, so a workflow that did this before fails at its next run.
  One did - the SD.cpp generator that 35photo2anime calls - and has been
  granted it.
- "files" could not be granted from the workflow settings at all: the list
  offers collections and the file store is not one. It has its own row
  now, saying what it is, rather than hiding as a name in a dropdown of
  every collection. Requiring a permission the interface cannot give is a
  fault in itself.

The refusal named the setting but not where it lives, which sent me to the
wrong level of the settings object; it now names
storagePermissions.collections.

Then the metadata. The database names document metadata one fixed way -
_created_at, _updated_at, in nanoseconds - and four places read _createdAt
instead, matched nothing and returned zero. Credentials, users, sessions
and node definitions all reported no timestamp, and the Database page's
Created/Updated panel keyed off fields that never exist, so it never
appeared at all. One reader now, TimeUtils::documentStamp: a single key,
because guessing at spellings is what hid this, and the unit checked
rather than divided blindly.

Also: a document listing summarises large string values. A page of
twenty-five image documents was 54 MB and the browser gave up at thirty
seconds. Opening a document fetches it whole, because editing a summary
and saving it would write the placeholder over the content. The database
side of that slowness is being fixed upstream; a listing still should not
ship a payload it will not show.

Verified: a stored download makes a 913-byte document with a fileId and no
data field, and the bytes come back whole - 10522 bytes, image/jpeg;
credentials show a real date; the file store appears in the settings.
fszontagh 1 ヶ月 前
親
コミット
35fbb8acb0

+ 8 - 0
lib/common/time_utils.cpp

@@ -5,6 +5,14 @@
 
 namespace smartbotic::common {
 
+int64_t TimeUtils::documentStamp(const nlohmann::json& document, const char* key) {
+    if (!document.contains(key) || !document[key].is_number()) return 0;
+    const int64_t value = document[key].get<int64_t>();
+    if (value <= 0) return 0;
+    constexpr int64_t kNanosecondThreshold = 100000000000000LL;  // year 5138 in ms
+    return value > kNanosecondThreshold ? value / 1000000 : value;
+}
+
 int64_t TimeUtils::nowMs() {
     return std::chrono::duration_cast<std::chrono::milliseconds>(
         Clock::now().time_since_epoch()

+ 14 - 0
lib/common/time_utils.hpp

@@ -1,5 +1,7 @@
 #pragma once
 
+#include <nlohmann/json.hpp>
+
 #include <chrono>
 #include <string>
 #include <ctime>
@@ -15,6 +17,18 @@ public:
     // Get current timestamp in milliseconds
     static int64_t nowMs();
 
+    /**
+     * A document's stored timestamp, in milliseconds.
+     *
+     * The database names document metadata one fixed way - _created_at,
+     * _updated_at - and writes it in nanoseconds. Reading it with a guessed
+     * spelling silently yields zero, which every screen renders as "unknown";
+     * that was true of credentials, users, sessions and node definitions at
+     * once. The unit is checked rather than divided blindly, because a value
+     * already in milliseconds divided again lands in 1970.
+     */
+    static int64_t documentStamp(const nlohmann::json& document, const char* key);
+
     // Get current timestamp in seconds
     static int64_t nowSec();
 

+ 8 - 2
lib/credentials/credential_types.cpp

@@ -55,6 +55,8 @@ static std::string base64Encode(const std::string& input) {
     return result;
 }
 
+
+
 // HttpAuth
 nlohmann::json HttpAuth::toJson() const {
     return {
@@ -302,8 +304,12 @@ CredentialMetadata CredentialMetadata::fromJson(const nlohmann::json& j) {
     meta.created_by = j.value("createdBy", j.value("created_by", ""));
     meta.declared_type = j.value("declaredType", "");
     meta.project_id = j.value("projectId", "");
-    meta.created_at = j.value("_createdAt", j.value("createdAt", j.value("created_at", int64_t{0})));
-    meta.updated_at = j.value("_updatedAt", j.value("updatedAt", j.value("updated_at", int64_t{0})));
+    // The database writes _created_at and _updated_at - snake_case, with the
+    // leading underscore, in nanoseconds. The chain here tried _createdAt,
+    // createdAt and created_at and matched none of them, so every credential
+    // reported a timestamp of zero and the screen said "Created unknown".
+    meta.created_at = common::TimeUtils::documentStamp(j, "_created_at");
+    meta.updated_at = common::TimeUtils::documentStamp(j, "_updated_at");
     if (j.contains("allowedWorkflows") && j["allowedWorkflows"].is_array()) {
         for (const auto& wf : j["allowedWorkflows"]) {
             meta.allowed_workflows.push_back(wf.get<std::string>());

+ 26 - 5
nodes/core/http-request.js

@@ -140,7 +140,7 @@ const configSchema = {
     storeDownload: {
       type: 'boolean',
       title: 'Store Download',
-      description: 'Store downloaded file in database (only for binary mode)',
+      description: 'Keep the downloaded file. The bytes go to the file store and a document in the collection below records what it is and where it came from',
       default: false
     },
     downloadCollection: {
@@ -198,10 +198,11 @@ const outputSchema = {
     },
     storage: {
       type: 'object',
-      description: 'Storage reference (only when storeDownload is true)',
+      description: 'Storage reference (only when storeDownload is true). The bytes live in the file store; the document holds the metadata and the fileId',
       properties: {
         collection: { type: 'string' },
-        id: { type: 'string' }
+        id: { type: 'string' },
+        fileId: { type: 'string', description: 'Fetch the bytes with this' }
       }
     }
   }
@@ -467,6 +468,25 @@ async function execute(config, input, context) {
 
         const checksum = contentHash;
 
+        // The bytes go to the file store, and the document keeps only what
+        // somebody would search or read: the name, the type, where it came
+        // from, and the id to fetch it by.
+        //
+        // Putting base64 in the document made every row megabytes wide, which
+        // the database has to parse in full for any query touching the
+        // collection - a listing of 25 downloads came to 54 MB. Nothing ever
+        // read those bytes back out of the document; they were written once and
+        // carried for ever.
+        const stored = smartbotic.storage.uploadFile(base64Data, {
+          name: filename,
+          mimeType: mimeType,
+          fileType: 'download',
+          relatedId: checksum
+        });
+        if (!stored.success) {
+          throw new Error('Could not store the download: ' + (stored.error || 'unknown'));
+        }
+
         const storageDoc = {
           filename: filename,
           mimeType: mimeType,
@@ -474,7 +494,7 @@ async function execute(config, input, context) {
           sourceUrl: url,
           checksum: checksum,
           downloadedAt: Date.now(),
-          data: base64Data
+          fileId: stored.id
         };
 
         const insertResult = smartbotic.storage.insert(collection, storageDoc, null, ttlMs);
@@ -483,7 +503,8 @@ async function execute(config, input, context) {
         } else {
           result.storage = {
             collection: collection,
-            id: insertResult.id
+            id: insertResult.id,
+            fileId: stored.id
           };
           smartbotic.log.info(`Stored download in ${collection}/${insertResult.id}`);
         }

+ 17 - 1
nodes/sdcpp/sdcpp-fetch-output.js

@@ -334,6 +334,22 @@ async function execute(config, input, context) {
             const ttlHours = config.storageTtlHours ?? 24;
             const ttlMs = ttlHours > 0 ? ttlHours * 60 * 60 * 1000 : 0;
 
+            // The image goes to the file store; the document records what it
+            // is and how to fetch it. Base64 in the document made a collection
+            // of 78 generated images 333 MB, which the database parses in full
+            // for any query that touches it - and nothing ever read those bytes
+            // back out of the document.
+            const stored = smartbotic.storage.uploadFile(base64Data, {
+                name: filename,
+                mimeType: mimeType,
+                fileType: 'generated',
+                relatedId: path
+            });
+            if (!stored.success) {
+                throw new Error('SD.cpp: downloaded "' + path + '" but storing the image failed: ' +
+                    (stored.error || 'unknown error'));
+            }
+
             const insertResult = smartbotic.storage.insert(collection, {
                 filename: filename,
                 mimeType: mimeType,
@@ -341,7 +357,7 @@ async function execute(config, input, context) {
                 sourcePath: path,
                 sourceUrl: fileObject.url,
                 downloadedAt: Date.now(),
-                data: base64Data
+                fileId: stored.id
             }, null, ttlMs);
 
             if (!insertResult.success) {

+ 4 - 4
src/runner/workflow_engine.cpp

@@ -1453,7 +1453,7 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
         -> common::Result<engine::ScriptFileInfo> {
         if (!canWriteCollection("files")) {
             return common::Error(common::ErrorCode::PermissionDenied,
-                "No write access to the file store (grant \"files\" in storagePermissions)");
+                "No write access to the file store (grant \"files\" in the workflow settings, under storagePermissions.collections)");
         }
 
         storage::FileMeta upstream;
@@ -1485,7 +1485,7 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
         -> common::Result<std::vector<uint8_t>> {
         if (!canReadCollection("files")) {
             return common::Error(common::ErrorCode::PermissionDenied,
-                "No read access to the file store (grant \"files\" in storagePermissions)");
+                "No read access to the file store (grant \"files\" in the workflow settings, under storagePermissions.collections)");
         }
         return storage_.downloadFile(id);
     };
@@ -1494,7 +1494,7 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
         -> common::Result<engine::ScriptFileInfo> {
         if (!canReadCollection("files")) {
             return common::Error(common::ErrorCode::PermissionDenied,
-                "No read access to the file store (grant \"files\" in storagePermissions)");
+                "No read access to the file store (grant \"files\" in the workflow settings, under storagePermissions.collections)");
         }
 
         auto result = storage_.getFileInfo(id);
@@ -1521,7 +1521,7 @@ NodeExecutionResult WorkflowEngine::executeNode(const WorkflowNode& node,
         -> common::Result<void> {
         if (!canWriteCollection("files")) {
             return common::Error(common::ErrorCode::PermissionDenied,
-                "No write access to the file store (grant \"files\" in storagePermissions)");
+                "No write access to the file store (grant \"files\" in the workflow settings, under storagePermissions.collections)");
         }
         return storage_.deleteFile(id);
     };

+ 36 - 1
src/webserver/api/database_controller.cpp

@@ -320,6 +320,35 @@ void DatabaseController::createCollection(const httplib::Request& req, httplib::
     sendJson(res, {{"name", name}, {"success", true}}, 201);
 }
 
+namespace {
+// Big enough that ordinary text, a prompt or a small config survives whole;
+// small enough that a page of documents cannot become tens of megabytes.
+constexpr size_t kListValueLimit = 4096;
+
+// A listing is for finding a document, not for carrying its contents. One
+// collection here holds a base64 image per document, which made page two of
+// the database browser a 42 MB response and page three 54 MB - the server
+// answered in about two seconds and the browser gave up at thirty.
+//
+// Large values are replaced by a note saying what was left out, and the
+// document says which fields those were, so nothing silently looks empty.
+// Whoever opens one fetches it whole by id.
+bool summariseForListing(nlohmann::json& document) {
+    nlohmann::json omitted = nlohmann::json::array();
+    for (auto& [key, value] : document.items()) {
+        if (!value.is_string()) continue;
+        const auto& text = value.get_ref<const std::string&>();
+        if (text.size() <= kListValueLimit) continue;
+        const size_t kb = text.size() / 1024;
+        value = "<<" + std::to_string(kb) + " KB not loaded - open this document to see it>>";
+        omitted.push_back(key);
+    }
+    if (omitted.empty()) return false;
+    document["_omittedFields"] = omitted;
+    return true;
+}
+}  // namespace
+
 void DatabaseController::getDocuments(const httplib::Request& req, httplib::Response& res,
                                        const auth::AuthContext& ctx) {
     std::string collection = req.matches[1];
@@ -347,11 +376,17 @@ void DatabaseController::getDocuments(const httplib::Request& req, httplib::Resp
     }
 
     nlohmann::json documents = nlohmann::json::array();
+    bool anySummarised = false;
     for (const auto& doc : result.value().documents) {
-        documents.push_back(doc);
+        nlohmann::json entry = doc;
+        anySummarised = summariseForListing(entry) || anySummarised;
+        documents.push_back(entry);
     }
 
     sendJson(res, {
+        // Says plainly that this listing is not the whole truth, so a caller
+        // cannot mistake it for one and write it back.
+        {"summarised", anySummarised},
         {"documents", documents},
         {"total", result.value().total_count},
         {"page", page},

+ 3 - 3
src/webserver/auth/auth_store.cpp

@@ -31,8 +31,8 @@ User User::fromJson(const nlohmann::json& j) {
     user.password_hash = j.value("passwordHash", "");
     user.role = j.value("role", "user");
     user.active = j.value("active", true);
-    user.created_at = j.value("_createdAt", int64_t{0});
-    user.updated_at = j.value("_updatedAt", int64_t{0});
+    user.created_at = common::TimeUtils::documentStamp(j, "_created_at");
+    user.updated_at = common::TimeUtils::documentStamp(j, "_updated_at");
     user.last_login = j.value("lastLogin", int64_t{0});
     return user;
 }
@@ -57,7 +57,7 @@ Session Session::fromJson(const nlohmann::json& j) {
     session.refresh_token = j.value("refreshToken", "");
     session.ip_address = j.value("ipAddress", "");
     session.user_agent = j.value("userAgent", "");
-    session.created_at = j.value("_createdAt", int64_t{0});
+    session.created_at = common::TimeUtils::documentStamp(j, "_created_at");
     session.expires_at = j.value("expiresAt", int64_t{0});
     return session;
 }

+ 2 - 2
src/webserver/nodes/node_store.cpp

@@ -235,8 +235,8 @@ StoredNode StoredNode::fromJson(const nlohmann::json& j) {
     node.is_trigger = j.value("isTrigger", false);
     node.is_scheduled = j.value("isScheduled", false);
     node.owner_id = j.value("ownerId", "");
-    node.created_at = j.value("_createdAt", int64_t{0});
-    node.updated_at = j.value("_updatedAt", int64_t{0});
+    node.created_at = common::TimeUtils::documentStamp(j, "_created_at");
+    node.updated_at = common::TimeUtils::documentStamp(j, "_updated_at");
 
     if (j.contains("configSchema")) {
         node.config_schema = j["configSchema"];

+ 48 - 9
tests/nodes/config-defaults-ttl-zero-http-request.json

@@ -2,19 +2,58 @@
   "name": "verify-config-defaults-ttl-zero-http-request",
   "settings": {
     "storagePermissions": {
-      "collections": { "ttl0_download_probe": "read-write" }
+      "collections": {
+        "ttl0_download_probe": "read-write",
+        "files": "read-write"
+      }
     }
   },
   "nodes": [
-    {"id": "n1", "name": "Trigger", "type": "click-trigger", "position": {"x": 0, "y": 0}, "config": {}},
-    {"id": "dl", "name": "Download", "type": "http-request", "position": {"x": 0, "y": 100},
-     "config": {"method": "GET", "url": "http://localhost:8090/index.html", "responseMode": "binary",
-       "storeDownload": true, "downloadCollection": "ttl0_download_probe", "downloadTtlHours": 0}}
+    {
+      "id": "n1",
+      "name": "Trigger",
+      "type": "click-trigger",
+      "position": {
+        "x": 0,
+        "y": 0
+      },
+      "config": {}
+    },
+    {
+      "id": "dl",
+      "name": "Download",
+      "type": "http-request",
+      "position": {
+        "x": 0,
+        "y": 100
+      },
+      "config": {
+        "method": "GET",
+        "url": "http://localhost:8090/index.html",
+        "responseMode": "binary",
+        "storeDownload": true,
+        "downloadCollection": "ttl0_download_probe",
+        "downloadTtlHours": 0
+      }
+    }
   ],
   "connections": [
-    {"sourceNodeId": "n1", "sourceOutput": "main", "targetNodeId": "dl", "targetInput": "data"}
+    {
+      "sourceNodeId": "n1",
+      "sourceOutput": "main",
+      "targetNodeId": "dl",
+      "targetInput": "data"
+    }
   ],
   "expect": {
-    "dl": {"status": "completed", "output": {"storage": {"collection": "ttl0_download_probe"}}}
-  }
-}
+    "dl": {
+      "status": "completed",
+      "output": {
+        "storage": {
+          "collection": "ttl0_download_probe"
+        }
+      }
+    }
+  },
+  "_comment": "The bytes of a stored download go to the file store now, not into the document - so this needs write access to \"files\" as well as to the collection. A document that carried its own base64 made a listing of twenty-five downloads 54 MB."
+}

+ 3 - 2
webui/src/api/credentials.ts

@@ -105,8 +105,9 @@ function transformCredential(data: any): CredentialInfo {
     ...data,
     id: data.id || data._id,
     createdBy: data.createdBy || data.ownerId,
-    createdAt: data.createdAt || data._createdAt,
-    updatedAt: data.updatedAt || data._updatedAt,
+    // Sent by the server in milliseconds; it reads the document metadata itself.
+    createdAt: data.createdAt,
+    updatedAt: data.updatedAt,
     declaredType: data.declaredType || undefined,
     allowedWorkflows: data.allowedWorkflows || [],
   }

+ 2 - 2
webui/src/api/workflowGroups.ts

@@ -20,8 +20,8 @@ function transformGroup(data: any): WorkflowGroup {
     parentId: data.parentId || undefined,
     ownerId: data.ownerId,
     // Same snake_case nanosecond metadata as workflows - see utils/timestamps.
-    createdAt: readTimestamp(data, '_created_at', '_createdAt'),
-    updatedAt: readTimestamp(data, '_updated_at', '_updatedAt'),
+    createdAt: readTimestamp(data, '_created_at'),
+    updatedAt: readTimestamp(data, '_updated_at'),
   }
 }
 

+ 6 - 6
webui/src/api/workflows.ts

@@ -84,12 +84,12 @@ function transformWorkflow(data: any): Workflow {
     nodes: data.nodes || [],
     connections: data.connections || [],
     settings: data.settings || {},
-    // The database writes _created_at / _updated_at in snake_case and in
-    // nanoseconds. Reading _createdAt returned undefined, which is why every
-    // workflow showed "unknown". Both spellings are tried and the unit is
-    // normalised - see utils/timestamps.
-    createdAt: readTimestamp(data, '_created_at', '_createdAt'),
-    updatedAt: readTimestamp(data, '_updated_at', '_updatedAt'),
+    // The database names document metadata one fixed way - _created_at,
+    // _updated_at - and writes it in nanoseconds. utils/timestamps normalises
+    // the unit; the key is not guessed at, because guessing is what made every
+    // workflow show "unknown" here in the first place.
+    createdAt: readTimestamp(data, '_created_at'),
+    updatedAt: readTimestamp(data, '_updated_at'),
     ownerId: data.ownerId,
     // _created_by is present on the document but has been empty on every record
     // inspected so far, so ownerId is the field that actually identifies a

+ 49 - 4
webui/src/components/workflow/WorkflowStorageSettings.tsx

@@ -1,6 +1,10 @@
 import { useState, useEffect, useMemo } from 'react'
 import { useQuery } from '@tanstack/react-query'
-import { Plus, Trash2, Database, ShieldAlert } from 'lucide-react'
+import { Plus, Trash2, Database, ShieldAlert, HardDrive } from 'lucide-react'
+
+// The reserved permission name for the file store. Nodes that keep a download
+// or a generated image write the bytes here and put only the id in a document.
+const FILE_STORE = 'files'
 
 interface StoragePermission {
   collection: string
@@ -31,8 +35,12 @@ export function WorkflowStorageSettings({ permissions, onChange }: WorkflowStora
   })
 
   const availableCollections = useMemo(() => {
-    if (!collectionsData?.collections) return []
-    return collectionsData.collections as string[]
+    const collections = (collectionsData?.collections as string[]) || []
+    // "files" is not a collection - it is the reserved name for the file store,
+    // where nodes put the bytes of a download or a generated image. It was
+    // impossible to grant from here at all, so a node that needed it could only
+    // be fixed by editing the workflow over the API.
+    return collections.filter((c) => c !== FILE_STORE)
   }, [collectionsData])
 
   // System collections that cannot be accessed by workflows
@@ -55,6 +63,15 @@ export function WorkflowStorageSettings({ permissions, onChange }: WorkflowStora
     }
   }, [permissions])
 
+  // The file store is set by its own row above, so it is not shown among the
+  // collections as well - one setting, one place.
+  const handleFileStoreChange = (access: string) => {
+    const others = localCollections.filter((p) => p.collection !== FILE_STORE)
+    const next = access === 'none' ? others : [...others, { collection: FILE_STORE, access: access as any }]
+    setLocalCollections(next)
+    updatePermissions(next, defaultAccess)
+  }
+
   // Notify parent of changes
   const updatePermissions = (newCollections: StoragePermission[], newDefaultAccess: string) => {
     const collectionsObj: Record<string, string> = {}
@@ -133,6 +150,34 @@ export function WorkflowStorageSettings({ permissions, onChange }: WorkflowStora
         </div>
       </div>
 
+      {/* The file store, as its own row.
+          It was reachable only by finding "files" in a dropdown of every
+          collection, with nothing to say it existed or that a node needed it -
+          which is the same as not being there. */}
+      <div className="p-3 bg-gray-50 dark:bg-slate-900 rounded-lg border border-gray-200 dark:border-slate-700">
+        <div className="flex items-center justify-between gap-3">
+          <div className="min-w-0">
+            <div className="text-sm font-medium text-gray-700 dark:text-gray-300 flex items-center gap-2">
+              <HardDrive className="w-4 h-4 text-gray-400" />
+              File store
+            </div>
+            <div className="text-xs text-gray-500 dark:text-gray-400">
+              Where a download or a generated image keeps its bytes. Nodes that store a
+              file need write access here as well as to the collection that records it.
+            </div>
+          </div>
+          <select
+            value={localCollections.find((p) => p.collection === FILE_STORE)?.access || 'none'}
+            onChange={(e) => handleFileStoreChange(e.target.value)}
+            className="px-3 py-1.5 text-sm border border-gray-200 dark:border-slate-600 rounded-lg bg-white dark:bg-slate-800 text-gray-900 dark:text-gray-100 focus:ring-2 focus:ring-primary-500 shrink-0"
+          >
+            <option value="none">No Access</option>
+            <option value="read-only">Read Only</option>
+            <option value="read-write">Read & Write</option>
+          </select>
+        </div>
+      </div>
+
       {/* Collection permissions list */}
       <div className="space-y-2">
         <div className="text-sm font-medium text-gray-700 dark:text-gray-300">Collection Permissions</div>
@@ -143,7 +188,7 @@ export function WorkflowStorageSettings({ permissions, onChange }: WorkflowStora
           </div>
         ) : (
           <div className="border border-gray-200 dark:border-slate-700 rounded-lg divide-y divide-gray-200 dark:divide-slate-700">
-            {localCollections.map((perm, idx) => (
+            {localCollections.filter((p) => p.collection !== FILE_STORE).map((perm, idx) => (
               <div key={perm.collection} className="flex items-center gap-3 p-3">
                 <Database className="w-4 h-4 text-gray-400 dark:text-gray-500" />
                 <span className="flex-1 text-sm font-mono text-gray-900 dark:text-gray-100">{perm.collection}</span>

+ 47 - 14
webui/src/pages/DatabasePage.tsx

@@ -28,10 +28,15 @@ import { useTheme } from '../contexts/ThemeContext'
 
 interface Document {
   _id: string
-  _createdAt?: number
-  _updatedAt?: number
-  _createdBy?: string
-  _updatedBy?: string
+  // The database names these one fixed way, in nanoseconds. Declaring the
+  // camelCase spellings meant the Created/Updated panel below keyed off fields
+  // that never exist, so it never appeared at all.
+  /** Fields the listing left out because they were large. Set only in a listing. */
+  _omittedFields?: string[]
+  _created_at?: number
+  _updated_at?: number
+  _created_by?: string
+  _updated_by?: string
   [key: string]: any
 }
 
@@ -70,6 +75,7 @@ export default function DatabasePage() {
   const [selectedCollection, setSelectedCollection] = useState<string | null>(null)
   const [selectedDocuments, setSelectedDocuments] = useState<Set<string>>(new Set())
   const [editingDocument, setEditingDocument] = useState<Document | null>(null)
+  const [loadingDocument, setLoadingDocument] = useState(false)
   const [editingJson, setEditingJson] = useState('')
   const [originalJson, setOriginalJson] = useState('')
   const [jsonError, setJsonError] = useState<string | null>(null)
@@ -306,8 +312,25 @@ export default function DatabasePage() {
       // Single selection - also open document for editing
       setSelectedDocuments(new Set([doc._id]))
       guardUnsavedChanges(() => {
-        setEditingDocument(doc)
         setShowNewDocument(false)
+        // The listing leaves large values out, so the row in hand may be a
+        // summary. Editing it as-is would show a placeholder where the content
+        // should be, and saving would write that placeholder over the real
+        // thing - so the document is fetched whole before it can be edited.
+        if (doc._omittedFields && selectedCollection) {
+          setEditingDocument(null)
+          setLoadingDocument(true)
+          databaseApi
+            .getDocument(selectedCollection, doc._id)
+            .then((full) => setEditingDocument(full))
+            .catch(() => {
+              setEditingDocument(null)
+              setJsonError('Could not read that document in full - it has not been opened, so nothing can overwrite it')
+            })
+            .finally(() => setLoadingDocument(false))
+        } else {
+          setEditingDocument(doc)
+        }
       })
     }
   }
@@ -783,24 +806,24 @@ export default function DatabasePage() {
                 />
               </div>
               {/* Document metadata */}
-              {editingDocument && (editingDocument._createdAt || editingDocument._updatedAt) && (
+              {editingDocument && (editingDocument._created_at || editingDocument._updated_at) && (
                 <div className="px-4 pb-4 flex-shrink-0">
                   <div className="p-3 bg-gray-50 dark:bg-slate-700 rounded-lg text-xs text-gray-600 dark:text-gray-400 grid grid-cols-2 gap-2">
-                    {editingDocument._createdAt && (
+                    {editingDocument._created_at && (
                       <div className="flex items-center gap-2">
                         <Clock className="w-3 h-3" />
-                        <span>Created: {formatTimestamp(editingDocument._createdAt)}</span>
-                        {editingDocument._createdBy && (
-                          <span className="text-gray-400 dark:text-gray-500">by {editingDocument._createdBy}</span>
+                        <span>Created: {formatTimestamp(editingDocument._created_at)}</span>
+                        {editingDocument._created_by && (
+                          <span className="text-gray-400 dark:text-gray-500">by {editingDocument._created_by}</span>
                         )}
                       </div>
                     )}
-                    {editingDocument._updatedAt && (
+                    {editingDocument._updated_at && (
                       <div className="flex items-center gap-2">
                         <Clock className="w-3 h-3" />
-                        <span>Updated: {formatTimestamp(editingDocument._updatedAt)}</span>
-                        {editingDocument._updatedBy && (
-                          <span className="text-gray-400 dark:text-gray-500">by {editingDocument._updatedBy}</span>
+                        <span>Updated: {formatTimestamp(editingDocument._updated_at)}</span>
+                        {editingDocument._updated_by && (
+                          <span className="text-gray-400 dark:text-gray-500">by {editingDocument._updated_by}</span>
                         )}
                       </div>
                     )}
@@ -811,12 +834,22 @@ export default function DatabasePage() {
           ) : (
             <div className="h-full flex items-center justify-center text-gray-500 dark:text-gray-400">
               <div className="text-center">
+                {loadingDocument ? (
+                  <>
+                    <div className="animate-spin rounded-full h-8 w-8 border-b-2 border-primary-600 mx-auto mb-4" />
+                    <p>Reading the whole document...</p>
+                    <p className="text-sm">The list leaves large values out; this is fetching the real one.</p>
+                  </>
+                ) : (
+                <>
                 <FileJson className="w-12 h-12 mx-auto mb-4 opacity-50" />
                 <p>Select a document to edit</p>
                 <p className="text-sm">or create a new one</p>
                 <p className="text-xs mt-4 text-gray-400 dark:text-gray-500">
                   Tip: Ctrl+Click to multi-select, Shift+Click for range select
                 </p>
+                </>
+                )}
               </div>
             </div>
           )}