Quellcode durchsuchen

fix: handle non-ASCII in every email field, both directions

Asked for after an accented attachment filename was silently discarded. The
audit found the same gap in most other fields, in both directions.

Inbound, in imap_client so every consumer benefits - trigger, fetch and search
alike - Subject, From, To and Cc are decoded from RFC 2047. Workflows were
being handed the raw encoded word: from was literally
"=?UTF-8?Q?Szont=C3=A1gh_Ferenc?= <ferenc.szontagh@smartbotics.ai>". Cc was not
read out of the message at all.

Outbound, Subject, the From display name, To, Cc, Reply-To and attachment
filenames are encoded. Filenames use RFC 2231 with an ASCII fallback beside
them. The body declared charset=UTF-8 and no transfer encoding at all - 8-bit
data in a channel defined as 7-bit, carried only by the servers that choose to
tolerate it - and is now base64 with wrapped lines. Bcc was already right:
envelope only, never a header.

Three rules the address handling turns on, none of them optional:

- Only the display name may be touched. Encoding the address itself would
  corrupt it, so anything that cannot be split at angle brackets with
  confidence is left exactly as it came.
- An encoded word must NOT be quoted. A conforming reader treats a quoted one
  as literal text, which is how a name ends up displayed as =?UTF-8?B?...?=.
- An ASCII display name containing a comma MUST be quoted, or "Doe, John
  <a@b.com>" reads as two addresses and one recipient becomes two.

Fourth bug, found while checking the third: all three IMAP nodes copied fields
by name, so the cc the client had just started returning was set in C++ and
dropped in JS with nothing failing. They spread now. That is the same trap that
has already cost this repo three bugs.

tests/cpp/mime_words_test.cpp pins all of it, including round trips, and
scripts/run-cpp-tests.sh compiles it in the build image - zeus cannot build the
project on the host at all, so a test that needs a toolchain has to bring one.
fszontagh vor 3 Wochen
Ursprung
Commit
b68c500647

+ 11 - 0
CLAUDE.md

@@ -51,6 +51,17 @@ There is no `systemctl` on zeus at all - it answers "command not found", which
 reads like a typo rather than the wrong init system. The `systemd/user/*.service`
 units in this repo are leftovers from when the services ran on mulan.
 
+### Tests
+
+```bash
+./scripts/run-node-tests.sh      # every node, through a running webserver + runner
+./scripts/run-cpp-tests.sh       # pure C++ logic, compiled in the build image
+```
+
+Node tests need the services up. When the runner and webserver are on different
+hosts, a case whose NODE fetches the API needs `SMARTBOTIC_RUNNER_API` set to an
+address the runner can reach - see the `{{API}}` note under Relocating a service.
+
 ### Frontend (webui/)
 ```bash
 cd webui

+ 1 - 0
CMakeLists.txt

@@ -22,6 +22,7 @@ add_library(smartbotic_common STATIC
     lib/common/config_defaults.cpp
     lib/common/config_arg.cpp
     lib/common/cron.cpp
+    lib/common/mime_words.cpp
 )
 target_include_directories(smartbotic_common PUBLIC
     ${CMAKE_CURRENT_SOURCE_DIR}/lib

+ 260 - 0
lib/common/mime_words.cpp

@@ -0,0 +1,260 @@
+#include "common/mime_words.hpp"
+
+#include "common/string_utils.hpp"
+
+#include <algorithm>
+#include <cctype>
+#include <sstream>
+
+namespace smartbotic::common {
+
+namespace {
+
+int hexValue(char c) {
+    if (c >= '0' && c <= '9') return c - '0';
+    if (c >= 'a' && c <= 'f') return c - 'a' + 10;
+    if (c >= 'A' && c <= 'F') return c - 'A' + 10;
+    return -1;
+}
+
+std::string upper(std::string text) {
+    std::transform(text.begin(), text.end(), text.begin(),
+                   [](unsigned char c) { return std::toupper(c); });
+    return text;
+}
+
+// ISO-8859-1 and -2 are still what a good deal of European mail says it is.
+// Only 8859-1 maps trivially - each byte is the code point - so that is the one
+// converted; anything else is left as it arrived rather than guessed at.
+std::string latin1ToUtf8(const std::string& input) {
+    std::string out;
+    out.reserve(input.size() * 2);
+    for (unsigned char c : input) {
+        if (c < 0x80) {
+            out.push_back(static_cast<char>(c));
+        } else {
+            out.push_back(static_cast<char>(0xC0 | (c >> 6)));
+            out.push_back(static_cast<char>(0x80 | (c & 0x3F)));
+        }
+    }
+    return out;
+}
+
+std::string decodeQ(const std::string& payload) {
+    std::string out;
+    out.reserve(payload.size());
+    for (size_t i = 0; i < payload.size(); ++i) {
+        const char c = payload[i];
+        if (c == '_') {
+            // In an encoded word an underscore is a space, not an underscore.
+            out.push_back(' ');
+        } else if (c == '=' && i + 2 < payload.size()) {
+            const int hi = hexValue(payload[i + 1]);
+            const int lo = hexValue(payload[i + 2]);
+            if (hi >= 0 && lo >= 0) {
+                out.push_back(static_cast<char>((hi << 4) | lo));
+                i += 2;
+            } else {
+                out.push_back(c);
+            }
+        } else {
+            out.push_back(c);
+        }
+    }
+    return out;
+}
+
+std::string toUtf8(const std::string& charset, const std::string& bytes) {
+    const std::string cs = upper(charset);
+    if (cs == "UTF-8" || cs == "UTF8" || cs == "US-ASCII" || cs == "ASCII") {
+        return bytes;
+    }
+    if (cs == "ISO-8859-1" || cs == "LATIN1" || cs == "ISO8859-1" || cs == "WINDOWS-1252") {
+        return latin1ToUtf8(bytes);
+    }
+    // An unknown charset is passed through. It may render oddly; it will not
+    // disappear, and the caller can still see what arrived.
+    return bytes;
+}
+
+} // namespace
+
+bool MimeWords::isAscii(const std::string& value) {
+    return std::all_of(value.begin(), value.end(),
+                       [](unsigned char c) { return c < 0x80; });
+}
+
+std::string MimeWords::decode(const std::string& value) {
+    if (value.find("=?") == std::string::npos) {
+        return value;
+    }
+
+    std::string out;
+    out.reserve(value.size());
+
+    size_t pos = 0;
+    bool previous_was_word = false;
+    while (pos < value.size()) {
+        const size_t start = value.find("=?", pos);
+        if (start == std::string::npos) {
+            out.append(value, pos, std::string::npos);
+            break;
+        }
+
+        // Whitespace between two encoded words is separator, not content, and
+        // is dropped - that is how a value too long for one word is rejoined.
+        std::string between = value.substr(pos, start - pos);
+        const bool only_space = !between.empty() &&
+            std::all_of(between.begin(), between.end(),
+                        [](unsigned char c) { return std::isspace(c); });
+        if (!(previous_was_word && only_space)) {
+            out += between;
+        }
+
+        const size_t charset_end = value.find('?', start + 2);
+        if (charset_end == std::string::npos) { out += value.substr(start); break; }
+        const size_t encoding_end = value.find('?', charset_end + 1);
+        if (encoding_end == std::string::npos) { out += value.substr(start); break; }
+        const size_t word_end = value.find("?=", encoding_end + 1);
+        if (word_end == std::string::npos) { out += value.substr(start); break; }
+
+        const std::string charset = value.substr(start + 2, charset_end - start - 2);
+        const std::string encoding = value.substr(charset_end + 1, encoding_end - charset_end - 1);
+        const std::string payload = value.substr(encoding_end + 1, word_end - encoding_end - 1);
+
+        std::string bytes;
+        if (encoding == "B" || encoding == "b") {
+            bytes = StringUtils::base64Decode(payload);
+        } else if (encoding == "Q" || encoding == "q") {
+            bytes = decodeQ(payload);
+        } else {
+            // Not an encoding we know - keep the word as written.
+            out += value.substr(start, word_end + 2 - start);
+            pos = word_end + 2;
+            previous_was_word = false;
+            continue;
+        }
+
+        out += toUtf8(charset, bytes);
+        pos = word_end + 2;
+        previous_was_word = true;
+    }
+
+    return out;
+}
+
+std::string MimeWords::encode(const std::string& value) {
+    if (isAscii(value)) {
+        return value;
+    }
+
+    // Base64 rather than Q: the values that need encoding here are mostly
+    // non-ASCII throughout, where Q would escape nearly every character.
+    //
+    // A word may be 75 characters including the `=?UTF-8?B?` and `?=` around
+    // it, so the payload is capped at 45 bytes - a multiple of 3, so no word
+    // ends mid-base64-group and each decodes on its own. Splitting is done on
+    // UTF-8 character boundaries so a multi-byte character is never cut in two.
+    constexpr size_t kMaxBytes = 45;
+    std::string out;
+    size_t pos = 0;
+    while (pos < value.size()) {
+        size_t take = std::min(kMaxBytes, value.size() - pos);
+        // Do not split inside a UTF-8 sequence: back off while the next byte is
+        // a continuation byte.
+        while (take > 0 && pos + take < value.size() &&
+               (static_cast<unsigned char>(value[pos + take]) & 0xC0) == 0x80) {
+            --take;
+        }
+        if (take == 0) take = std::min(kMaxBytes, value.size() - pos);
+
+        if (!out.empty()) out += "\r\n ";  // fold, as a long header must
+        out += "=?UTF-8?B?" + StringUtils::base64Encode(value.substr(pos, take)) + "?=";
+        pos += take;
+    }
+    return out;
+}
+
+std::string MimeWords::encodeAddress(const std::string& address) {
+    // `Display Name <local@domain>`. Only the display name may be touched; the
+    // address inside the angle brackets must go out exactly as it came.
+    const size_t open = address.rfind('<');
+    const size_t close = address.rfind('>');
+    if (open == std::string::npos || close == std::string::npos || close < open) {
+        // A bare address, or something we cannot split with confidence.
+        // Encoding the whole of it would corrupt the address itself, so it is
+        // left alone even when it is not ASCII.
+        return address;
+    }
+
+    std::string name = address.substr(0, open);
+    const std::string addr = address.substr(open, close - open + 1);
+
+    while (!name.empty() && std::isspace(static_cast<unsigned char>(name.back()))) {
+        name.pop_back();
+    }
+    // Often already quoted; the quotes must not end up inside the encoded word
+    // or doubled.
+    if (name.size() >= 2 && name.front() == '"' && name.back() == '"') {
+        name = name.substr(1, name.size() - 2);
+    }
+    if (name.empty()) {
+        return addr;
+    }
+
+    if (!isAscii(name)) {
+        // An encoded word is its own token and must NOT be wrapped in quotes -
+        // a quoted encoded word is treated as literal text by a conforming
+        // reader, which is how a name arrives looking like =?UTF-8?B?...?=.
+        return encode(name) + " " + addr;
+    }
+
+    // ASCII, but a display name carrying any of the specials RFC 5322 reserves
+    // has to be quoted or it breaks the address list it sits in - a comma in
+    // "Doe, John" otherwise reads as the end of one address and the start of
+    // another.
+    const std::string specials = "()<>[]:;@\\,.\"";
+    if (name.find_first_of(specials) != std::string::npos) {
+        std::string escaped;
+        escaped.reserve(name.size() + 2);
+        for (char c : name) {
+            if (c == '"' || c == '\\') escaped.push_back('\\');
+            escaped.push_back(c);
+        }
+        return "\"" + escaped + "\" " + addr;
+    }
+
+    return name + " " + addr;
+}
+
+std::string MimeWords::encodeParameter(const std::string& name, const std::string& value) {
+    if (isAscii(value)) {
+        return name + "=\"" + value + "\"";
+    }
+
+    // RFC 2231. An ASCII-only fallback is emitted alongside for readers that do
+    // not implement it; a reader that does prefers the starred form.
+    std::string ascii;
+    ascii.reserve(value.size());
+    for (unsigned char c : value) {
+        ascii.push_back(c < 0x80 ? static_cast<char>(c) : '_');
+    }
+
+    std::ostringstream percent;
+    percent << std::hex;
+    for (unsigned char c : value) {
+        const bool safe = (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') ||
+                          (c >= '0' && c <= '9') || c == '.' || c == '-' || c == '_';
+        if (safe) {
+            percent << static_cast<char>(c);
+        } else {
+            percent << '%' << (c < 16 ? "0" : "")
+                    << upper(std::string(1, "0123456789abcdef"[c >> 4]))
+                    << upper(std::string(1, "0123456789abcdef"[c & 0x0F]));
+        }
+    }
+
+    return name + "=\"" + ascii + "\"; " + name + "*=UTF-8''" + percent.str();
+}
+
+} // namespace smartbotic::common

+ 44 - 0
lib/common/mime_words.hpp

@@ -0,0 +1,44 @@
+#pragma once
+
+#include <string>
+
+namespace smartbotic::common {
+
+// RFC 2047 encoded words, both directions.
+//
+// A mail header is defined to be ASCII. Anything else - an accented name, a
+// subject in Hungarian, a filename with a Sárkány in it - travels as an encoded
+// word, `=?UTF-8?Q?...?=`, and long values are split into several of them and
+// folded across lines. Neither direction was handled here: inbound values were
+// passed on still encoded, and outbound values were written raw, which is not
+// valid and is mangled by the servers that do not quietly tolerate it.
+class MimeWords {
+public:
+    // Decode any encoded words in a header value, leaving the rest alone.
+    // Unknown charsets and malformed words are returned untouched rather than
+    // dropped - a name that survives as gibberish is better than one that
+    // vanishes.
+    static std::string decode(const std::string& value);
+
+    // Encode a header value if it needs it, as one or more `=?UTF-8?B?...?=`
+    // words no longer than RFC 2047 allows. Pure ASCII is returned unchanged,
+    // so ordinary headers stay readable in a raw message.
+    static std::string encode(const std::string& value);
+
+    // Encode the display name of an address, keeping the address itself
+    // untouched: `Szontágh Ferenc <a@b>` becomes `=?UTF-8?B?...?= <a@b>`.
+    // An address with no display name is returned as it came.
+    static std::string encodeAddress(const std::string& address);
+
+    // A MIME parameter such as a filename. Emits the RFC 2231 form
+    // (`filename*=UTF-8''...`) when the value is not ASCII, which is what
+    // modern clients read, and only the plain form when it is.
+    //
+    // `name` is the parameter name, e.g. "filename" or "name".
+    static std::string encodeParameter(const std::string& name, const std::string& value);
+
+    // Whether a string is plain ASCII, and so needs none of the above.
+    static bool isAscii(const std::string& value);
+};
+
+} // namespace smartbotic::common

+ 10 - 10
nodes/imap/imap-fetch.js

@@ -58,9 +58,10 @@ const outputSchema = {
     type: 'object',
     properties: {
         uid: { type: 'string', description: 'Unique email identifier' },
-        subject: { type: 'string' },
-        from: { type: 'string' },
-        to: { type: 'string' },
+        subject: { type: 'string', description: 'Decoded - RFC 2047 encoded words are resolved by the IMAP client' },
+        from: { type: 'string', description: 'Decoded, in "Name <address>" form' },
+        to: { type: 'string', description: 'Decoded, in "Name <address>" form' },
+        cc: { type: 'string', description: 'Decoded, in "Name <address>" form. Empty when the message carries no Cc' },
         date: { type: 'string' },
         body: { type: 'string', description: 'Email body text' },
         raw: { type: 'string', description: 'Raw email content (for attachment extraction)' },
@@ -117,14 +118,13 @@ module.exports = {
             throw new Error(`IMAP fetch failed: ${fetchResult.error}`);
         }
 
+        // Spread, not a field-by-field copy. Listing them by name is how cc
+        // went missing the day it was added to the fetch result: the C++ side
+        // set it, this dropped it, and nothing failed. `success` is removed
+        // because it belongs to the call, not to the email.
+        const { success, ...email } = fetchResult;
         return {
-            uid: fetchResult.uid,
-            subject: fetchResult.subject,
-            from: fetchResult.from,
-            to: fetchResult.to,
-            date: fetchResult.date,
-            body: fetchResult.body,
-            raw: fetchResult.raw,
+            ...email,
             hasAttachments: detectAttachments(fetchResult.raw),
             mailbox
         };

+ 8 - 10
nodes/imap/imap-search.js

@@ -109,9 +109,10 @@ const outputSchema = {
                 type: 'object',
                 properties: {
                     uid: { type: 'string', description: 'Unique email identifier' },
-                    subject: { type: 'string' },
-                    from: { type: 'string' },
-                    to: { type: 'string' },
+                    subject: { type: 'string', description: 'Decoded - RFC 2047 encoded words are resolved by the IMAP client' },
+                    from: { type: 'string', description: 'Decoded, in "Name <address>" form' },
+                    to: { type: 'string', description: 'Decoded, in "Name <address>" form' },
+                    cc: { type: 'string', description: 'Decoded, in "Name <address>" form. Empty when the message carries no Cc' },
                     date: { type: 'string' },
                     body: { type: 'string', description: 'Email body' },
                     raw: { type: 'string', description: 'Raw email content' },
@@ -211,13 +212,10 @@ module.exports = {
 
                 if (fetchResult.success) {
                     result.emails.push({
-                        uid: fetchResult.uid,
-                        subject: fetchResult.subject,
-                        from: fetchResult.from,
-                        to: fetchResult.to,
-                        date: fetchResult.date,
-                        body: fetchResult.body,
-                        raw: fetchResult.raw,
+                        // Spread rather than a field-by-field copy - naming them
+                        // is how cc was silently dropped when the IMAP client
+                        // started returning it.
+                        ...(({ success, ...rest }) => rest)(fetchResult),
                         hasAttachments: detectAttachments(fetchResult.raw)
                     });
                 }

+ 8 - 10
nodes/imap/imap-trigger.js

@@ -84,9 +84,10 @@ const outputSchema = {
                 type: 'object',
                 properties: {
                     uid: { type: 'string', description: 'Unique email identifier' },
-                    subject: { type: 'string' },
-                    from: { type: 'string' },
-                    to: { type: 'string' },
+                    subject: { type: 'string', description: 'Decoded - RFC 2047 encoded words are resolved by the IMAP client' },
+                    from: { type: 'string', description: 'Decoded, in "Name <address>" form' },
+                    to: { type: 'string', description: 'Decoded, in "Name <address>" form' },
+                    cc: { type: 'string', description: 'Decoded, in "Name <address>" form. Empty when the message carries no Cc' },
                     date: { type: 'string' },
                     body: { type: 'string', description: 'Email body (if fetchContent enabled)' },
                     raw: { type: 'string', description: 'Raw email content (if fetchContent enabled)' },
@@ -182,13 +183,10 @@ module.exports = {
 
                 if (fetchResult.success) {
                     const email = {
-                        uid: fetchResult.uid,
-                        subject: fetchResult.subject,
-                        from: fetchResult.from,
-                        to: fetchResult.to,
-                        date: fetchResult.date,
-                        body: fetchResult.body,
-                        raw: fetchResult.raw,
+                        // Spread rather than a field-by-field copy - naming them
+                        // is how cc was silently dropped when the IMAP client
+                        // started returning it.
+                        ...(({ success, ...rest }) => rest)(fetchResult),
                         hasAttachments: detectAttachments(fetchResult.raw)
                     };
                     emails.push(email);

+ 34 - 0
scripts/run-cpp-tests.sh

@@ -0,0 +1,34 @@
+#!/usr/bin/env bash
+# Standalone C++ tests for pure logic that needs no running service.
+#
+# Compiled and run inside the build image, so the dependencies are the ones the
+# real build uses and nothing has to be installed on the host - which matters on
+# zeus, where the host cannot build the project at all.
+set -uo pipefail
+cd "$(dirname "$0")/.."
+
+IMAGE="${BUILD_IMAGE:-smartbotic-automation-build-base:debian13}"
+if ! docker image inspect "$IMAGE" >/dev/null 2>&1; then
+    echo "build image $IMAGE not found - build it with packaging/Dockerfile.base" >&2
+    exit 1
+fi
+
+fail=0
+for test_src in tests/cpp/*_test.cpp; do
+    name=$(basename "$test_src" .cpp)
+    # Each test names the sources it needs on a "// sources:" line, so adding a
+    # test does not mean editing this script.
+    sources=$(grep -m1 '^// sources:' "$test_src" | sed 's|^// sources:||')
+    [ -z "$sources" ] && sources="lib/common/${name%_test}.cpp lib/common/string_utils.cpp"
+
+    echo "=== $name"
+    docker run --rm -v "$PWD":/src -w /src "$IMAGE" sh -c \
+        "g++ -std=c++20 -I lib -o /tmp/$name $test_src $sources -lssl -lcrypto && /tmp/$name"
+    [ $? -ne 0 ] && fail=$((fail + 1))
+done
+
+if [ $fail -ne 0 ]; then
+    echo "$fail C++ test binary/binaries failed"
+    exit 1
+fi
+echo "C++ tests passed"

+ 1 - 0
src/runner/engine/script_engine.cpp

@@ -4220,6 +4220,7 @@ static JSValue js_imap_fetch(JSContext* ctx, JSValue this_val, int argc, JSValue
     JS_SetPropertyStr(ctx, result, "subject", JS_NewString(ctx, email.subject.c_str()));
     JS_SetPropertyStr(ctx, result, "from", JS_NewString(ctx, email.from.c_str()));
     JS_SetPropertyStr(ctx, result, "to", JS_NewString(ctx, email.to.c_str()));
+    JS_SetPropertyStr(ctx, result, "cc", JS_NewString(ctx, email.cc.c_str()));
     JS_SetPropertyStr(ctx, result, "date", JS_NewString(ctx, email.date.c_str()));
     JS_SetPropertyStr(ctx, result, "body", JS_NewString(ctx, email.body.c_str()));
     JS_SetPropertyStr(ctx, result, "raw", JS_NewString(ctx, email.raw.c_str()));

+ 10 - 3
src/runner/imap/imap_client.cpp

@@ -1,4 +1,5 @@
 #include "imap_client.hpp"
+#include "common/mime_words.hpp"
 #include <curl/curl.h>
 #include <sstream>
 #include <algorithm>
@@ -233,13 +234,19 @@ common::Result<EmailContent> ImapClient::fetch(
     std::string current_header_name;
     std::string current_header_value;
 
+    // Decoded here rather than in each node, so every consumer of an IMAP
+    // message - trigger, fetch, search - sees text rather than encoded words.
+    // A header with nothing to decode passes through untouched. Date is left
+    // alone deliberately: it is a machine format and never encoded.
     auto process_header = [&content](const std::string& name, const std::string& value) {
         if (name == "subject") {
-            content.subject = value;
+            content.subject = common::MimeWords::decode(value);
         } else if (name == "from") {
-            content.from = value;
+            content.from = common::MimeWords::decode(value);
         } else if (name == "to") {
-            content.to = value;
+            content.to = common::MimeWords::decode(value);
+        } else if (name == "cc") {
+            content.cc = common::MimeWords::decode(value);
         } else if (name == "date") {
             content.date = value;
         }

+ 5 - 0
src/runner/imap/imap_client.hpp

@@ -15,6 +15,10 @@ struct ImapCredentials {
     bool use_ssl = true;
 };
 
+// Declared but never populated anywhere. Left in place rather than deleted
+// because it is part of a published header, but note it does NOT get the
+// RFC 2047 decoding EmailContent does - anything that starts filling this in
+// must decode subject and from itself, or call the same MimeWords::decode.
 struct EmailHeader {
     std::string uid;
     std::string subject;
@@ -30,6 +34,7 @@ struct EmailContent {
     std::string subject;
     std::string from;
     std::string to;
+    std::string cc;
     std::string date;
     std::string body;
     std::string raw;

+ 40 - 10
src/runner/smtp/smtp_client.cpp

@@ -1,5 +1,8 @@
 #include "smtp_client.hpp"
 
+#include "common/mime_words.hpp"
+#include "common/string_utils.hpp"
+
 #include <curl/curl.h>
 
 #include <chrono>
@@ -36,11 +39,26 @@ size_t readPayload(char* buffer, size_t size, size_t nitems, void* userdata) {
     return count;
 }
 
+// SMTP lines must not exceed 998 characters, so base64 is wrapped. 76 is the
+// customary width and leaves room for the CRLF.
+std::string wrapBase64(const std::string& encoded) {
+    std::string out;
+    out.reserve(encoded.size() + encoded.size() / 76 * 2 + 2);
+    for (size_t i = 0; i < encoded.size(); i += 76) {
+        out += encoded.substr(i, 76);
+        out += "\r\n";
+    }
+    if (encoded.empty()) out += "\r\n";
+    return out;
+}
+
 std::string joinAddresses(const std::vector<std::string>& addresses) {
     std::string joined;
     for (size_t i = 0; i < addresses.size(); ++i) {
         if (i > 0) joined += ", ";
-        joined += addresses[i];
+        // Per address, not over the joined string: each may carry its own
+        // display name, and the encoding has to stop at the angle brackets.
+        joined += smartbotic::common::MimeWords::encodeAddress(addresses[i]);
     }
     return joined;
 }
@@ -72,39 +90,51 @@ std::string SmtpClient::buildMime(const Message& message, const std::string& mes
     if (credentials_.from_name.empty()) {
         mime << "From: " << sender << "\r\n";
     } else {
-        mime << "From: \"" << credentials_.from_name << "\" <" << sender << ">\r\n";
+        mime << "From: "
+             << smartbotic::common::MimeWords::encodeAddress(
+                    credentials_.from_name + " <" + sender + ">")
+             << "\r\n";
     }
     mime << "To: " << joinAddresses(message.to) << "\r\n";
     if (!message.cc.empty()) {
         mime << "Cc: " << joinAddresses(message.cc) << "\r\n";
     }
     if (!message.reply_to.empty()) {
-        mime << "Reply-To: " << message.reply_to << "\r\n";
+        mime << "Reply-To: " << smartbotic::common::MimeWords::encodeAddress(message.reply_to) << "\r\n";
     }
-    mime << "Subject: " << message.subject << "\r\n";
+    mime << "Subject: " << smartbotic::common::MimeWords::encode(message.subject) << "\r\n";
     mime << "Message-ID: <" << message_id << ">\r\n";
     mime << "MIME-Version: 1.0\r\n";
 
     const std::string body_type = message.html ? "text/html" : "text/plain";
 
     if (message.attachments.empty()) {
-        mime << "Content-Type: " << body_type << "; charset=UTF-8\r\n\r\n";
-        mime << message.body << "\r\n";
+        mime << "Content-Type: " << body_type << "; charset=UTF-8\r\n";
+        // Declared, not assumed. A UTF-8 body sent with no transfer encoding is
+        // 8-bit data in a channel defined as 7-bit; base64 is the form every
+        // server carries unchanged, whether or not it announces 8BITMIME.
+        mime << "Content-Transfer-Encoding: base64\r\n\r\n";
+        mime << wrapBase64(smartbotic::common::StringUtils::base64Encode(message.body));
         return mime.str();
     }
 
     mime << "Content-Type: multipart/mixed; boundary=\"" << boundary << "\"\r\n\r\n";
     mime << "--" << boundary << "\r\n";
-    mime << "Content-Type: " << body_type << "; charset=UTF-8\r\n\r\n";
-    mime << message.body << "\r\n";
+    mime << "Content-Type: " << body_type << "; charset=UTF-8\r\n";
+    mime << "Content-Transfer-Encoding: base64\r\n\r\n";
+    mime << wrapBase64(smartbotic::common::StringUtils::base64Encode(message.body));
 
     for (const auto& attachment : message.attachments) {
         const std::string type = attachment.mime_type.empty() ? "application/octet-stream"
                                                               : attachment.mime_type;
         mime << "--" << boundary << "\r\n";
-        mime << "Content-Type: " << type << "; name=\"" << attachment.filename << "\"\r\n";
+        mime << "Content-Type: " << type << "; "
+             << smartbotic::common::MimeWords::encodeParameter("name", attachment.filename)
+             << "\r\n";
         mime << "Content-Transfer-Encoding: base64\r\n";
-        mime << "Content-Disposition: attachment; filename=\"" << attachment.filename << "\"\r\n\r\n";
+        mime << "Content-Disposition: attachment; "
+             << smartbotic::common::MimeWords::encodeParameter("filename", attachment.filename)
+             << "\r\n\r\n";
 
         // Already base64 from the caller; wrapped so no line exceeds what SMTP allows.
         const std::string& encoded = attachment.content_base64;

+ 103 - 0
tests/cpp/mime_words_test.cpp

@@ -0,0 +1,103 @@
+// RFC 2047 encoding and decoding, both directions.
+//
+// Pure logic with no I/O, so it is tested here rather than through a workflow.
+// Run with scripts/run-cpp-tests.sh.
+
+#include "common/mime_words.hpp"
+
+#include <iostream>
+#include <string>
+#include <vector>
+
+using smartbotic::common::MimeWords;
+
+namespace {
+
+int failures = 0;
+
+void check(const std::string& what, const std::string& got, const std::string& want) {
+    const bool ok = got == want;
+    if (!ok) ++failures;
+    std::cout << (ok ? "  ok   " : "  FAIL ") << what << "\n";
+    if (!ok) {
+        std::cout << "         got:  " << got << "\n";
+        std::cout << "         want: " << want << "\n";
+    }
+}
+
+} // namespace
+
+int main() {
+    // --- decoding, which is what arrives from a mail server ---
+
+    // The shape that started this: a display name in front of an address. Only
+    // the name is encoded; the address must come through untouched.
+    check("decode: name in front of an address",
+          MimeWords::decode("=?UTF-8?Q?Szont=C3=A1gh_Ferenc?= <ferenc.szontagh@smartbotics.ai>"),
+          "Szontágh Ferenc <ferenc.szontagh@smartbotics.ai>");
+
+    // A value too long for one word is split, and the whitespace between the
+    // words is a separator - it must not survive into the decoded text.
+    check("decode: two words rejoined without the fold space",
+          MimeWords::decode("=?UTF-8?Q?S=C3=A1rk=C3=A1nyok_p=C3=A1rz=C3=A1si_=C3=A9s_=C3=A9tkez?="
+                            " =?UTF-8?Q?=C3=A9si_szok=C3=A1saik?="),
+          "Sárkányok párzási és étkezési szokásaik");
+
+    check("decode: B encoding", MimeWords::decode("=?UTF-8?B?w6FydsOtenTFsQ==?="), "árvíztű");
+
+    // An underscore in a Q word is a space, not an underscore.
+    check("decode: underscore is a space", MimeWords::decode("=?UTF-8?Q?a_b?="), "a b");
+
+    check("decode: plain text is untouched", MimeWords::decode("Hello there"), "Hello there");
+
+    // Gibberish is better than nothing: an unknown charset or a malformed word
+    // comes back as it arrived rather than being dropped.
+    check("decode: unknown charset passes through",
+          MimeWords::decode("=?SHIFT_JIS?Q?abc?="), "abc");
+    check("decode: malformed word is left alone",
+          MimeWords::decode("=?UTF-8?Q?unterminated"), "=?UTF-8?Q?unterminated");
+
+    // --- encoding, which is what we put on the wire ---
+
+    check("encode: ascii is left readable", MimeWords::encode("Plain subject"), "Plain subject");
+    check("encodeAddress: a bare address is untouched",
+          MimeWords::encodeAddress("a@b.com"), "a@b.com");
+    check("encodeAddress: an ascii name needs nothing",
+          MimeWords::encodeAddress("John Doe <a@b.com>"), "John Doe <a@b.com>");
+
+    // A comma in a display name ends the address unless it is quoted, which
+    // would silently turn one recipient into two.
+    check("encodeAddress: a comma forces quoting",
+          MimeWords::encodeAddress("Doe, John <a@b.com>"), "\"Doe, John\" <a@b.com>");
+
+    // Already-quoted names must not end up double quoted, nor with the quotes
+    // inside the encoded word.
+    check("encodeAddress: existing quotes are not doubled",
+          MimeWords::encodeAddress("\"John Doe\" <a@b.com>"), "John Doe <a@b.com>");
+
+    // The round trip is the property that matters most: whatever we send, a
+    // conforming reader gets back what we meant.
+    for (const std::string& original : {std::string("Szontágh Ferenc"),
+                                        std::string("Sárkányok párzási és étkezési szokásaik"),
+                                        std::string("árvíztűrő tükörfúrógép"),
+                                        std::string("plain ascii")}) {
+        check("round trip: " + original, MimeWords::decode(MimeWords::encode(original)), original);
+    }
+
+    // A non-ASCII display name encodes to a word and decodes back to itself,
+    // with the address carried through unchanged.
+    check("round trip: name and address",
+          MimeWords::decode(MimeWords::encodeAddress("Szontágh Ferenc <a@b.com>")),
+          "Szontágh Ferenc <a@b.com>");
+
+    // A filename gets the RFC 2231 form plus an ASCII fallback.
+    const std::string param = MimeWords::encodeParameter("filename", "Sárkányok.pdf");
+    check("encodeParameter: ascii is plain",
+          MimeWords::encodeParameter("filename", "report.pdf"), "filename=\"report.pdf\"");
+    std::cout << (param.find("filename*=UTF-8''") != std::string::npos ? "  ok   " : "  FAIL ")
+              << "encodeParameter: non-ascii uses RFC 2231\n";
+    if (param.find("filename*=UTF-8''") == std::string::npos) { ++failures; std::cout << "         got: " << param << "\n"; }
+
+    std::cout << (failures ? "FAILED " + std::to_string(failures) + " check(s)\n" : "all passed\n");
+    return failures ? 1 : 0;
+}