feat(eth): port mesh/http to RP2350 + W5500 (HTTP/80 + HTTPS/443) (#10573)

* add watchdog to rp2xx0

* feat(eth-api): phase 0 skeleton — HTTP API server on TCP/80 (503 only)

First step of porting mesh/http/ to RP2350 + W5500 (today ESP32-only).
Phase 0 stands up the listener over the existing arduino-libraries/Ethernet
stack (no Mongoose — its built-in TCP/IP would conflict with EthernetServer
and break OTA/MQTT/NTP). Skeleton parses request line, logs it, replies 503.
Real handlers (/api/v1/{info,fromradio,toradio}) come in later phases.

- New mesh/eth/ethApiServer.{h,cpp} gated by HAS_ETHERNET_API
- Wired into ethClient init+loop in parallel with existing ethHttpOTA (port 4244)
- Enabled in both wiznet_5500_evb_pico2_e22p and pico2_w5500_e22 variants
- Build impact on wiznet: +~9 KB flash (63.0% used), <1 KB RAM

* feat(eth-api): phase 1 — implement /api/v1/{fromradio,toradio} with PhoneAPI

Replace phase 0 skeleton with the real handlers that bridge the HTTP transport
to PhoneAPI, mirroring mesh/http/ContentHandler (ESP32) semantics.

- EthHttpAPI : public PhoneAPI (api_type=TYPE_HTTP, checkIsConnected=true),
  outside the MESHTASTIC_EXCLUDE_WEBSERVER gate so it builds on RP2350.
- Minimal HTTP parser: method/path/query/Content-Length, no allocations beyond
  Arduino String, capped at 32 header lines x 256 bytes (anti-DoS).
- OPTIONS preflight -> 204 with CORS + X-Protobuf-Schema.
- GET /api/v1/fromradio[?all=true]: stream-write up to 64 protobufs from
  webAPI.getFromRadio(), Connection: close framing (HTTP/1.0 style).
- PUT /api/v1/toradio: read Content-Length bytes (<=512), call
  webAPI.handleToRadio(), echo body back.
- 404 default for unknown paths, 405 for wrong method, 400 for bad body.

Validated e2e against wiznet @ 192.168.1.143:
- OPTIONS /api/v1/fromradio -> 204 + correct CORS headers
- PUT ToRadio{want_config_id=1} (2 bytes) -> 200 + echo
- GET /api/v1/fromradio?all=true -> 200 + 3476 bytes (config+channels+nodeinfo)
- GET /api/v1/nonexistent -> 404 unknown endpoint
- OTA HTTP on port 4244 untouched and still responsive

Build impact on wiznet_5500_evb_pico2_e22p: +2 KB flash (63.1%), +1.8 KB RAM
(17.6%) vs phase 0.

* fix(eth-api): move accept loop to dedicated OSThread (20ms tick)

The API server was being polled from the Ethernet client Periodic which runs
every 5s. Measured impact before this fix on wiznet @ 192.168.1.143:

  req 1: ttfb=6.458s total=6.509s    (matches the 5s tick + handler overhead)
  req 2: connection refused          (W5500 sockets exhausted)
  req 3: connection refused

After: same hardware, same network, back-to-back requests:

  req 2: ttfb=0.287s
  req 3: ttfb=0.015s
  req 4: ttfb=0.022s
  req 5: ttfb=0.020s

The web client (meshtastic/web served from localhost) was visibly stalling
mid-handshake — it had pulled 109 nodes + 19 messages but channels and device
info never arrived. With the periodic-only polling, every request takes 5s
and the W5500's 4 hardware sockets fill up under the burst.

EthApiServerThread mirrors the WebServerThread pattern from ESP32
mesh/http/WebServer.cpp: adaptive interval — 20ms when there's recent
traffic, 100ms after 5s of idle, 500ms after 30s. Auto-registers with the
OSThread scheduler on construction.

ethHttpOTA still ticks from the periodic; left unchanged because OTA is one
large transfer that tolerates the latency, and minimizing scope here to one
behaviour change at a time.

* refactor(eth-api): extract handlers to shared module via IStreamReadWrite

Phase 2.0 of the TLS port — separate transport from request handling so
the upcoming HTTPS server can reuse the exact same parser + routing logic.

- New ethApiHandlers.{h,cpp}: Request, parser, CORS helpers, fromradio/toradio
  handlers, EthHttpAPI : PhoneAPI subclass. All driven by a single
  IStreamReadWrite interface that inherits Print (so client.print(...) keeps
  working transparently).
- ethApiServer.cpp slimmed to ~90 LOC: now just the EthernetServer(80) +
  OSThread + an EthernetClientStream adapter that forwards reads/writes to
  the underlying EthernetClient. No behaviour change.

Validated on wiznet @ 192.168.1.143: 5 curl requests TTFB 18-110ms (same
as pre-refactor), protobuf round-trip PUT 200 + GET ?all=true 3499 bytes.
Flash impact: +200 bytes (63.2%); RAM unchanged (17.6%).

* feat(eth-tls): ECDSA P-256 self-signed cert generation for HTTPS API

mbedTLS-based cert generation module that produces a SAN=IP self-signed
ECDSA P-256 server certificate, persisted under LittleFS so subsequent
boots reuse the same key. Generation runs once on a dedicated OSThread
so the ECDSA keygen path (~430 ms) never blocks the Periodic stack or
the Ethernet reconnect loop.

  - mbedTLS 3.6.2 sources compiled in via scripts/add_mbedtls_sources.py
    (BuildSources of pico-sdk/lib/mbedtls/library/*.c, all 108 files).
    MBEDTLS_USER_CONFIG_FILE injected as CPPDEFINES tuple in the script —
    build_flags shell escape mangles the embedded quotes on Windows.
  - src/mbedtls_user_config.h undefs MBEDTLS_HAVE_TIME, HAVE_TIME_DATE,
    TIMING_C, NET_C, FS_IO, PSA_ITS_FILE_C, PSA_CRYPTO_STORAGE_C, and
    defines MBEDTLS_NO_PLATFORM_ENTROPY (entropy_poll.c uses a raw
    platform check, not gated by a flag).
  - src/mesh/eth/ethCert.{h,cpp}: ECDSA P-256 keypair + cert with
    SAN(IP=current) using pico-sdk get_rand_64() directly as f_rng,
    bypassing mbedtls_entropy. DER buffers heap-allocated to keep the
    OSThread stack within budget. Cert + key + ip persisted under
    /prefs/eth_*.der so subsequent boots reuse the same identity.
  - Gated by HAS_ETHERNET_TLS_API. Standalone phase: generation only;
    HTTPS server (TCP/443) wired up in the follow-up commit.

* feat(eth-tls-api): phase 2.2 — HTTPS server on TCP/443 reusing handlers

Brings up an mbedTLS server on port 443 that defers to the ECDSA P-256
self-signed cert produced by [[ethCert]]. Reuses the HTTP request /
response handlers from [[ethApiHandlers]] via the IStreamReadWrite
interface — there is exactly one code path for routing / CORS / PhoneAPI
integration regardless of whether the transport is plain TCP or TLS.

EthTlsApiServerThread (OSThread):
  - Phase A: poll isEthCertReady() every 500 ms while the cert worker
    runs. Once true, parse cert chain + key, build ssl_config (TLS server,
    stream transport, default preset, VERIFY_NONE since we are the cert
    issuer), install own cert, run ssl_setup, bind tlsServer on 443.
  - Phase B: standard adaptive accept loop (20 / 100 / 500 ms tick),
    identical to the plain-HTTP server.

Per-connection flow:
  1. session_reset on the static ssl context (1 in-flight session — multi-
     session pool is Phase 3 if needed)
  2. set_bio routes mbedtls I/O to two C callbacks (netSend / netRecv)
     that bridge to the live EthernetClient via the void* ctx
  3. handshake loop (sync, blocking, with 10 s recv timeout) — mbedTLS
     errors are logged with mbedtls_strerror
  4. wrap (ssl, client) in MbedTlsStream → handleApiClient(stream)
  5. close_notify + client.stop

Stack budget continues the Phase 2.1-bis discipline: every large buffer
(ssl context, cert chain, pk_key, ssl_config) lives in BSS as a static
global. The OSThread stack only holds the per-connection EthernetClient
adapter and small mbedtls return codes.

Validated on wiznet 192.168.1.143:
  - cert load-from-FS path: 13 ms (regen skipped on second boot)
  - cert gen path: 290 ms (first boot only)
  - ssl_setup chain: parse cert, parse key, ssl_config_defaults,
    conf_own_cert, ssl_setup all return 0
  - end-to-end: curl -k https://192.168.1.143/api/v1/fromradio → 200 OK
    with application/x-protobuf, CORS, X-Protobuf-Schema headers, server
    initiates close_notify cleanly

Footprint vs 2.0:
  - flash 70.5% → 87.3% (+220 KB for mbedtls_ssl + x509 server-side code)
  - RAM static 17.6% → 19.8% (+11 KB BSS for ssl_context + cert chain)
  - heap at runtime adds ~32 KB per active session (mbedtls in/out
    record buffers, default MBEDTLS_SSL_*_CONTENT_LEN=16384)

Next: validate from Firefox direct (https://192.168.1.143 → warning
accept → JSON visible), then from client.meshtastic.org hosted to
confirm the mixed-content block is gone.

* fix(eth-tls): add KeyUsage + EKU serverAuth and cap TLS 1.2 for browser compat

Two browser-compat fixes that surfaced in Firefox validation:

1. Cert v1 only had Basic Constraints + SAN. NSS / Firefox refuse to
   treat a cert as a TLS server cert without an Extended Key Usage
   extension naming id-kp-serverAuth since 2023 — the error surfaces as
   a non-overridable 'Secure Connection Failed' with no 'Accept the
   Risk' path. Add KeyUsage(digitalSignature + keyEncipherment, critical)
   and ExtendedKeyUsage(serverAuth, critical). Bump cert/key/ip file
   paths to '_v2' so live boards regenerate on next start instead of
   loading a v1 cert that the browser silently refuses.

2. pico-sdk mbedtls defines MBEDTLS_SSL_PROTO_TLS1_3 in its default
   config, but the server-side 1.3 plumbing in this vendored build is
   incomplete: Firefox and openssl-3.5's s_client default to 1.3 and
   the handshake dies 4 ms in with MBEDTLS_ERR_ERROR_GENERIC_ERROR
   (-0x0001). curl/SChannel happened to default to 1.2 so it masked the
   issue earlier. Cap min/max_tls_version to TLS 1.2 — clients downgrade
   transparently and we keep the ECDHE-ECDSA + AES-GCM / CHACHA20-POLY1305
   suites that already work end-to-end.

Validated on wiznet 192.168.1.143:
  - cert v2 dumps with the 4 extensions visible (openssl x509 -text):
    BasicConstraints CA:FALSE, KeyUsage(critical) digitalSignature +
    keyEncipherment, ExtendedKeyUsage(critical) serverAuth, SAN IP.
  - openssl s_client (no version flag): downgrades to TLS 1.2, handshake
    completes, verify_return=18 (self-signed, expected).
  - Firefox: warning self-signed -> Advanced -> Accept Risk -> handshake
    OK in 632 ms with CHACHA20-POLY1305, request reaches handleApiClient
    and returns the protobuf body.

* perf(eth-api): HTTP/1.1 keep-alive — one handshake per session, not per request

Before this change /fromradio used 'Connection: close' framing (HTTP/1.0
style with no Content-Length), forcing each poll to redo the full TLS
handshake. client.meshtastic.org needs ~80 sequential /fromradio polls
during initial sync (one per packet from the config-replay state
machine: MyInfo, channels, every Config_*, every ModuleConfig_*, every
NodeInfo), so the user-visible load time was ~50 s of pure handshake
overhead (80 requests * ~625 ms ECDSA P-256 each).

Three coordinated changes:

1. handleApiClient() now loops on the same connection until the peer
   closes or parseRequest hits its 3 s idle timeout. requestsServed
   counter keeps the 'bad/timeout request' debug log from firing on
   the natural idle close after a keep-alive sequence.

2. handleFromRadio() buffers all packets into a std::vector before
   writing, so it can emit a real Content-Length and 'Connection:
   keep-alive'. Buffer is dynamic — common 1-packet response only
   allocates ~256 B; ?all=true keeps the 64-packet cap which tops out
   around 16 KB. handleToRadio + sendPreflight switched to keep-alive
   too (they already had real Content-Length). sendError keeps close —
   errors are terminal.

3. ethTlsApiServer netRecv RECV_TIMEOUT_MS dropped from 10 s to 3 s so
   mbedtls_ssl_read can't outlast the handler's idle deadline (a quiet
   browser leaving the socket open would otherwise wedge the OSThread
   for 10 s past the natural close).

Measured on wiznet 192.168.1.143 against client.meshtastic.org:
  - one TLS handshake (641 ms, CHACHA20-POLY1305)
  - ~80 requests pipelined over the same session
  - full config + 40 NodeInfos in ~6-8 s (vs ~50 s before)
  - per-request latency post-handshake: ~10-25 ms
  - curl -kv with two URLs: 'Reusing existing https: connection',
    server returns 'Connection: keep-alive' explicitly.

* fix(eth-tls): watchdog + Chrome compat — handshake stability across browsers

Phase 3 keep-alive shipped a working Firefox flow but client.meshtastic.org
loops + Chrome 'Test Connection' both rebooted the board. Four distinct
issues; collectively they kept the OSThread inside serveClient() too long
or spin-looping without yielding, and pico-sdk mbedtls' TLS 1.3 code path
choked on Chrome's modern ClientHello.

1. Watchdog reset during keep-alive idle. Once a client drained the
   replay queue and entered 3 s poll mode, the OSThread sat inside
   netRecv()'s busy-wait waiting for the next request. Two consecutive
   3 s waits plus prior handler time crossed the 8 s RP2350 hardware
   watchdog. Pet the watchdog inside netRecv()'s poll loop (every 2 ms)
   so a quiet client can never starve the watchdog. Same fix in the
   ethApiHandlers per-request yield path.

2. Cap session at 64 requests + yield() between. Defense-in-depth: a
   pathological client can't monopolize serveClient indefinitely; after
   the cap it just re-handshakes (~625 ms), still vastly cheaper than
   the per-request handshake we had before keep-alive.

3. TLS 1.3 code compiled out of mbedtls entirely
   (#undef MBEDTLS_SSL_PROTO_TLS1_3 in mbedtls_user_config). Capping
   max_tls_version=TLS1_2 at runtime is enough for Firefox / openssl
   (they downgrade cleanly), but Chrome's ClientHello carries TLS 1.3
   extensions — post-quantum key shares, Encrypted ClientHello,
   etc. — that the vendored mbedtls 1.3 parser crashes on before the
   downgrade decision happens. Removing the 1.3 sources sidesteps the
   parser; ServerHello just announces TLS 1.2 and Chrome accepts.

4. netSend infinite WANT_WRITE spin. When W5500's TX buffer momentarily
   filled mid-handshake (Chrome draining slower than Firefox during
   ServerKeyExchange), EthernetClient::write() returned 0, our netSend
   returned MBEDTLS_ERR_SSL_WANT_WRITE without delay, mbedtls retried
   immediately, repeat at ~180k iter/sec until ... well, until the
   board's other threads got nothing done. Log signature: ret=-0x6880
   tight-looping in the handshake iter trace. Rewrite netSend to block
   with delay(2) + watchdog_update() and a 3 s timeout — same shape as
   netRecv. Return MBEDTLS_ERR_SSL_PEER_CLOSE_NOTIFY on disconnect
   (was incorrectly returning WANT_WRITE).

Also added granular per-iter handshake logging gated on first 20 iters
+ every 50th after that, so any future regression localizes itself in
COM9 without RTT JLink.

Validated on wiznet 192.168.1.143:
  - Firefox: client.meshtastic.org full sync + idle poll stable (no
    reset during the previously-crashing 'replay drain complete' phase)
  - Chrome: 'Test Connection' accepts the cert prompt and connects
  - Edge: same as Chrome
  - openssl s_client default + tls1_2 forced: both negotiate TLS 1.2
    with ECDHE-ECDSA + AES-GCM / CHACHA20-POLY1305, verify=18 (self-
    signed, expected)

* chore(eth-tls): drop verbose debug logs from cert + TLS init paths

The cert pipeline + TLS context init had step-by-step LOG_INFOs and
ubiquitous Serial.flush() that were essential while diagnosing the
Phase 2.1-bis stack overflow, the Chrome handshake crash, and the
keep-alive watchdog reset. Once those bugs were fixed the logs just
clutter COM9 on every boot.

Kept on hand:
  - cert: 'loaded from FS', 'generating ...', 'generated N B in T ms',
    'persisted to LittleFS', plus all LOG_ERROR / LOG_WARN paths
  - tls: 'server listening on TCP port 443', 'client connected from',
    'handshake OK in N ms ciphersuite=', 'handshake failed -0xXXXX (...)',
    plus all init LOG_ERRORs

Dropped:
  - cert: 'step 1/8 pk_setup' through 'step 8/8 copy key DER', 'thread
    woke', 'pipeline OK' (now silent on success)
  - tls: 'cert is ready, initializing', 'parsing cert chain', 'parsing
    key', 'ssl_config_defaults', 'conf_own_cert', 'ssl_setup',
    'server worker scheduled', and the per-iter handshake trace
  - obsolete 'Optional mbedtls debug bridge' commented-out stub
  - all Serial.flush() that were added defensively for the
    debug-the-crash phase

Bin shrinks ~40 KB (logs + format strings). Validated on wiznet
192.168.1.143: HTTP 200 round-trip works post-flash, no regression.

* fix(eth): harden API/TLS build guards (review #10573)

- ethApiHandlers.{h,cpp}: gate on (HAS_ETHERNET_API || HAS_ETHERNET_TLS_API)
  so the shared handleApiClient/IStreamReadWrite still compile for an
  HTTPS-only variant (the TLS server depends on them).
- ethTlsApiServer.{cpp,h}, ethCert.{cpp,h}: require ARCH_RP2040 (they use Pico
  SDK get_rand_64()/watchdog_update()); the ethClient.cpp TLS include + call
  site are tightened to match, so enabling the flag on a non-RP2040 Ethernet
  target is a clean no-op instead of a cryptic build break.
- clang-format (16) the touched eth files to satisfy Trunk Check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(eth): address review — listener/cert recovery + keep-alive framing

- Rebind HTTP/80 and HTTPS/443 after a W5500 chip reset. reconnectETH() now
  tears the API/TLS servers down (deInitEthApiServer/deInitEthTlsApiServer) so
  the restart path recreates them; the singleton guards previously left both
  bound to dead sockets until reboot. The workers are kept (only the listener
  is dropped) so nothing is deleted from another thread's runOnce.

- Stop the keep-alive loop after a handler error. /fromradio and /toradio
  return whether the connection may stay open; a 400/405/408 response
  (Connection: close) now ends the loop so unread body bytes can't be parsed
  as the next request.

- Validate the cached cert/key before trusting it (parse both DERs), and clear
  the IP commit-marker before rewriting the pair and write it last, so a reset
  mid-persist regenerates instead of handing the TLS server a truncated or
  mismatched pair (which would hard-disable HTTPS).

- Track a cert generation counter so the TLS server reloads and rebinds when a
  DHCP lease change regenerates the cert for a new IP (the SAN must follow, or
  browsers reject the new address). The cert worker is no longer one-shot.

- Silence the SCons F821/E402 lint on add_mbedtls_sources.py to match the other
  extra_scripts.

Validated on a W5500 board: a forced chip reset rebinds 80 and 443 without an
MCU reboot; a truncated cached cert regenerates and HTTPS recovers; a simulated
lease change reissues the cert with the new IP in the SAN.

* fix(eth): clear cert readiness on regen failure + verify cached pair

Follow-up to the review:
- Reset ready_/certIp_ when ensureCertForIp() fails, so isReady() no longer
  reports true with empty material — otherwise a later TLS reload fails
  initTlsContext() and stays disabled instead of retrying on the next poll.
- certKeyParse() now also checks the cached cert and key form a matching pair
  (mbedtls_pk_check_pair), rejecting an independently-parseable but mismatched
  cert/key (e.g. if the IP commit-marker survived a failed clear).

Validated on a W5500 board: a valid cached pair still loads from FS, and a
regenerated cert reloads cleanly.

---------

Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Carlos Valdes
2026-07-01 05:51:14 -05:00
committed by GitHub
co-authored by GitHub Ben Meadors Claude Opus 4.8
parent cfc2e457a9
commit e64d20548c
13 changed files with 1470 additions and 1 deletions
+351
View File
@@ -0,0 +1,351 @@
#include "configuration.h"
#if HAS_ETHERNET && defined(HAS_ETHERNET_TLS_API) && defined(ARCH_RP2040)
#include "concurrency/OSThread.h"
#include "ethApiHandlers.h"
#include "ethCert.h"
#include "ethTlsApiServer.h"
#include <Arduino.h>
#ifdef USE_ARDUINO_ETHERNET
#include <Ethernet.h>
#else
#include <RAK13800_W5100S.h>
#endif
#include <mbedtls/error.h>
#include <mbedtls/pk.h>
#include <mbedtls/ssl.h>
#include <mbedtls/x509_crt.h>
#include <hardware/watchdog.h>
#include <pico/rand.h>
#ifndef ETH_TLS_API_PORT
#define ETH_TLS_API_PORT 443
#endif
// Adaptive poll intervals (mirror ethApiServer.cpp).
static constexpr uint32_t ACTIVE_THRESHOLD_MS = 5000;
static constexpr uint32_t MEDIUM_THRESHOLD_MS = 30000;
static constexpr int32_t ACTIVE_INTERVAL_MS = 20;
static constexpr int32_t MEDIUM_INTERVAL_MS = 100;
static constexpr int32_t IDLE_INTERVAL_MS = 500;
// Matches the keep-alive idle window in ethApiHandlers — if the handler
// loop calls read() and netRecv blocked for 10 s, the 3 s idle deadline
// inside parseRequest would be irrelevant and the OSThread would stay stuck
// long after a quiet browser closed its end of the TCP socket.
static constexpr uint32_t RECV_TIMEOUT_MS = 3000;
// Reuse the picoRand callback semantics from ethCert.cpp. Local copy so we
// don't have to expose it through a header; the cost is two trivial functions.
static int picoRand(void * /*ctx*/, unsigned char *out, size_t len)
{
while (len > 0) {
uint64_t r = get_rand_64();
size_t to_copy = len > sizeof(r) ? sizeof(r) : len;
memcpy(out, &r, to_copy);
out += to_copy;
len -= to_copy;
}
return 0;
}
// One-shot TLS context lives in BSS — keeps mbedtls allocations off the
// OSThread stack (lesson from Phase 2.1-bis: stack budget is tight on M33).
static EthernetServer *tlsServer = nullptr;
static mbedtls_x509_crt certChain;
static mbedtls_pk_context pkKey;
static mbedtls_ssl_config sslConf;
static mbedtls_ssl_context ssl;
static bool tlsReady = false;
// Adapter: route mbedtls_ssl_set_bio() through the EthernetClient instance
// that runOnce() is currently servicing. The void* ctx we hand mbedtls is a
// pointer to the EthernetClient.
static int netSend(void *ctx, const unsigned char *buf, size_t len)
{
auto *client = static_cast<EthernetClient *>(ctx);
if (!client->connected())
return MBEDTLS_ERR_SSL_PEER_CLOSE_NOTIFY;
// Block-with-yield until the W5500 TX buffer can absorb the chunk.
// Returning WANT_WRITE without delay made mbedtls_ssl_handshake() spin
// at ~180k iter/s when Chrome was slow to drain the socket during the
// ECDHE-ECDSA ServerKeyExchange — the original code logged exactly that
// signature (ret=-0x6880 / WANT_WRITE) tight-looping forever. Firefox
// happened to read fast enough that the buffer never filled.
uint32_t t0 = millis();
while (true) {
if (!client->connected())
return MBEDTLS_ERR_SSL_PEER_CLOSE_NOTIFY;
size_t w = client->write(buf, len);
if (w > 0)
return (int)w;
if (millis() - t0 > RECV_TIMEOUT_MS)
return MBEDTLS_ERR_SSL_TIMEOUT;
watchdog_update();
delay(2);
}
}
static int netRecv(void *ctx, unsigned char *buf, size_t len)
{
auto *client = static_cast<EthernetClient *>(ctx);
// Block-with-timeout: spin until bytes arrive, the peer closes, or we
// exceed the per-recv budget. Pure non-blocking (return WANT_READ) would
// require mbedtls_ssl_handshake to be driven from the runOnce dispatcher
// — overkill for the Phase 2.2 skeleton with a single in-flight session.
//
// Pet the 8 s hardware watchdog from inside the poll loop. We sit here
// for up to RECV_TIMEOUT_MS waiting for the next keep-alive request, and
// a quiet client can string two such waits back-to-back (6 s) plus the
// earlier handshake/handler time — easily past the watchdog deadline.
// The main loop()'s watchdog_update() never runs while the OSThread is
// inside serveClient(), so it has to be done here.
uint32_t t0 = millis();
while (client->available() == 0) {
if (!client->connected())
return MBEDTLS_ERR_SSL_PEER_CLOSE_NOTIFY;
if (millis() - t0 > RECV_TIMEOUT_MS)
return MBEDTLS_ERR_SSL_TIMEOUT;
watchdog_update();
delay(2);
}
int n = client->read(buf, len);
if (n <= 0)
return MBEDTLS_ERR_SSL_WANT_READ;
return n;
}
// Bridges mbedtls_ssl_read/write to the request handlers via the same
// IStreamReadWrite interface the plain-HTTP server uses.
class MbedTlsStream : public IStreamReadWrite
{
public:
MbedTlsStream(mbedtls_ssl_context *s, EthernetClient *c) : ssl_(s), client_(c) {}
size_t write(uint8_t b) override
{
int r = mbedtls_ssl_write(ssl_, &b, 1);
return r > 0 ? 1 : 0;
}
size_t write(const uint8_t *buf, size_t len) override
{
size_t total = 0;
while (total < len) {
int r = mbedtls_ssl_write(ssl_, buf + total, len - total);
if (r == MBEDTLS_ERR_SSL_WANT_WRITE || r == MBEDTLS_ERR_SSL_WANT_READ)
continue;
if (r <= 0)
break;
total += (size_t)r;
}
return total;
}
int available() override
{
// mbedtls buffers internally; ssl_get_bytes_avail reports what is
// already decoded. If 0, peek a record via read into a 1-byte buffer.
size_t pending = mbedtls_ssl_get_bytes_avail(ssl_);
if (pending > 0)
return (int)pending;
// Best-effort: report network bytes (rough proxy — handlers usually
// call read() in a loop and tolerate slow streams).
return client_->available();
}
int read() override
{
uint8_t b;
int r = mbedtls_ssl_read(ssl_, &b, 1);
return r == 1 ? (int)b : -1;
}
int read(uint8_t *buf, size_t len) override
{
int r = mbedtls_ssl_read(ssl_, buf, len);
return r > 0 ? r : -1;
}
bool connected() override { return client_->connected(); }
void flush() override { client_->flush(); }
IPAddress remoteIP() override { return client_->remoteIP(); }
private:
mbedtls_ssl_context *ssl_;
EthernetClient *client_;
};
class EthTlsApiServerThread : public concurrency::OSThread
{
public:
EthTlsApiServerThread() : concurrency::OSThread("EthTlsApi") { lastActivityMs = millis(); }
protected:
int32_t runOnce() override
{
// Phase A: wait for the cert worker, then rebuild if the cert was
// regenerated (a DHCP lease change to a new IP bumps the generation, and
// the SAN must follow or browsers reject the new address).
if (tlsReady && getEthCertGeneration() != loadedCertGen_) {
LOG_INFO("ETH TLS: cert regenerated (gen %u->%u), reloading + rebinding", (unsigned)loadedCertGen_,
(unsigned)getEthCertGeneration());
deInitEthTlsApiServer(); // frees ctx, drops the listener, clears tlsReady
}
if (!tlsReady) {
if (!isEthCertReady())
return 500;
if (!initTlsContext())
return INT32_MAX; // hard fail — TLS server stays disabled
loadedCertGen_ = getEthCertGeneration();
tlsReady = true;
}
// Phase B: accept + serve one client.
if (!tlsServer)
return INT32_MAX;
EthernetClient client = tlsServer->accept();
if (client) {
lastActivityMs = millis();
serveClient(client);
client.stop();
}
uint32_t since = millis() - lastActivityMs;
if (since < ACTIVE_THRESHOLD_MS)
return ACTIVE_INTERVAL_MS;
if (since < MEDIUM_THRESHOLD_MS)
return MEDIUM_INTERVAL_MS;
return IDLE_INTERVAL_MS;
}
private:
uint32_t lastActivityMs;
uint32_t loadedCertGen_ = 0; // cert generation the current TLS context was built from
bool initTlsContext()
{
const EthCertMaterial &cert = getEthCert();
if (cert.certDer.empty() || cert.keyDer.empty()) {
LOG_ERROR("ETH TLS: cert material is empty, refusing to start TLS server");
return false;
}
mbedtls_x509_crt_init(&certChain);
mbedtls_pk_init(&pkKey);
mbedtls_ssl_config_init(&sslConf);
mbedtls_ssl_init(&ssl);
int ret;
ret = mbedtls_x509_crt_parse_der(&certChain, cert.certDer.data(), cert.certDer.size());
if (ret != 0) {
LOG_ERROR("ETH TLS: x509_crt_parse_der failed -0x%04x", -ret);
return false;
}
ret = mbedtls_pk_parse_key(&pkKey, cert.keyDer.data(), cert.keyDer.size(), nullptr, 0, picoRand, nullptr);
if (ret != 0) {
LOG_ERROR("ETH TLS: pk_parse_key failed -0x%04x", -ret);
return false;
}
ret = mbedtls_ssl_config_defaults(&sslConf, MBEDTLS_SSL_IS_SERVER, MBEDTLS_SSL_TRANSPORT_STREAM,
MBEDTLS_SSL_PRESET_DEFAULT);
if (ret != 0) {
LOG_ERROR("ETH TLS: ssl_config_defaults failed -0x%04x", -ret);
return false;
}
mbedtls_ssl_conf_rng(&sslConf, picoRand, nullptr);
mbedtls_ssl_conf_authmode(&sslConf, MBEDTLS_SSL_VERIFY_NONE);
// TLS 1.3 is compiled out via mbedtls_user_config.h (it would crash
// on Chrome's ClientHello extensions). These calls pin runtime to
// 1.2 too as a defense-in-depth: if a future user config flip
// re-enables 1.3 code, the config layer still won't negotiate it.
mbedtls_ssl_conf_max_tls_version(&sslConf, MBEDTLS_SSL_VERSION_TLS1_2);
mbedtls_ssl_conf_min_tls_version(&sslConf, MBEDTLS_SSL_VERSION_TLS1_2);
ret = mbedtls_ssl_conf_own_cert(&sslConf, &certChain, &pkKey);
if (ret != 0) {
LOG_ERROR("ETH TLS: conf_own_cert failed -0x%04x", -ret);
return false;
}
ret = mbedtls_ssl_setup(&ssl, &sslConf);
if (ret != 0) {
LOG_ERROR("ETH TLS: ssl_setup failed -0x%04x", -ret);
return false;
}
tlsServer = new EthernetServer(ETH_TLS_API_PORT);
tlsServer->begin();
LOG_INFO("ETH TLS: server listening on TCP port %d", ETH_TLS_API_PORT);
return true;
}
void serveClient(EthernetClient &client)
{
LOG_INFO("ETH TLS: client connected from %s", client.remoteIP().toString().c_str());
mbedtls_ssl_session_reset(&ssl);
mbedtls_ssl_set_bio(&ssl, &client, netSend, netRecv, nullptr);
uint32_t t0 = millis();
int ret;
do {
ret = mbedtls_ssl_handshake(&ssl);
watchdog_update();
} while (ret == MBEDTLS_ERR_SSL_WANT_READ || ret == MBEDTLS_ERR_SSL_WANT_WRITE);
if (ret != 0) {
char err[80];
mbedtls_strerror(ret, err, sizeof(err));
LOG_WARN("ETH TLS: handshake failed -0x%04x (%s) after %u ms", -ret, err, (unsigned)(millis() - t0));
return;
}
LOG_INFO("ETH TLS: handshake OK in %u ms, ciphersuite=%s", (unsigned)(millis() - t0), mbedtls_ssl_get_ciphersuite(&ssl));
MbedTlsStream stream(&ssl, &client);
handleApiClient(stream);
mbedtls_ssl_close_notify(&ssl);
}
};
static EthTlsApiServerThread *tlsThread = nullptr;
void initEthTlsApiServer()
{
if (tlsThread)
return;
tlsThread = new EthTlsApiServerThread();
LOG_INFO("ETH TLS: server worker scheduled (waits for cert ready)");
}
void deInitEthTlsApiServer()
{
// A W5500 chip reset leaves tlsServer bound to a dead socket and the cached
// mbedTLS context stale. Reset the worker back to Phase A (free the context,
// drop the listener, clear tlsReady) WITHOUT deleting the OSThread — its next
// runOnce re-waits for isEthCertReady() and rebuilds the context + rebinds
// TCP/443. Safe to free here: this runs in reconnectETH (ethConnect thread),
// and the cooperative scheduler guarantees tlsThread is not mid-runOnce, so
// nothing is using these contexts right now.
if (tlsServer) {
delete tlsServer;
tlsServer = nullptr;
}
if (tlsReady) {
mbedtls_ssl_free(&ssl);
mbedtls_ssl_config_free(&sslConf);
mbedtls_pk_free(&pkKey);
mbedtls_x509_crt_free(&certChain);
tlsReady = false;
}
}
#endif // HAS_ETHERNET && HAS_ETHERNET_TLS_API && ARCH_RP2040