feat(eth): port mesh/http to RP2350 + W5500 (HTTP/80 + HTTPS/443) (#10573)

* add watchdog to rp2xx0

* feat(eth-api): phase 0 skeleton — HTTP API server on TCP/80 (503 only)

First step of porting mesh/http/ to RP2350 + W5500 (today ESP32-only).
Phase 0 stands up the listener over the existing arduino-libraries/Ethernet
stack (no Mongoose — its built-in TCP/IP would conflict with EthernetServer
and break OTA/MQTT/NTP). Skeleton parses request line, logs it, replies 503.
Real handlers (/api/v1/{info,fromradio,toradio}) come in later phases.

- New mesh/eth/ethApiServer.{h,cpp} gated by HAS_ETHERNET_API
- Wired into ethClient init+loop in parallel with existing ethHttpOTA (port 4244)
- Enabled in both wiznet_5500_evb_pico2_e22p and pico2_w5500_e22 variants
- Build impact on wiznet: +~9 KB flash (63.0% used), <1 KB RAM

* feat(eth-api): phase 1 — implement /api/v1/{fromradio,toradio} with PhoneAPI

Replace phase 0 skeleton with the real handlers that bridge the HTTP transport
to PhoneAPI, mirroring mesh/http/ContentHandler (ESP32) semantics.

- EthHttpAPI : public PhoneAPI (api_type=TYPE_HTTP, checkIsConnected=true),
  outside the MESHTASTIC_EXCLUDE_WEBSERVER gate so it builds on RP2350.
- Minimal HTTP parser: method/path/query/Content-Length, no allocations beyond
  Arduino String, capped at 32 header lines x 256 bytes (anti-DoS).
- OPTIONS preflight -> 204 with CORS + X-Protobuf-Schema.
- GET /api/v1/fromradio[?all=true]: stream-write up to 64 protobufs from
  webAPI.getFromRadio(), Connection: close framing (HTTP/1.0 style).
- PUT /api/v1/toradio: read Content-Length bytes (<=512), call
  webAPI.handleToRadio(), echo body back.
- 404 default for unknown paths, 405 for wrong method, 400 for bad body.

Validated e2e against wiznet @ 192.168.1.143:
- OPTIONS /api/v1/fromradio -> 204 + correct CORS headers
- PUT ToRadio{want_config_id=1} (2 bytes) -> 200 + echo
- GET /api/v1/fromradio?all=true -> 200 + 3476 bytes (config+channels+nodeinfo)
- GET /api/v1/nonexistent -> 404 unknown endpoint
- OTA HTTP on port 4244 untouched and still responsive

Build impact on wiznet_5500_evb_pico2_e22p: +2 KB flash (63.1%), +1.8 KB RAM
(17.6%) vs phase 0.

* fix(eth-api): move accept loop to dedicated OSThread (20ms tick)

The API server was being polled from the Ethernet client Periodic which runs
every 5s. Measured impact before this fix on wiznet @ 192.168.1.143:

  req 1: ttfb=6.458s total=6.509s    (matches the 5s tick + handler overhead)
  req 2: connection refused          (W5500 sockets exhausted)
  req 3: connection refused

After: same hardware, same network, back-to-back requests:

  req 2: ttfb=0.287s
  req 3: ttfb=0.015s
  req 4: ttfb=0.022s
  req 5: ttfb=0.020s

The web client (meshtastic/web served from localhost) was visibly stalling
mid-handshake — it had pulled 109 nodes + 19 messages but channels and device
info never arrived. With the periodic-only polling, every request takes 5s
and the W5500's 4 hardware sockets fill up under the burst.

EthApiServerThread mirrors the WebServerThread pattern from ESP32
mesh/http/WebServer.cpp: adaptive interval — 20ms when there's recent
traffic, 100ms after 5s of idle, 500ms after 30s. Auto-registers with the
OSThread scheduler on construction.

ethHttpOTA still ticks from the periodic; left unchanged because OTA is one
large transfer that tolerates the latency, and minimizing scope here to one
behaviour change at a time.

* refactor(eth-api): extract handlers to shared module via IStreamReadWrite

Phase 2.0 of the TLS port — separate transport from request handling so
the upcoming HTTPS server can reuse the exact same parser + routing logic.

- New ethApiHandlers.{h,cpp}: Request, parser, CORS helpers, fromradio/toradio
  handlers, EthHttpAPI : PhoneAPI subclass. All driven by a single
  IStreamReadWrite interface that inherits Print (so client.print(...) keeps
  working transparently).
- ethApiServer.cpp slimmed to ~90 LOC: now just the EthernetServer(80) +
  OSThread + an EthernetClientStream adapter that forwards reads/writes to
  the underlying EthernetClient. No behaviour change.

Validated on wiznet @ 192.168.1.143: 5 curl requests TTFB 18-110ms (same
as pre-refactor), protobuf round-trip PUT 200 + GET ?all=true 3499 bytes.
Flash impact: +200 bytes (63.2%); RAM unchanged (17.6%).

* feat(eth-tls): ECDSA P-256 self-signed cert generation for HTTPS API

mbedTLS-based cert generation module that produces a SAN=IP self-signed
ECDSA P-256 server certificate, persisted under LittleFS so subsequent
boots reuse the same key. Generation runs once on a dedicated OSThread
so the ECDSA keygen path (~430 ms) never blocks the Periodic stack or
the Ethernet reconnect loop.

  - mbedTLS 3.6.2 sources compiled in via scripts/add_mbedtls_sources.py
    (BuildSources of pico-sdk/lib/mbedtls/library/*.c, all 108 files).
    MBEDTLS_USER_CONFIG_FILE injected as CPPDEFINES tuple in the script —
    build_flags shell escape mangles the embedded quotes on Windows.
  - src/mbedtls_user_config.h undefs MBEDTLS_HAVE_TIME, HAVE_TIME_DATE,
    TIMING_C, NET_C, FS_IO, PSA_ITS_FILE_C, PSA_CRYPTO_STORAGE_C, and
    defines MBEDTLS_NO_PLATFORM_ENTROPY (entropy_poll.c uses a raw
    platform check, not gated by a flag).
  - src/mesh/eth/ethCert.{h,cpp}: ECDSA P-256 keypair + cert with
    SAN(IP=current) using pico-sdk get_rand_64() directly as f_rng,
    bypassing mbedtls_entropy. DER buffers heap-allocated to keep the
    OSThread stack within budget. Cert + key + ip persisted under
    /prefs/eth_*.der so subsequent boots reuse the same identity.
  - Gated by HAS_ETHERNET_TLS_API. Standalone phase: generation only;
    HTTPS server (TCP/443) wired up in the follow-up commit.

* feat(eth-tls-api): phase 2.2 — HTTPS server on TCP/443 reusing handlers

Brings up an mbedTLS server on port 443 that defers to the ECDSA P-256
self-signed cert produced by [[ethCert]]. Reuses the HTTP request /
response handlers from [[ethApiHandlers]] via the IStreamReadWrite
interface — there is exactly one code path for routing / CORS / PhoneAPI
integration regardless of whether the transport is plain TCP or TLS.

EthTlsApiServerThread (OSThread):
  - Phase A: poll isEthCertReady() every 500 ms while the cert worker
    runs. Once true, parse cert chain + key, build ssl_config (TLS server,
    stream transport, default preset, VERIFY_NONE since we are the cert
    issuer), install own cert, run ssl_setup, bind tlsServer on 443.
  - Phase B: standard adaptive accept loop (20 / 100 / 500 ms tick),
    identical to the plain-HTTP server.

Per-connection flow:
  1. session_reset on the static ssl context (1 in-flight session — multi-
     session pool is Phase 3 if needed)
  2. set_bio routes mbedtls I/O to two C callbacks (netSend / netRecv)
     that bridge to the live EthernetClient via the void* ctx
  3. handshake loop (sync, blocking, with 10 s recv timeout) — mbedTLS
     errors are logged with mbedtls_strerror
  4. wrap (ssl, client) in MbedTlsStream → handleApiClient(stream)
  5. close_notify + client.stop

Stack budget continues the Phase 2.1-bis discipline: every large buffer
(ssl context, cert chain, pk_key, ssl_config) lives in BSS as a static
global. The OSThread stack only holds the per-connection EthernetClient
adapter and small mbedtls return codes.

Validated on wiznet 192.168.1.143:
  - cert load-from-FS path: 13 ms (regen skipped on second boot)
  - cert gen path: 290 ms (first boot only)
  - ssl_setup chain: parse cert, parse key, ssl_config_defaults,
    conf_own_cert, ssl_setup all return 0
  - end-to-end: curl -k https://192.168.1.143/api/v1/fromradio → 200 OK
    with application/x-protobuf, CORS, X-Protobuf-Schema headers, server
    initiates close_notify cleanly

Footprint vs 2.0:
  - flash 70.5% → 87.3% (+220 KB for mbedtls_ssl + x509 server-side code)
  - RAM static 17.6% → 19.8% (+11 KB BSS for ssl_context + cert chain)
  - heap at runtime adds ~32 KB per active session (mbedtls in/out
    record buffers, default MBEDTLS_SSL_*_CONTENT_LEN=16384)

Next: validate from Firefox direct (https://192.168.1.143 → warning
accept → JSON visible), then from client.meshtastic.org hosted to
confirm the mixed-content block is gone.

* fix(eth-tls): add KeyUsage + EKU serverAuth and cap TLS 1.2 for browser compat

Two browser-compat fixes that surfaced in Firefox validation:

1. Cert v1 only had Basic Constraints + SAN. NSS / Firefox refuse to
   treat a cert as a TLS server cert without an Extended Key Usage
   extension naming id-kp-serverAuth since 2023 — the error surfaces as
   a non-overridable 'Secure Connection Failed' with no 'Accept the
   Risk' path. Add KeyUsage(digitalSignature + keyEncipherment, critical)
   and ExtendedKeyUsage(serverAuth, critical). Bump cert/key/ip file
   paths to '_v2' so live boards regenerate on next start instead of
   loading a v1 cert that the browser silently refuses.

2. pico-sdk mbedtls defines MBEDTLS_SSL_PROTO_TLS1_3 in its default
   config, but the server-side 1.3 plumbing in this vendored build is
   incomplete: Firefox and openssl-3.5's s_client default to 1.3 and
   the handshake dies 4 ms in with MBEDTLS_ERR_ERROR_GENERIC_ERROR
   (-0x0001). curl/SChannel happened to default to 1.2 so it masked the
   issue earlier. Cap min/max_tls_version to TLS 1.2 — clients downgrade
   transparently and we keep the ECDHE-ECDSA + AES-GCM / CHACHA20-POLY1305
   suites that already work end-to-end.

Validated on wiznet 192.168.1.143:
  - cert v2 dumps with the 4 extensions visible (openssl x509 -text):
    BasicConstraints CA:FALSE, KeyUsage(critical) digitalSignature +
    keyEncipherment, ExtendedKeyUsage(critical) serverAuth, SAN IP.
  - openssl s_client (no version flag): downgrades to TLS 1.2, handshake
    completes, verify_return=18 (self-signed, expected).
  - Firefox: warning self-signed -> Advanced -> Accept Risk -> handshake
    OK in 632 ms with CHACHA20-POLY1305, request reaches handleApiClient
    and returns the protobuf body.

* perf(eth-api): HTTP/1.1 keep-alive — one handshake per session, not per request

Before this change /fromradio used 'Connection: close' framing (HTTP/1.0
style with no Content-Length), forcing each poll to redo the full TLS
handshake. client.meshtastic.org needs ~80 sequential /fromradio polls
during initial sync (one per packet from the config-replay state
machine: MyInfo, channels, every Config_*, every ModuleConfig_*, every
NodeInfo), so the user-visible load time was ~50 s of pure handshake
overhead (80 requests * ~625 ms ECDSA P-256 each).

Three coordinated changes:

1. handleApiClient() now loops on the same connection until the peer
   closes or parseRequest hits its 3 s idle timeout. requestsServed
   counter keeps the 'bad/timeout request' debug log from firing on
   the natural idle close after a keep-alive sequence.

2. handleFromRadio() buffers all packets into a std::vector before
   writing, so it can emit a real Content-Length and 'Connection:
   keep-alive'. Buffer is dynamic — common 1-packet response only
   allocates ~256 B; ?all=true keeps the 64-packet cap which tops out
   around 16 KB. handleToRadio + sendPreflight switched to keep-alive
   too (they already had real Content-Length). sendError keeps close —
   errors are terminal.

3. ethTlsApiServer netRecv RECV_TIMEOUT_MS dropped from 10 s to 3 s so
   mbedtls_ssl_read can't outlast the handler's idle deadline (a quiet
   browser leaving the socket open would otherwise wedge the OSThread
   for 10 s past the natural close).

Measured on wiznet 192.168.1.143 against client.meshtastic.org:
  - one TLS handshake (641 ms, CHACHA20-POLY1305)
  - ~80 requests pipelined over the same session
  - full config + 40 NodeInfos in ~6-8 s (vs ~50 s before)
  - per-request latency post-handshake: ~10-25 ms
  - curl -kv with two URLs: 'Reusing existing https: connection',
    server returns 'Connection: keep-alive' explicitly.

* fix(eth-tls): watchdog + Chrome compat — handshake stability across browsers

Phase 3 keep-alive shipped a working Firefox flow but client.meshtastic.org
loops + Chrome 'Test Connection' both rebooted the board. Four distinct
issues; collectively they kept the OSThread inside serveClient() too long
or spin-looping without yielding, and pico-sdk mbedtls' TLS 1.3 code path
choked on Chrome's modern ClientHello.

1. Watchdog reset during keep-alive idle. Once a client drained the
   replay queue and entered 3 s poll mode, the OSThread sat inside
   netRecv()'s busy-wait waiting for the next request. Two consecutive
   3 s waits plus prior handler time crossed the 8 s RP2350 hardware
   watchdog. Pet the watchdog inside netRecv()'s poll loop (every 2 ms)
   so a quiet client can never starve the watchdog. Same fix in the
   ethApiHandlers per-request yield path.

2. Cap session at 64 requests + yield() between. Defense-in-depth: a
   pathological client can't monopolize serveClient indefinitely; after
   the cap it just re-handshakes (~625 ms), still vastly cheaper than
   the per-request handshake we had before keep-alive.

3. TLS 1.3 code compiled out of mbedtls entirely
   (#undef MBEDTLS_SSL_PROTO_TLS1_3 in mbedtls_user_config). Capping
   max_tls_version=TLS1_2 at runtime is enough for Firefox / openssl
   (they downgrade cleanly), but Chrome's ClientHello carries TLS 1.3
   extensions — post-quantum key shares, Encrypted ClientHello,
   etc. — that the vendored mbedtls 1.3 parser crashes on before the
   downgrade decision happens. Removing the 1.3 sources sidesteps the
   parser; ServerHello just announces TLS 1.2 and Chrome accepts.

4. netSend infinite WANT_WRITE spin. When W5500's TX buffer momentarily
   filled mid-handshake (Chrome draining slower than Firefox during
   ServerKeyExchange), EthernetClient::write() returned 0, our netSend
   returned MBEDTLS_ERR_SSL_WANT_WRITE without delay, mbedtls retried
   immediately, repeat at ~180k iter/sec until ... well, until the
   board's other threads got nothing done. Log signature: ret=-0x6880
   tight-looping in the handshake iter trace. Rewrite netSend to block
   with delay(2) + watchdog_update() and a 3 s timeout — same shape as
   netRecv. Return MBEDTLS_ERR_SSL_PEER_CLOSE_NOTIFY on disconnect
   (was incorrectly returning WANT_WRITE).

Also added granular per-iter handshake logging gated on first 20 iters
+ every 50th after that, so any future regression localizes itself in
COM9 without RTT JLink.

Validated on wiznet 192.168.1.143:
  - Firefox: client.meshtastic.org full sync + idle poll stable (no
    reset during the previously-crashing 'replay drain complete' phase)
  - Chrome: 'Test Connection' accepts the cert prompt and connects
  - Edge: same as Chrome
  - openssl s_client default + tls1_2 forced: both negotiate TLS 1.2
    with ECDHE-ECDSA + AES-GCM / CHACHA20-POLY1305, verify=18 (self-
    signed, expected)

* chore(eth-tls): drop verbose debug logs from cert + TLS init paths

The cert pipeline + TLS context init had step-by-step LOG_INFOs and
ubiquitous Serial.flush() that were essential while diagnosing the
Phase 2.1-bis stack overflow, the Chrome handshake crash, and the
keep-alive watchdog reset. Once those bugs were fixed the logs just
clutter COM9 on every boot.

Kept on hand:
  - cert: 'loaded from FS', 'generating ...', 'generated N B in T ms',
    'persisted to LittleFS', plus all LOG_ERROR / LOG_WARN paths
  - tls: 'server listening on TCP port 443', 'client connected from',
    'handshake OK in N ms ciphersuite=', 'handshake failed -0xXXXX (...)',
    plus all init LOG_ERRORs

Dropped:
  - cert: 'step 1/8 pk_setup' through 'step 8/8 copy key DER', 'thread
    woke', 'pipeline OK' (now silent on success)
  - tls: 'cert is ready, initializing', 'parsing cert chain', 'parsing
    key', 'ssl_config_defaults', 'conf_own_cert', 'ssl_setup',
    'server worker scheduled', and the per-iter handshake trace
  - obsolete 'Optional mbedtls debug bridge' commented-out stub
  - all Serial.flush() that were added defensively for the
    debug-the-crash phase

Bin shrinks ~40 KB (logs + format strings). Validated on wiznet
192.168.1.143: HTTP 200 round-trip works post-flash, no regression.

* fix(eth): harden API/TLS build guards (review #10573)

- ethApiHandlers.{h,cpp}: gate on (HAS_ETHERNET_API || HAS_ETHERNET_TLS_API)
  so the shared handleApiClient/IStreamReadWrite still compile for an
  HTTPS-only variant (the TLS server depends on them).
- ethTlsApiServer.{cpp,h}, ethCert.{cpp,h}: require ARCH_RP2040 (they use Pico
  SDK get_rand_64()/watchdog_update()); the ethClient.cpp TLS include + call
  site are tightened to match, so enabling the flag on a non-RP2040 Ethernet
  target is a clean no-op instead of a cryptic build break.
- clang-format (16) the touched eth files to satisfy Trunk Check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(eth): address review — listener/cert recovery + keep-alive framing

- Rebind HTTP/80 and HTTPS/443 after a W5500 chip reset. reconnectETH() now
  tears the API/TLS servers down (deInitEthApiServer/deInitEthTlsApiServer) so
  the restart path recreates them; the singleton guards previously left both
  bound to dead sockets until reboot. The workers are kept (only the listener
  is dropped) so nothing is deleted from another thread's runOnce.

- Stop the keep-alive loop after a handler error. /fromradio and /toradio
  return whether the connection may stay open; a 400/405/408 response
  (Connection: close) now ends the loop so unread body bytes can't be parsed
  as the next request.

- Validate the cached cert/key before trusting it (parse both DERs), and clear
  the IP commit-marker before rewriting the pair and write it last, so a reset
  mid-persist regenerates instead of handing the TLS server a truncated or
  mismatched pair (which would hard-disable HTTPS).

- Track a cert generation counter so the TLS server reloads and rebinds when a
  DHCP lease change regenerates the cert for a new IP (the SAN must follow, or
  browsers reject the new address). The cert worker is no longer one-shot.

- Silence the SCons F821/E402 lint on add_mbedtls_sources.py to match the other
  extra_scripts.

Validated on a W5500 board: a forced chip reset rebinds 80 and 443 without an
MCU reboot; a truncated cached cert regenerates and HTTPS recovers; a simulated
lease change reissues the cert with the new IP in the SAN.

* fix(eth): clear cert readiness on regen failure + verify cached pair

Follow-up to the review:
- Reset ready_/certIp_ when ensureCertForIp() fails, so isReady() no longer
  reports true with empty material — otherwise a later TLS reload fails
  initTlsContext() and stays disabled instead of retrying on the next poll.
- certKeyParse() now also checks the cached cert and key form a matching pair
  (mbedtls_pk_check_pair), rejecting an independently-parseable but mismatched
  cert/key (e.g. if the IP commit-marker survived a failed clear).

Validated on a W5500 board: a valid cached pair still loads from FS, and a
regenerated cert reloads cleanly.

---------

Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Carlos Valdes
2026-07-01 05:51:14 -05:00
committed by GitHub
co-authored by GitHub Ben Meadors Claude Opus 4.8
parent cfc2e457a9
commit e64d20548c
13 changed files with 1470 additions and 1 deletions
+3
View File
@@ -1325,6 +1325,9 @@ void loop()
#endif
#ifdef ARCH_NRF54L15
nrf54l15Loop();
#endif
#ifdef ARCH_RP2040
rp2040Loop();
#endif
power->powerCommandsCheck();