fe6bf91343867fb52610d86aebdac823e9fb018f
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3becaf2d95 | emdashes begone (#10847) | ||
|
|
e64d20548c |
feat(eth): port mesh/http to RP2350 + W5500 (HTTP/80 + HTTPS/443) (#10573)
* add watchdog to rp2xx0
* feat(eth-api): phase 0 skeleton — HTTP API server on TCP/80 (503 only)
First step of porting mesh/http/ to RP2350 + W5500 (today ESP32-only).
Phase 0 stands up the listener over the existing arduino-libraries/Ethernet
stack (no Mongoose — its built-in TCP/IP would conflict with EthernetServer
and break OTA/MQTT/NTP). Skeleton parses request line, logs it, replies 503.
Real handlers (/api/v1/{info,fromradio,toradio}) come in later phases.
- New mesh/eth/ethApiServer.{h,cpp} gated by HAS_ETHERNET_API
- Wired into ethClient init+loop in parallel with existing ethHttpOTA (port 4244)
- Enabled in both wiznet_5500_evb_pico2_e22p and pico2_w5500_e22 variants
- Build impact on wiznet: +~9 KB flash (63.0% used), <1 KB RAM
* feat(eth-api): phase 1 — implement /api/v1/{fromradio,toradio} with PhoneAPI
Replace phase 0 skeleton with the real handlers that bridge the HTTP transport
to PhoneAPI, mirroring mesh/http/ContentHandler (ESP32) semantics.
- EthHttpAPI : public PhoneAPI (api_type=TYPE_HTTP, checkIsConnected=true),
outside the MESHTASTIC_EXCLUDE_WEBSERVER gate so it builds on RP2350.
- Minimal HTTP parser: method/path/query/Content-Length, no allocations beyond
Arduino String, capped at 32 header lines x 256 bytes (anti-DoS).
- OPTIONS preflight -> 204 with CORS + X-Protobuf-Schema.
- GET /api/v1/fromradio[?all=true]: stream-write up to 64 protobufs from
webAPI.getFromRadio(), Connection: close framing (HTTP/1.0 style).
- PUT /api/v1/toradio: read Content-Length bytes (<=512), call
webAPI.handleToRadio(), echo body back.
- 404 default for unknown paths, 405 for wrong method, 400 for bad body.
Validated e2e against wiznet @ 192.168.1.143:
- OPTIONS /api/v1/fromradio -> 204 + correct CORS headers
- PUT ToRadio{want_config_id=1} (2 bytes) -> 200 + echo
- GET /api/v1/fromradio?all=true -> 200 + 3476 bytes (config+channels+nodeinfo)
- GET /api/v1/nonexistent -> 404 unknown endpoint
- OTA HTTP on port 4244 untouched and still responsive
Build impact on wiznet_5500_evb_pico2_e22p: +2 KB flash (63.1%), +1.8 KB RAM
(17.6%) vs phase 0.
* fix(eth-api): move accept loop to dedicated OSThread (20ms tick)
The API server was being polled from the Ethernet client Periodic which runs
every 5s. Measured impact before this fix on wiznet @ 192.168.1.143:
req 1: ttfb=6.458s total=6.509s (matches the 5s tick + handler overhead)
req 2: connection refused (W5500 sockets exhausted)
req 3: connection refused
After: same hardware, same network, back-to-back requests:
req 2: ttfb=0.287s
req 3: ttfb=0.015s
req 4: ttfb=0.022s
req 5: ttfb=0.020s
The web client (meshtastic/web served from localhost) was visibly stalling
mid-handshake — it had pulled 109 nodes + 19 messages but channels and device
info never arrived. With the periodic-only polling, every request takes 5s
and the W5500's 4 hardware sockets fill up under the burst.
EthApiServerThread mirrors the WebServerThread pattern from ESP32
mesh/http/WebServer.cpp: adaptive interval — 20ms when there's recent
traffic, 100ms after 5s of idle, 500ms after 30s. Auto-registers with the
OSThread scheduler on construction.
ethHttpOTA still ticks from the periodic; left unchanged because OTA is one
large transfer that tolerates the latency, and minimizing scope here to one
behaviour change at a time.
* refactor(eth-api): extract handlers to shared module via IStreamReadWrite
Phase 2.0 of the TLS port — separate transport from request handling so
the upcoming HTTPS server can reuse the exact same parser + routing logic.
- New ethApiHandlers.{h,cpp}: Request, parser, CORS helpers, fromradio/toradio
handlers, EthHttpAPI : PhoneAPI subclass. All driven by a single
IStreamReadWrite interface that inherits Print (so client.print(...) keeps
working transparently).
- ethApiServer.cpp slimmed to ~90 LOC: now just the EthernetServer(80) +
OSThread + an EthernetClientStream adapter that forwards reads/writes to
the underlying EthernetClient. No behaviour change.
Validated on wiznet @ 192.168.1.143: 5 curl requests TTFB 18-110ms (same
as pre-refactor), protobuf round-trip PUT 200 + GET ?all=true 3499 bytes.
Flash impact: +200 bytes (63.2%); RAM unchanged (17.6%).
* feat(eth-tls): ECDSA P-256 self-signed cert generation for HTTPS API
mbedTLS-based cert generation module that produces a SAN=IP self-signed
ECDSA P-256 server certificate, persisted under LittleFS so subsequent
boots reuse the same key. Generation runs once on a dedicated OSThread
so the ECDSA keygen path (~430 ms) never blocks the Periodic stack or
the Ethernet reconnect loop.
- mbedTLS 3.6.2 sources compiled in via scripts/add_mbedtls_sources.py
(BuildSources of pico-sdk/lib/mbedtls/library/*.c, all 108 files).
MBEDTLS_USER_CONFIG_FILE injected as CPPDEFINES tuple in the script —
build_flags shell escape mangles the embedded quotes on Windows.
- src/mbedtls_user_config.h undefs MBEDTLS_HAVE_TIME, HAVE_TIME_DATE,
TIMING_C, NET_C, FS_IO, PSA_ITS_FILE_C, PSA_CRYPTO_STORAGE_C, and
defines MBEDTLS_NO_PLATFORM_ENTROPY (entropy_poll.c uses a raw
platform check, not gated by a flag).
- src/mesh/eth/ethCert.{h,cpp}: ECDSA P-256 keypair + cert with
SAN(IP=current) using pico-sdk get_rand_64() directly as f_rng,
bypassing mbedtls_entropy. DER buffers heap-allocated to keep the
OSThread stack within budget. Cert + key + ip persisted under
/prefs/eth_*.der so subsequent boots reuse the same identity.
- Gated by HAS_ETHERNET_TLS_API. Standalone phase: generation only;
HTTPS server (TCP/443) wired up in the follow-up commit.
* feat(eth-tls-api): phase 2.2 — HTTPS server on TCP/443 reusing handlers
Brings up an mbedTLS server on port 443 that defers to the ECDSA P-256
self-signed cert produced by [[ethCert]]. Reuses the HTTP request /
response handlers from [[ethApiHandlers]] via the IStreamReadWrite
interface — there is exactly one code path for routing / CORS / PhoneAPI
integration regardless of whether the transport is plain TCP or TLS.
EthTlsApiServerThread (OSThread):
- Phase A: poll isEthCertReady() every 500 ms while the cert worker
runs. Once true, parse cert chain + key, build ssl_config (TLS server,
stream transport, default preset, VERIFY_NONE since we are the cert
issuer), install own cert, run ssl_setup, bind tlsServer on 443.
- Phase B: standard adaptive accept loop (20 / 100 / 500 ms tick),
identical to the plain-HTTP server.
Per-connection flow:
1. session_reset on the static ssl context (1 in-flight session — multi-
session pool is Phase 3 if needed)
2. set_bio routes mbedtls I/O to two C callbacks (netSend / netRecv)
that bridge to the live EthernetClient via the void* ctx
3. handshake loop (sync, blocking, with 10 s recv timeout) — mbedTLS
errors are logged with mbedtls_strerror
4. wrap (ssl, client) in MbedTlsStream → handleApiClient(stream)
5. close_notify + client.stop
Stack budget continues the Phase 2.1-bis discipline: every large buffer
(ssl context, cert chain, pk_key, ssl_config) lives in BSS as a static
global. The OSThread stack only holds the per-connection EthernetClient
adapter and small mbedtls return codes.
Validated on wiznet 192.168.1.143:
- cert load-from-FS path: 13 ms (regen skipped on second boot)
- cert gen path: 290 ms (first boot only)
- ssl_setup chain: parse cert, parse key, ssl_config_defaults,
conf_own_cert, ssl_setup all return 0
- end-to-end: curl -k https://192.168.1.143/api/v1/fromradio → 200 OK
with application/x-protobuf, CORS, X-Protobuf-Schema headers, server
initiates close_notify cleanly
Footprint vs 2.0:
- flash 70.5% → 87.3% (+220 KB for mbedtls_ssl + x509 server-side code)
- RAM static 17.6% → 19.8% (+11 KB BSS for ssl_context + cert chain)
- heap at runtime adds ~32 KB per active session (mbedtls in/out
record buffers, default MBEDTLS_SSL_*_CONTENT_LEN=16384)
Next: validate from Firefox direct (https://192.168.1.143 → warning
accept → JSON visible), then from client.meshtastic.org hosted to
confirm the mixed-content block is gone.
* fix(eth-tls): add KeyUsage + EKU serverAuth and cap TLS 1.2 for browser compat
Two browser-compat fixes that surfaced in Firefox validation:
1. Cert v1 only had Basic Constraints + SAN. NSS / Firefox refuse to
treat a cert as a TLS server cert without an Extended Key Usage
extension naming id-kp-serverAuth since 2023 — the error surfaces as
a non-overridable 'Secure Connection Failed' with no 'Accept the
Risk' path. Add KeyUsage(digitalSignature + keyEncipherment, critical)
and ExtendedKeyUsage(serverAuth, critical). Bump cert/key/ip file
paths to '_v2' so live boards regenerate on next start instead of
loading a v1 cert that the browser silently refuses.
2. pico-sdk mbedtls defines MBEDTLS_SSL_PROTO_TLS1_3 in its default
config, but the server-side 1.3 plumbing in this vendored build is
incomplete: Firefox and openssl-3.5's s_client default to 1.3 and
the handshake dies 4 ms in with MBEDTLS_ERR_ERROR_GENERIC_ERROR
(-0x0001). curl/SChannel happened to default to 1.2 so it masked the
issue earlier. Cap min/max_tls_version to TLS 1.2 — clients downgrade
transparently and we keep the ECDHE-ECDSA + AES-GCM / CHACHA20-POLY1305
suites that already work end-to-end.
Validated on wiznet 192.168.1.143:
- cert v2 dumps with the 4 extensions visible (openssl x509 -text):
BasicConstraints CA:FALSE, KeyUsage(critical) digitalSignature +
keyEncipherment, ExtendedKeyUsage(critical) serverAuth, SAN IP.
- openssl s_client (no version flag): downgrades to TLS 1.2, handshake
completes, verify_return=18 (self-signed, expected).
- Firefox: warning self-signed -> Advanced -> Accept Risk -> handshake
OK in 632 ms with CHACHA20-POLY1305, request reaches handleApiClient
and returns the protobuf body.
* perf(eth-api): HTTP/1.1 keep-alive — one handshake per session, not per request
Before this change /fromradio used 'Connection: close' framing (HTTP/1.0
style with no Content-Length), forcing each poll to redo the full TLS
handshake. client.meshtastic.org needs ~80 sequential /fromradio polls
during initial sync (one per packet from the config-replay state
machine: MyInfo, channels, every Config_*, every ModuleConfig_*, every
NodeInfo), so the user-visible load time was ~50 s of pure handshake
overhead (80 requests * ~625 ms ECDSA P-256 each).
Three coordinated changes:
1. handleApiClient() now loops on the same connection until the peer
closes or parseRequest hits its 3 s idle timeout. requestsServed
counter keeps the 'bad/timeout request' debug log from firing on
the natural idle close after a keep-alive sequence.
2. handleFromRadio() buffers all packets into a std::vector before
writing, so it can emit a real Content-Length and 'Connection:
keep-alive'. Buffer is dynamic — common 1-packet response only
allocates ~256 B; ?all=true keeps the 64-packet cap which tops out
around 16 KB. handleToRadio + sendPreflight switched to keep-alive
too (they already had real Content-Length). sendError keeps close —
errors are terminal.
3. ethTlsApiServer netRecv RECV_TIMEOUT_MS dropped from 10 s to 3 s so
mbedtls_ssl_read can't outlast the handler's idle deadline (a quiet
browser leaving the socket open would otherwise wedge the OSThread
for 10 s past the natural close).
Measured on wiznet 192.168.1.143 against client.meshtastic.org:
- one TLS handshake (641 ms, CHACHA20-POLY1305)
- ~80 requests pipelined over the same session
- full config + 40 NodeInfos in ~6-8 s (vs ~50 s before)
- per-request latency post-handshake: ~10-25 ms
- curl -kv with two URLs: 'Reusing existing https: connection',
server returns 'Connection: keep-alive' explicitly.
* fix(eth-tls): watchdog + Chrome compat — handshake stability across browsers
Phase 3 keep-alive shipped a working Firefox flow but client.meshtastic.org
loops + Chrome 'Test Connection' both rebooted the board. Four distinct
issues; collectively they kept the OSThread inside serveClient() too long
or spin-looping without yielding, and pico-sdk mbedtls' TLS 1.3 code path
choked on Chrome's modern ClientHello.
1. Watchdog reset during keep-alive idle. Once a client drained the
replay queue and entered 3 s poll mode, the OSThread sat inside
netRecv()'s busy-wait waiting for the next request. Two consecutive
3 s waits plus prior handler time crossed the 8 s RP2350 hardware
watchdog. Pet the watchdog inside netRecv()'s poll loop (every 2 ms)
so a quiet client can never starve the watchdog. Same fix in the
ethApiHandlers per-request yield path.
2. Cap session at 64 requests + yield() between. Defense-in-depth: a
pathological client can't monopolize serveClient indefinitely; after
the cap it just re-handshakes (~625 ms), still vastly cheaper than
the per-request handshake we had before keep-alive.
3. TLS 1.3 code compiled out of mbedtls entirely
(#undef MBEDTLS_SSL_PROTO_TLS1_3 in mbedtls_user_config). Capping
max_tls_version=TLS1_2 at runtime is enough for Firefox / openssl
(they downgrade cleanly), but Chrome's ClientHello carries TLS 1.3
extensions — post-quantum key shares, Encrypted ClientHello,
etc. — that the vendored mbedtls 1.3 parser crashes on before the
downgrade decision happens. Removing the 1.3 sources sidesteps the
parser; ServerHello just announces TLS 1.2 and Chrome accepts.
4. netSend infinite WANT_WRITE spin. When W5500's TX buffer momentarily
filled mid-handshake (Chrome draining slower than Firefox during
ServerKeyExchange), EthernetClient::write() returned 0, our netSend
returned MBEDTLS_ERR_SSL_WANT_WRITE without delay, mbedtls retried
immediately, repeat at ~180k iter/sec until ... well, until the
board's other threads got nothing done. Log signature: ret=-0x6880
tight-looping in the handshake iter trace. Rewrite netSend to block
with delay(2) + watchdog_update() and a 3 s timeout — same shape as
netRecv. Return MBEDTLS_ERR_SSL_PEER_CLOSE_NOTIFY on disconnect
(was incorrectly returning WANT_WRITE).
Also added granular per-iter handshake logging gated on first 20 iters
+ every 50th after that, so any future regression localizes itself in
COM9 without RTT JLink.
Validated on wiznet 192.168.1.143:
- Firefox: client.meshtastic.org full sync + idle poll stable (no
reset during the previously-crashing 'replay drain complete' phase)
- Chrome: 'Test Connection' accepts the cert prompt and connects
- Edge: same as Chrome
- openssl s_client default + tls1_2 forced: both negotiate TLS 1.2
with ECDHE-ECDSA + AES-GCM / CHACHA20-POLY1305, verify=18 (self-
signed, expected)
* chore(eth-tls): drop verbose debug logs from cert + TLS init paths
The cert pipeline + TLS context init had step-by-step LOG_INFOs and
ubiquitous Serial.flush() that were essential while diagnosing the
Phase 2.1-bis stack overflow, the Chrome handshake crash, and the
keep-alive watchdog reset. Once those bugs were fixed the logs just
clutter COM9 on every boot.
Kept on hand:
- cert: 'loaded from FS', 'generating ...', 'generated N B in T ms',
'persisted to LittleFS', plus all LOG_ERROR / LOG_WARN paths
- tls: 'server listening on TCP port 443', 'client connected from',
'handshake OK in N ms ciphersuite=', 'handshake failed -0xXXXX (...)',
plus all init LOG_ERRORs
Dropped:
- cert: 'step 1/8 pk_setup' through 'step 8/8 copy key DER', 'thread
woke', 'pipeline OK' (now silent on success)
- tls: 'cert is ready, initializing', 'parsing cert chain', 'parsing
key', 'ssl_config_defaults', 'conf_own_cert', 'ssl_setup',
'server worker scheduled', and the per-iter handshake trace
- obsolete 'Optional mbedtls debug bridge' commented-out stub
- all Serial.flush() that were added defensively for the
debug-the-crash phase
Bin shrinks ~40 KB (logs + format strings). Validated on wiznet
192.168.1.143: HTTP 200 round-trip works post-flash, no regression.
* fix(eth): harden API/TLS build guards (review #10573)
- ethApiHandlers.{h,cpp}: gate on (HAS_ETHERNET_API || HAS_ETHERNET_TLS_API)
so the shared handleApiClient/IStreamReadWrite still compile for an
HTTPS-only variant (the TLS server depends on them).
- ethTlsApiServer.{cpp,h}, ethCert.{cpp,h}: require ARCH_RP2040 (they use Pico
SDK get_rand_64()/watchdog_update()); the ethClient.cpp TLS include + call
site are tightened to match, so enabling the flag on a non-RP2040 Ethernet
target is a clean no-op instead of a cryptic build break.
- clang-format (16) the touched eth files to satisfy Trunk Check.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(eth): address review — listener/cert recovery + keep-alive framing
- Rebind HTTP/80 and HTTPS/443 after a W5500 chip reset. reconnectETH() now
tears the API/TLS servers down (deInitEthApiServer/deInitEthTlsApiServer) so
the restart path recreates them; the singleton guards previously left both
bound to dead sockets until reboot. The workers are kept (only the listener
is dropped) so nothing is deleted from another thread's runOnce.
- Stop the keep-alive loop after a handler error. /fromradio and /toradio
return whether the connection may stay open; a 400/405/408 response
(Connection: close) now ends the loop so unread body bytes can't be parsed
as the next request.
- Validate the cached cert/key before trusting it (parse both DERs), and clear
the IP commit-marker before rewriting the pair and write it last, so a reset
mid-persist regenerates instead of handing the TLS server a truncated or
mismatched pair (which would hard-disable HTTPS).
- Track a cert generation counter so the TLS server reloads and rebinds when a
DHCP lease change regenerates the cert for a new IP (the SAN must follow, or
browsers reject the new address). The cert worker is no longer one-shot.
- Silence the SCons F821/E402 lint on add_mbedtls_sources.py to match the other
extra_scripts.
Validated on a W5500 board: a forced chip reset rebinds 80 and 443 without an
MCU reboot; a truncated cached cert regenerates and HTTPS recovers; a simulated
lease change reissues the cert with the new IP in the SAN.
* fix(eth): clear cert readiness on regen failure + verify cached pair
Follow-up to the review:
- Reset ready_/certIp_ when ensureCertForIp() fails, so isReady() no longer
reports true with empty material — otherwise a later TLS reload fails
initTlsContext() and stays disabled instead of retrying on the next poll.
- certKeyParse() now also checks the cached cert and key form a matching pair
(mbedtls_pk_check_pair), rejecting an independently-parseable but mismatched
cert/key (e.g. if the IP commit-marker survived a failed clear).
Validated on a W5500 board: a valid cached pair still loads from FS, and a
regenerated cert reloads cleanly.
---------
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|