Commit Graph
12335 Commits
Author SHA1 Message Date
Ben MeadorsandGitHub 5c5cb094e6 fix(admin): stop the admin dispatcher from overflowing the ESP32 loopTask stack (#11239)
AdminModule::handleReceivedProtobuf carried a 3008-byte stack frame - 37% of
the 8 KB Arduino loopTask stack. GCC inlines the eight handleGet* helpers into
it, each of which builds a whole meshtastic_AdminMessage (480 B), so the
dispatcher reserved a slot for every one of them at once even though only one
case ever runs. Two more full AdminMessages lived directly in the function.

That left any admin write short of stack for what runs beneath it:
NodeDB::saveToDisk -> saveProto -> pb_encode -> SafeFile -> LittleFS -> flash.
On heltec-v4 the canary fired at 7824 of 8192 bytes while writing
/prefs/config.proto, so no setting ever persisted and the device rebooted.

Mark the response builders NOINLINE and move the module-API and observer
responses into their own functions. Pure refactor; measured on heltec-v4 the
dispatcher frame drops 3008 -> 384 bytes, taking the crashing path from 368
bytes of headroom to 2992.

Fixes #11237
2026-07-26 21:34:38 +00:00
AustinandGitHub 69b68e2c42 ESP32: Don't run pkg install on build (#11243)
pioarduino runs this anyways. No need to do it twice.
2026-07-26 17:11:58 -04:00
AustinandGitHub 2c8a2a8cb8 Docker: Only cache qemu/buildx on default branch (#11241)
Docker qemu / buildx are eating up lots of cache for little benefit. Remove unless we're building from the main branch.
2026-07-26 16:47:53 -04:00
AustinandGitHub 67e12dc8c0 Trunk: Only cache nightly runs, base on develop (#11240)
These trunk caches are eating us for dinner. Remove them for PRs so we aren't storing 20 copies
2026-07-26 14:53:42 -04:00
4800484d41 lint: ignore userPrefs.jsonc in trunk (#11174)
checkov/CKV_SECRET_6 flags the commented-out "large4cats" example, which
is the public default credential for the meshtastic.org MQTT broker rather
than a real secret. Ignore the file in .trunk/trunk.yaml so the example
config stays free of lint-suppression comments.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-26 08:26:27 +00:00
JorropoandGitHub 6be7d9ceeb lint: hoist stdlib imports above Import(env) in add_mbedtls_sources (#11172)
glob and os do not depend on the SCons-injected env, so importing them at
the top of the file removes the flake8 E402 findings and lets the E402
ignore-all directives (which were no longer suppressing anything) go away.
2026-07-26 05:32:54 +00:00
JorropoandGitHub e700cc7a8c lint: drop redundant bandit ignores in c3 exception decoder (#11173)
B404 and B603 are already skipped globally in .trunk/configs/.bandit, so
the file-level ignore-all directives never suppressed anything.
2026-07-26 05:32:05 +00:00
JorropoandGitHub 6111a54bcc lint: drop redundant quotes in .coderabbit.yaml (#11168)
yamllint quoted-strings (only-when-needed) flags the plain-scalar values
".*" and "bin/config.d/**", which do not require quoting.
2026-07-26 04:52:46 +00:00
8e104a909b fix(admin): persist TAK module config (team color / member role) (#11216)
Setting TAK team/role ACKed and rebooted but stored nothing:
handleSetModuleConfig had no tak case, saveToDisk never set has_tak,
and handleGetModuleConfig had no TAK_CONFIG case. Add all three, plus
a native suite sweeping every ModuleConfig submessage through
set -> save -> load -> get and a TAK value-fidelity suite.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 23:03:06 +00:00
AustinandGitHub 6d944e44d0 Switch to forked pioarduino platform (with penv fixes for CI issues) (#11219) 2026-07-25 15:31:56 -05:00
b9a5443480 Fix multicast address (#8612)
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
Co-authored-by: Jason P <applewiz@mac.com>
2026-07-25 11:54:39 +00:00
JorropoandGitHub 9e529da460 replace screens with unique_ptr (#11163) 2026-07-25 11:51:13 +00:00
abf217c38d fix(mesh): don't assert on malloc() failure in MemoryDynamic::alloc() (#11197)
* fix(mesh): don't assert on malloc() failure in MemoryDynamic::alloc()

MemoryDynamic<T>::alloc() called assert(p) right after malloc(), instead
of returning nullptr like the static MemoryPool<T,N>::alloc() already
does on exhaustion. All of packetPool's callers were already hardened to
null-check allocCopy()/allocZeroed() (#10948, #10951), but that path is
unreachable on real OOM here: assert() fires first, one level down,
before the caller ever gets a chance to check anything.

On STM32WL (packetPool is MemoryDynamic there - not enough static RAM
for the fixed pool), assert() failures are wrapped to an infinite
`while(true);` loop rather than aborting or resetting, so the thread
just hangs forever instead of returning nullptr. Reproduced on wio-e5
hardware under a burst of incoming DMs, with free heap draining toward
zero right before the hang.

Assisted-by: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Andrew Yong <me@ndoo.sg>

* fix: null-check the two allocUnique*() call sites missed by prior audits

#10948/#10951 null-checked every raw-pointer packetPool.allocCopy()/
allocZeroed() call site, but missed the two spots using the
UniquePacketPoolPacket (unique_ptr) wrapper: UdpMulticastHandler::onReceive()
and MQTT::onReceive() both dereferenced the allocation result unconditionally.

Previously unreachable in practice: MemoryDynamic::alloc() asserted before
ever returning null, so nothing downstream saw it. Fixed alongside that
assert removal (this branch) since these two are now reachable with a
genuine null - same fix as everywhere else, just an unchecked pointer that
turns up under OOM.

Assisted-by: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Andrew Yong <me@ndoo.sg>

* fix(SimRadio): don't leave isReceiving stuck true on allocation failure

startReceive() set isReceiving = true before allocCopy(), so a failed
allocation (now reachable with a real nullptr instead of hanging in
assert()) left the simulated radio permanently marked as receiving:
handleReceiveInterrupt() returns immediately on a null receivingPacket,
before ever reaching the code that would clear isReceiving.

Assisted-by: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Andrew Yong <me@ndoo.sg>

---------

Signed-off-by: Andrew Yong <me@ndoo.sg>
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
2026-07-25 11:28:42 +00:00
Ben MeadorsandGitHub 3f4e7cc622 Fix nRF52 freeze + watchdog reset when saving config over BLE (#11202)
* Fix nRF52 freeze + watchdog reset when saving config over BLE

Since #10967 phone-originated admin messages are handled synchronously on
Bluefruit's BLE event task. NRF52Bluetooth::disconnect() busy-waited for
BLE_GAP_EVT_DISCONNECTED, which only that same task can process, so any
config save that requires a reboot (e.g. position) deadlocked the device
until the 90s watchdog fired. Bound the wait to 1s and sleep instead of
spinning so lower-priority tasks (including the watchdog feed) keep
running; the SoftDevice completes the link termination on its own.

* Address review: use Throttle helper, tighten comment

* Name the disconnect timeout constant

* Log unconfirmed BLE disconnect at WARN with elapsed time
2026-07-25 11:20:04 +00:00
Thomas GöttgensandGitHub b46c8c9f80 Router: defer nested local sends to avoid handleReceived re-entrancy (#11191)
Prevents an nRF52 task-stack overflow on config save. Replaces draft #11155.
2026-07-24 19:30:48 +00:00
JorropoandGitHub f4f35bca93 lint: drop dead CKV2_GHA_1 ignore in main_matrix workflow (#11170)
The trunk-ignore no longer matches any finding (the job sets its own
permissions), so it only produced an ignore-does-nothing note.
2026-07-24 16:45:01 +00:00
renovate[bot]GitHubrenovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
36fdbddd57 Update LovyanGFX to v1.2.26 (#11157)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-07-24 14:31:29 +00:00
AustinandGitHub 5b7ee514ac Actions: Also cache wasm (#11200) 2026-07-24 14:15:45 -04:00
AustinandGitHub e50572b7c0 Save/Restore seperately in MacOS/Windows caching workflows. (#11199)
Only *save* cache on the default branch (develop), restore it everywhere it fits
2026-07-24 13:47:27 -04:00
AustinandGitHub 88e287f778 Remove docker buildkit layer caching (for now) (#11195)
This is taking up all our cache space, we need room for activities!
2026-07-24 12:28:09 -04:00
AustinandGitHub 0ef375dff4 Actions: Add caching for MacOS and Windows build workflows (#11194) 2026-07-24 12:09:12 -04:00
vidplace7 7754068582 Only build 2 docker images upon PR
Only build debian and alpine
2026-07-24 08:28:10 -04:00
Ben MeadorsandGitHub becfb770c1 nrf52: fix BLE-task stack overflow crashing pairing and wiping LittleFS (#11190)
Since #10967 made Router::sendLocal handle self-addressed packets
synchronously, the entire phone-API chain for a BLE client runs inline in
the Bluefruit characteristic write callback: toRadioWriteCb ->
PhoneAPI::handleToRadio -> admin set-config -> radio reconfigure ->
NodeDB::saveToDisk. That callback executes on the Bluefruit BLE FreeRTOS
task, whose stock stack is 5 KB (CFG_BLE_TASK_STACKSIZE = 256*5 words) -
not the Arduino loop task that #10944 already raised to 8 KB. The loop-task
fix therefore protects the wrong task for BLE-originated writes.

On a Seeed Wio Tracker L1 the 5 KB stack overflows during pairing
first-sync, resetting the device mid-LittleFS-write, every single time.
Repeated mid-write resets tear the LittleFS metadata, lfs_assert fires on
the next boot, and the corruption handler formats the whole filesystem:
region, channels, module config, and the node's keypair are all lost
(critical fault #13, new node identity on next region set). Reproduced
end-to-end tonight on stock develop 6908d27; with this change the same
device pairs, serves config screens, and survives back-to-back
config.proto saves over BLE.

Raise the BLE task to the same 2048 words (8 KB) as LOOP_STACK_SZ, for the
same reason. bluefruit.cpp's #ifndef guard makes the -D take effect with no
framework patch. Costs 3 KB of RAM on nrf52840 targets only.

Credit where due: Ixitxachitl independently established in #11155 testing
that the save-path crash persists after #11185 and that re-queueing
sendLocal (moving the pipeline back to the Router thread) makes it go away
- which corroborates this diagnosis from the other direction. This commit
is the minimal capacity-side fix; #11155's relocation of the pipeline off
the BLE task remains the right architectural follow-up, and this guard
stays correct even after it lands.

Likely also explains #10905 (L1 display-thread crash when a client
requests full configuration) and the 2.8 field reports of idle nodes
losing region and keys after a BLE session.
2026-07-24 06:38:46 -05:00
Ben MeadorsandGitHub 01d1873bfc Deliver locally-generated replies addressed to us to the phone (#11185)
* Deliver locally-generated replies addressed to us to the phone

Config get/set from the phone times out on every device: the client sends an
admin request, the node handles it, and the response is silently dropped
before it reaches the phone queue.

#10967 changed Router::sendLocal's isToUs branch from enqueueReceivedMessage()
to handleReceived(p, src), so a local packet keeps its RxSource instead of
being relabeled RX_SRC_RADIO by the queue round-trip. That is the right call
for the new policy gates, but module replies go out through
MeshService::sendToMesh() with the default RX_SRC_LOCAL, and a reply to a
phone-originated request is addressed to our own node (setReplyTo resolves
from == 0 to ourNodeNum). Those replies now re-enter callModules as
RX_SRC_LOCAL, where the loopback gate skips every module whose loopbackOk is
false - including RoutingModule, whose promiscuous sniff is the only path that
moves a received packet into toPhoneQueue. The reply is released, never sent.

Requests still work, because the phone's own packets arrive as RX_SRC_USER and
pass the gate, so a set_config is applied and only its acknowledgement is lost.
That is why a client can connect and download config but times out on every
config screen and every setter.

Deliver the phone's copy from sendToMesh() instead: for a local packet
addressed to us, the loopback gate is doing its job in keeping the packet away
from module re-dispatch, and the phone copy is exactly what is missing. Setting
loopbackOk on RoutingModule would instead echo every locally-generated
broadcast back to the phone, and relabeling replies RX_SRC_RADIO would undo the
origin separation #10967 added.

Also stop reporting ERRNO_SHOULD_RELEASE (35) to the phone in the QueueStatus
for these packets. It means "caller frees", not a send failure, and the same
hunk changed it from the 0 the phone used to see.

* Address review: trim comments, assert the QueueStatus count

Condense the added comments to the one-or-two-line house style; the rationale
lives in the commit message and PR.

The reply test drained QueueStatus records in a while loop, which would have
passed just as happily on an empty queue. Count them and require both the
request's and the reply's.
2026-07-24 09:02:51 +00:00
JorropoandGitHub 4c17dda3e6 lint: drop dead terrascan ignore in devcontainer Dockerfile (#11167)
terrascan is not an enabled linter, so the trunk-ignore-all directive only
produced a "disabled linter" note. Remove it.
2026-07-24 08:57:16 +00:00
JorropoandGitHub ee74c3c446 lint: suppress intentional root USER in runtime Dockerfiles (#11166)
meshtasticd must run as root to reach GPIO/I2C/SPI hardware, so the final
USER root is deliberate. The existing file-level ignore used the old trivy
code (DS002); rename it to DS-0002 and add checkov/CKV_DOCKER_8 so the
last-USER-root findings are suppressed with an explicit justification.
2026-07-24 08:28:21 +00:00
JorropoandGitHub 419e8ee11f lint: fix stale trivy rule codes in clusterfuzzlite Dockerfile (#11165)
The trunk upgrade (#11140) renamed trivy's Dockerfile rules from DSxxx to
DS-xxxx, so the trunk-ignore-all directives for DS026/DS002 stopped
suppressing and the HEALTHCHECK / non-root-USER findings resurfaced.
Update the codes to DS-0026 / DS-0002, and drop the checkov/CKV_DOCKER_8,
hadolint/DL3002 and disabled terrascan ignores that no longer match any
finding.
2026-07-24 07:45:06 +00:00
JorropoandGitHub 79ade464a5 lint: clean up pr_tests workflow lint findings (#11169)
Suppress checkov/CKV_GHA_7 (the workflow_dispatch reason input is a
free-text run label that never reaches the build) and drop the redundant
quotes yamllint flags on the description/default/name scalars.
2026-07-24 07:38:18 +00:00
Ben MeadorsandGitHub b08c04f264 Add Meshnology W12 (WiFi LoRa 32 V5) LR2021 board support (#10910)
* Add Meshnology W12 (WiFi LoRa 32 V5) variant with LR2021 radio

ESP32-S3R8 + 16 MB flash + Semtech LR2021 dual-band LoRa + SSD1315 OLED,
Heltec V3-style pinout.

Pin map comes from the vendor pin sheet and vendor Arduino demos, confirmed
on hardware: I2C scan finds the OLED at 0x3C on SDA17/SCL18, and the LR2021
answers on NSS=8 SCK=9 MISO=11 MOSI=10 with BUSY=13/NRESET=12 verified by
reset-pulse probing.

The IRQ line needs LR2021_IRQ_DIO_NUM 8. RadioLib defaults irqDioNum to 5,
but DIO5 is not bonded out on this board - the two IRQ lines are DIO7 -> IO7
and DIO8 -> IO14, established by driving each radio DIO as a GPIO output and
watching which ESP32 pin followed. Left at the default, every interrupt is
routed to a floating pin: RX/TX still appear to work because
RadioLibInterface::pollMissedIrqs() falls back to polling the IRQ status
register, but the blocking scanChannel() waits on the pin itself and never
returns.

A shadow pins_arduino.h is needed because the generic esp32s3 one defines
RGB_BUILTIN, which pulls the Arduino RMT HAL into the link and fails against
this build's FreeRTOS config.

* Address review: add W12 variant metadata, trim comments

Declares the custom_meshtastic_* metadata the build manifest expects. Without
it the emitted .mt.json carried no hwModel, slug, display name or support
level, so the board was invisible to the flasher - and board_level = extra
kept it out of PR builds entirely, which is why it never appeared in this
PR's board list. It now matches the W10 sibling at board_level = pr with
support level 1. hwModel stays 255/PRIVATE_HW until a W12 enum lands in the
protobufs.

Also trims the header comments to the one-or-two-line limit in AGENTS.md;
the vendor pin sheet and probing evidence live in the PR description.
2026-07-24 03:29:44 +00:00
AustinandGitHub 372b751dc8 Actions: Run --level pr on merge_group (for now) (#11186)
Other branches aren't set up for this (yet?) so go back to the old behaviour for now.
2026-07-23 22:41:38 -04:00
AustinandGitHub 6908d27660 Enable merge queues (#11162) 2026-07-23 19:02:57 +00:00
renovate[bot]GitHubrenovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
d0d029c8b1 Update python to v3.14.6 (#11063)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-07-23 13:59:36 +02:00
Ben Meadors 876962e8f2 Protos 2026-07-23 06:49:09 -05:00
Ben MeadorsandGitHub d9150e8dc7 ci: warn and skip the pio check job on transient network failures (#11145)
* ci: warn and skip the pio check job on transient network failures

The `check` matrix job intermittently fails before it analyses anything:
`pio check` has to fetch the platform package first, and that download gets
reset mid-flight by the CDN often enough to be a nuisance. Three of the last
~200 CI runs died this way, all identical and all ~1s into the run:

    Checking station-g3 > cppcheck (board: station-g3; platform:
    https://github.com/pioarduino/platform-espressif32/.../platform-espressif32.zip)
    requests.exceptions.ConnectionError: ('Connection aborted.',
        ConnectionResetError(104, 'Connection reset by peer'))

cppcheck never ran, so there is no signal at all - just a red X someone has to
re-run by hand.

`pio check` exits 1 for both a dead socket and a real defect, so only the output
can tell them apart. check-all.sh now tees the run and classifies the failure:

  * cppcheck reported something (summary table or a `file:line: [severity]`
    defect line) -> real, propagate the exit status. This is checked first, so
    no amount of network noise elsewhere in the log can mask a genuine defect.
  * transport-level error and nothing else -> log a ::warning:: annotation
    naming the board and the underlying error, and exit 0 under CI (exit 2
    locally, so a local run never quietly reports itself clean).
  * anything else -> propagate the exit status.

404 / 403 / UnknownPackageError are deliberately not treated as transient: a
missing or forbidden package is a reproducible break from a bad pin in a .ini,
not weather.

No workflow change is needed - gh-action-firmware's entrypoint runs
/workspace/bin/check-all.sh for MT_TARGET=check, so the whole fix lives in this
repo and applies to local runs too.

Verified by replaying the captured CI logs through the script with a stubbed
`pio`: the station-g3 connection-reset log skips with a warning under CI and
exits 2 locally, the rak11200 `Total 1 0 0` defect log still fails, that same
defect log with the reset traceback appended still fails, and a 404 still fails.

* ci: pass board names to pio as an argument array

Addresses CodeRabbit review feedback on the previous commit (and the SC2086 the
line has carried since it was written): BOARDS was a space-joined string and
$CHECK was expanded unquoted, so a board override containing whitespace or a
glob character could turn into extra pio arguments.

BOARDS and CHECK are arrays now, and pio is invoked with "${CHECK[@]}", so board
names reach it as literal arguments. The two skip messages use "${BOARDS[*]}" -
a bare "${BOARDS}" would have quietly named only the first board.

shellcheck -x is clean on the file. Re-verified with a stubbed pio that records
its argv: 'tb*' and 'tbeam extra' each arrive as one literal argument, the
default list still expands to 15 -e pairs, a two-board skip names both boards,
and all six classification cases (network/CI, network/local, real defects,
defects with a reset traceback appended, 404, clean pass) are unchanged.
2026-07-23 06:25:18 -05:00
ef1aedd569 Fixes a crash when you favorite a node (#11146)
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
2026-07-23 10:08:34 +02:00
efc9708236 Add userpref for the packet signature policy default (#11144)
Co-authored-by: Austin <vidplace7@gmail.com>
2026-07-23 10:05:02 +02:00
Carlos ValdesandGitHub 2c1fd6b7ea fix(ScanI2C): fetch the SE050 probe response with requestFrom() (#11133) 2026-07-23 09:57:14 +02:00
Ben MeadorsandGitHub 8a9980576f Merge pull request #11154 from RCGV1/codex/fix-packet-policy-excluded-test
test: cover packet policy coercion on excluded builds
2026-07-22 15:52:27 -05:00
Benjamin Faershtein 8acb7fec73 test: cover coerced packet signature policy 2026-07-22 12:25:19 -07:00
Ben MeadorsandGitHub 1804fd1886 Merge pull request #11150 from meshtastic/always-exclude-rangetest
Always exclude range test
2026-07-22 12:20:06 -05:00
Ben Meadors cad3a3a5c7 Always exclude range test
Range test is excluded on every target as of 2.8, so set
MESHTASTIC_EXCLUDE_RANGETEST=1 globally in [env] build_flags and drop the
now-redundant per-variant flags from esp32, rak11200 and stm32.

Guard RangeTestModule.cpp with the same condition used in Modules.cpp so the
translation unit compiles to nothing rather than relying on --gc-sections to
strip it, and always report RANGETEST_CONFIG in the device metadata
excluded_modules bitmask so clients hide the config.
2026-07-22 10:25:54 -05:00
Ben MeadorsandGitHub 9662c818b5 Merge pull request #11148 from NomDeTom/flash-clawback
claw back some space from signing policies
2026-07-22 08:54:23 -05:00
nomdetom 583248c19d claw back some space from signing policies 2026-07-22 14:36:06 +01:00
Ben MeadorsandGitHub f90c1340de Merge pull request #11117 from NomDeTom/nailed-test-number
Nailed native test numbers so PRs don't forget to increment
2026-07-22 06:06:08 -05:00
Thomas GöttgensandGitHub b9c7858fd0 enable autoLDRO on LR2021 (#11139) 2026-07-22 11:12:52 +02:00
Thomas GöttgensandGitHub f09f2b1366 Gate the PKI-decrypt key on signer-proven, not the opportunistic cache (#11116) 2026-07-22 11:12:06 +02:00
Thomas GöttgensandGitHub 11c510ff1b WIP: Big Display Node (#10926) 2026-07-22 09:35:36 +02:00
Ben MeadorsandGitHub c263a9e4a3 Merge pull request #11136 from meshtastic/renovate/rak13800-w5100s-1.x
Update RAK13800-W5100S to v1.0.3
2026-07-21 19:40:27 -05:00
Ben MeadorsandGitHub 1d6f8a8f34 Merge pull request #11130 from meshtastic/ci-native-tests-by-area
Run native PlatformIO tests by area for readable failures
2026-07-21 19:39:58 -05:00
Ben MeadorsandGitHub 5e3d1ea606 Merge pull request #11129 from RCGV1/agent/include-statusmessage-config
phoneapi: send status message config
2026-07-21 19:37:35 -05:00