Files
22072c5f4b Pr1.5 tmm nexthop (#10745)
* TrafficManagement: flat unified cache + persistent next-hop overflow store

Reworks the TrafficManagementModule cache layer (policing behaviour unchanged
from upstream) and adds a routing-hint overflow store:

- Flatten the ring: replace the cuckoo-hashed unified cache and the bucketed
  PSRAM NodeInfo index with plain flat arrays + linear scan (same idiom as
  WarmNodeStore). At LoRa packet rates an O(n) scan of the cache is negligible,
  and it removes a large amount of hashing/displacement complexity. The cache
  entry is 11 B; timestamps use a uniform +1 presence-offset so a 0 byte always
  means "empty" across every sub-store. Adds rebaseEpoch() so cached state
  survives the ~19 h relative-timestamp horizon instead of being flushed.

- Next-hop overflow cache: setNextHop/getNextHopHint store a confirmed last-byte
  relay for a destination, written only from NextHopRouter's ACK-confirmed
  decision (and mirrored from TraceRoute). NextHopRouter::getNextHop falls back
  to this cache when the hot NodeDB has no hint, so DMs/relays to long-tail
  nodes keep routing after the node ages out of NodeInfoLite.

- Persistence: preloadNextHopsFromNodeDB warm-starts the cache from persisted
  NodeInfoLite hints on first maintenance pass; next_hop entries are kept alive
  across the maintenance sweep (no TTL) and never clobbered by a stale preload.

All packet-policing logic (rate limit, position dedup, unknown-packet drop,
NodeInfo direct response, hop exhaustion) is the existing upstream behaviour,
untouched. HAS_TRAFFIC_MANAGEMENT defaults on so the module is compiled in. (see note).

Tests: upstream policing suite now actually runs (adds the MeshTypes.h include
that gates HAS_TRAFFIC_MANAGEMENT) plus 4 next-hop tests. Role-aware throttles,
politeness, precision clamp, port-interval and mesh-radius gating — and the
rate-limit >255 saturation fix — are deferred to the advanced-TMM branch.

Note: default dedup movement grid moves to ~91m, which also means 1.5km required to end up with the same signature position - coarser and therefore further than before.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* TrafficManagement: fix cppcheck constVariablePointer warning

`node` in preloadNextHopsFromNodeDB() is never written through — mark
it const to satisfy cppcheck's constVariablePointer check in CI.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Add multi-hop NextHop recovery tests and unit tests for routing reliability

- Introduced a new test suite for multi-hop NextHop directed-message delivery and relay recovery in `test_nexthop_multihop_recovery.py`. This includes tests for end-to-end delivery and recovery after relay drop.
- Implemented unit tests in `test_main.cpp` for NextHop routing reliability mitigations, covering:
  - M1: Ambiguity-aware last-byte resolution.
  - M2: NextHopRouter's strict-neighbor gate and hop limit checks.
  - M3: Route-health freshness and failure decay.
- Enhanced mock classes to facilitate controlled testing of node behaviors and routing logic.

* grafting fixed

* Address Copilot review for PR #10735 (NextHop improvements)

- docs/nexthop-routing-reliability.md: update status from "no code
  changes yet" to reflect that mitigations and tests are implemented

RAM pressure and MIGRATION_VERBOSE concerns addressed upstream in
PR2.5 (per-platform TRAFFIC_MANAGEMENT_CACHE_SIZE) and PR2 (verbose
default=0) respectively; (0,0) sentinel fixed in PR2.5.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* CI: fix cppcheck constVariablePointer and test include path

- NextHopRouter.cpp: qualify two RouteHealth *h locals as const — only
  read for stale-route checks, never mutated through the pointer
- Router.cpp: qualify meshtastic_NodeInfoLite *node as const in
  shouldDecrementHopLimit — only read for favorite/role predicate
- test_position_module/test_main.cpp: change bare PositionModule.h to
  modules/PositionModule.h — build_flags sets -Isrc, not -Isrc/modules,
  so the bare form fails to resolve in the native PlatformIO test env

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* WarmStore: cache device role + protected category in last_heard low bits

Steal the low 6 bits of WarmNodeEntry.last_heard to carry an evicted node's
device role (4 bits) and a protected category (2 bits) for the hop-trim path,
at zero record-size cost (entry stays 40 B; no RAM/flash growth). The high bits
remain a real unix-seconds timestamp, quantised to 64 s — ample for warm LRU
ordering of long-tail nodes.

- absorb() packs role/protectedCat; place()/ring replay store the raw word so
  metadata round-trips through flash. LRU compares masked time (warmTimeOf).
- take() rehydration masks the metadata bits and restores the cached role so a
  re-admitted node isn't stuck at CLIENT until its next NodeInfo.
- NodeDB classifies the category (favorite/ignored/verified -> Flag;
  tracker/sensor/tak_tracker -> Role) at each eviction site.
- WarmNodeStore::lookupMeta() exposes role/category to consumers.
- Bump WARM_RING_MAGIC (WRNG->WRN2): old rings read as erased and rebuild;
  warm data is a non-critical evictee cache, so discard-on-upgrade is safe.

Tests: test_warm_store 11/11 (new meta round-trip + quantisation-aware ordering);
NodeDB compiles (test_nodedb_blocked 4/4).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WarmStore: migrate v1 rings/files by discarding last_heard, not the data

Previously the WRNG->WRN2 magic bump treated old rings as erased, discarding all
warm entries — including the PKI public keys that let evicted nodes keep
decrypting DMs. Instead, read v1 (WRNG / WRM1) records and keep each node's
identity + public key, discarding only last_heard (its low bits would otherwise
be misread as the new role/protected metadata). Records re-rank and re-learn
their role on next contact.

- Ring backend (nRF52840): ringReadHeader accepts both magics and reports v1 via
  an out-param; replay zeroes last_heard for v1 records. If the active head page
  is v1, force a rotation so new v2 records never land in a v1-headered page
  (which would discard their freshly-set role on the next load). Legacy pages
  convert to v2 as the ring rotates.
- File backend (warm.dat): bump WARM_STORE_MAGIC WRM1->WRM2; accept WRM1, verify
  CRC against the stored bytes, then discard last_heard and mark dirty so the
  next save rewrites as v2.

Tests: test_warm_store 12/12 (adds test_ws_v1_migration_discardsLastHeard:
key survives, role/protected reset).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WarmStore: guard role bit-width + test eviction carries role/protected

- static_assert that the device role enum still fits the 4-bit warm metadata
  field (WARM_ROLE_MASK); fails the build loudly if a new role is added past 15
  rather than silently truncating role on eviction. (Max role today = 12.)
- Add test_migration_carriesRoleAndProtectedIntoWarm: a demoted TRACKER lands in
  the warm tier with its key, role=TRACKER and protected category=Role; a demoted
  CLIENT carries role=CLIENT/None. Exercises the NodeDB eviction path +
  warmProtectedCategory classification (the warm-store unit tests only cover
  absorb() directly).

Tests: test_nodedb_blocked 5/5.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix copilot comments

* fix(test): restore #if HAS_TRAFFIC_MANAGEMENT guard in TMM test

The rebase onto PR1.5 lost the top-level HAS_TRAFFIC_MANAGEMENT guard
that PR1.5 introduced, leaving the #else/#endif tail orphaned and
causing compile errors on non-TMM builds.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
2026-06-19 19:52:58 -05:00
..
2026-06-19 19:52:58 -05:00

Meshtastic MCP Server — Test Harness

Automated test suite for the MCP server, organized around real operator concerns rather than generic "unit vs hardware".

Tiers

Dir Hardware Question this tier answers
unit/ none Do the parsing / filtering / profile-generation primitives work?
provisioning/ 1 device, per-test bake Did my pre-bake recipe stick? Does it survive a factory reset?
admin/ 1 device, shared bake Do my daily admin ops (owner, channel URL, config writes) round-trip?
mesh/ 2 devices, shared bake Do my devices actually form a mesh? Send + receive? ACKs?
telemetry/ 2 devices, shared bake Is telemetry reporting? Is position broadcast correct?
monitor/ 1 device, shared bake Is the boot log clean (no panics)?
fleet/ varies Are my CI runs isolated from each other? Are reflashes idempotent?

Quick start

cd mcp-server
pip install -e ".[test]"

# No hardware — 33 unit tests, ~3 seconds
pytest tests/unit -v

# Hub attached (nRF52840 + ESP32-S3) — first run bakes, then exercises everything
pytest tests/ --html=report.html

# Hub already baked with session profile (dev loop) — skip bake
pytest tests/ --assume-baked --html=report.html

# Force a rebake (new firmware, new seed, etc.)
pytest tests/ --force-bake --html=report.html

CLI flags

  • --force-bake — always reflash both roles at session start, even if the current state matches the session profile.
  • --assume-baked — skip test_00_bake.py entirely. Use when you know the devices are already baked and want a fast dev loop.
  • --hub-profile=<yaml> — point at a YAML file for non-default hub hardware. Default targets VID 0x239a (nRF52) and 0x303a/0x10c4 (ESP32-S3).
  • --no-teardown-rebake — skip the session-end rebake that provisioning/ and fleet/ tests perform. Useful in rapid iteration.

Environment variables

  • MESHTASTIC_FIRMWARE_ROOT — firmware repo path (defaults to ../ from tests/)
  • MESHTASTIC_MCP_ENV_NRF52 — PlatformIO env for the nRF52 role (default rak4631)
  • MESHTASTIC_MCP_ENV_ESP32S3 — PlatformIO env for the ESP32-S3 role (default heltec-v3)
  • MESHTASTIC_MCP_SEED — override the session PSK seed (default: pytest-<unix-ts>). Set this to reproduce a specific failing run.

Fixtures you'll use when adding tests

All defined in conftest.py:

  • hub_devices{"nrf52": "/dev/cu.X", "esp32s3": "/dev/cu.Y"}. Auto- skips the test if a required role isn't present.
  • test_profile → USERPREFS dict for the session (build_testing_profile).
  • no_region_profile → variant without USERPREFS_CONFIG_LORA_REGION.
  • baked_mesh → verifies both devices are baked with the session profile (does NOT reflash — that's test_00_bake.py's job).
  • baked_single → single verified baked device; parametrize request.param to pick role.
  • serial_capture → factory; cap = serial_capture("esp32s3") starts a pio device monitor session, drains into a per-test buffer, attaches the buffer to the pytest-html report on failure.
  • wait_until → exponential-backoff polling helper; wait_until(lambda: predicate(), timeout=60) replaces flaky time.sleep() patterns.

Reports

pytest --html=report.html produces a self-contained HTML with:

  • Per-test pass/fail/skip with timings
  • On failure: serial log capture from any serial_capture fixture used
  • On failure: device_info + lora config JSON for every role on the hub
  • Session seed and session start time (for reproducibility)

pytest --junitxml=junit.xml produces CI-integration XML.

tool_coverage.json is emitted at session end in the tests directory — shows which of the 38 MCP tools the run exercised. Useful for closing test gaps.

Adding a new test

  1. Pick the category that matches the operator concern (not the technical surface). "Does my fleet's owner name persist" is admin/, not unit/.
  2. If you need both devices, depend on baked_mesh. If you need one, depend on baked_single. If you need to mutate hardware state, put it in provisioning/ or fleet/ and add a try/finally teardown that re-bakes the session profile.
  3. Use wait_until for anything involving LoRa timing — fixed sleep() produces flakes.
  4. Use serial_capture when you need to observe firmware log output (e.g. "did the packet get decoded?").
  5. Add a @pytest.mark.timeout(N) — mesh tests routinely hit LoRa-airtime waits; default pytest timeout is infinite.

Troubleshooting

  • All hardware tests SKIP → hub not detected. Plug in the USB hub, verify with pytest tests/ --collect-only or python -c "from meshtastic_mcp import devices; print(devices.list_devices())".
  • baked_mesh fails with "devices not baked" → run pytest tests/test_00_bake.py first, or pass --force-bake on the full run.
  • Mesh formation tests time out → check that both devices are on the same session profile (--force-bake forces both to the current seed).
  • Provisioning tests leave device in bad state → teardowns re-bake, but if a test crashes between "bake broken state" and "bake good state", run pytest tests/test_00_bake.py --force-bake to recover.