* TrafficManagement: flat unified cache + persistent next-hop overflow store Reworks the TrafficManagementModule cache layer (policing behaviour unchanged from upstream) and adds a routing-hint overflow store: - Flatten the ring: replace the cuckoo-hashed unified cache and the bucketed PSRAM NodeInfo index with plain flat arrays + linear scan (same idiom as WarmNodeStore). At LoRa packet rates an O(n) scan of the cache is negligible, and it removes a large amount of hashing/displacement complexity. The cache entry is 11 B; timestamps use a uniform +1 presence-offset so a 0 byte always means "empty" across every sub-store. Adds rebaseEpoch() so cached state survives the ~19 h relative-timestamp horizon instead of being flushed. - Next-hop overflow cache: setNextHop/getNextHopHint store a confirmed last-byte relay for a destination, written only from NextHopRouter's ACK-confirmed decision (and mirrored from TraceRoute). NextHopRouter::getNextHop falls back to this cache when the hot NodeDB has no hint, so DMs/relays to long-tail nodes keep routing after the node ages out of NodeInfoLite. - Persistence: preloadNextHopsFromNodeDB warm-starts the cache from persisted NodeInfoLite hints on first maintenance pass; next_hop entries are kept alive across the maintenance sweep (no TTL) and never clobbered by a stale preload. All packet-policing logic (rate limit, position dedup, unknown-packet drop, NodeInfo direct response, hop exhaustion) is the existing upstream behaviour, untouched. HAS_TRAFFIC_MANAGEMENT defaults on so the module is compiled in. (see note). Tests: upstream policing suite now actually runs (adds the MeshTypes.h include that gates HAS_TRAFFIC_MANAGEMENT) plus 4 next-hop tests. Role-aware throttles, politeness, precision clamp, port-interval and mesh-radius gating — and the rate-limit >255 saturation fix — are deferred to the advanced-TMM branch. Note: default dedup movement grid moves to ~91m, which also means 1.5km required to end up with the same signature position - coarser and therefore further than before. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * TrafficManagement: fix cppcheck constVariablePointer warning `node` in preloadNextHopsFromNodeDB() is never written through — mark it const to satisfy cppcheck's constVariablePointer check in CI. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * Add multi-hop NextHop recovery tests and unit tests for routing reliability - Introduced a new test suite for multi-hop NextHop directed-message delivery and relay recovery in `test_nexthop_multihop_recovery.py`. This includes tests for end-to-end delivery and recovery after relay drop. - Implemented unit tests in `test_main.cpp` for NextHop routing reliability mitigations, covering: - M1: Ambiguity-aware last-byte resolution. - M2: NextHopRouter's strict-neighbor gate and hop limit checks. - M3: Route-health freshness and failure decay. - Enhanced mock classes to facilitate controlled testing of node behaviors and routing logic. * grafting fixed * Address Copilot review for PR #10735 (NextHop improvements) - docs/nexthop-routing-reliability.md: update status from "no code changes yet" to reflect that mitigations and tests are implemented RAM pressure and MIGRATION_VERBOSE concerns addressed upstream in PR2.5 (per-platform TRAFFIC_MANAGEMENT_CACHE_SIZE) and PR2 (verbose default=0) respectively; (0,0) sentinel fixed in PR2.5. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * CI: fix cppcheck constVariablePointer and test include path - NextHopRouter.cpp: qualify two RouteHealth *h locals as const — only read for stale-route checks, never mutated through the pointer - Router.cpp: qualify meshtastic_NodeInfoLite *node as const in shouldDecrementHopLimit — only read for favorite/role predicate - test_position_module/test_main.cpp: change bare PositionModule.h to modules/PositionModule.h — build_flags sets -Isrc, not -Isrc/modules, so the bare form fails to resolve in the native PlatformIO test env Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * WarmStore: cache device role + protected category in last_heard low bits Steal the low 6 bits of WarmNodeEntry.last_heard to carry an evicted node's device role (4 bits) and a protected category (2 bits) for the hop-trim path, at zero record-size cost (entry stays 40 B; no RAM/flash growth). The high bits remain a real unix-seconds timestamp, quantised to 64 s — ample for warm LRU ordering of long-tail nodes. - absorb() packs role/protectedCat; place()/ring replay store the raw word so metadata round-trips through flash. LRU compares masked time (warmTimeOf). - take() rehydration masks the metadata bits and restores the cached role so a re-admitted node isn't stuck at CLIENT until its next NodeInfo. - NodeDB classifies the category (favorite/ignored/verified -> Flag; tracker/sensor/tak_tracker -> Role) at each eviction site. - WarmNodeStore::lookupMeta() exposes role/category to consumers. - Bump WARM_RING_MAGIC (WRNG->WRN2): old rings read as erased and rebuild; warm data is a non-critical evictee cache, so discard-on-upgrade is safe. Tests: test_warm_store 11/11 (new meta round-trip + quantisation-aware ordering); NodeDB compiles (test_nodedb_blocked 4/4). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * WarmStore: migrate v1 rings/files by discarding last_heard, not the data Previously the WRNG->WRN2 magic bump treated old rings as erased, discarding all warm entries — including the PKI public keys that let evicted nodes keep decrypting DMs. Instead, read v1 (WRNG / WRM1) records and keep each node's identity + public key, discarding only last_heard (its low bits would otherwise be misread as the new role/protected metadata). Records re-rank and re-learn their role on next contact. - Ring backend (nRF52840): ringReadHeader accepts both magics and reports v1 via an out-param; replay zeroes last_heard for v1 records. If the active head page is v1, force a rotation so new v2 records never land in a v1-headered page (which would discard their freshly-set role on the next load). Legacy pages convert to v2 as the ring rotates. - File backend (warm.dat): bump WARM_STORE_MAGIC WRM1->WRM2; accept WRM1, verify CRC against the stored bytes, then discard last_heard and mark dirty so the next save rewrites as v2. Tests: test_warm_store 12/12 (adds test_ws_v1_migration_discardsLastHeard: key survives, role/protected reset). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * WarmStore: guard role bit-width + test eviction carries role/protected - static_assert that the device role enum still fits the 4-bit warm metadata field (WARM_ROLE_MASK); fails the build loudly if a new role is added past 15 rather than silently truncating role on eviction. (Max role today = 12.) - Add test_migration_carriesRoleAndProtectedIntoWarm: a demoted TRACKER lands in the warm tier with its key, role=TRACKER and protected category=Role; a demoted CLIENT carries role=CLIENT/None. Exercises the NodeDB eviction path + warmProtectedCategory classification (the warm-store unit tests only cover absorb() directly). Tests: test_nodedb_blocked 5/5. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix copilot comments * fix(test): restore #if HAS_TRAFFIC_MANAGEMENT guard in TMM test The rebase onto PR1.5 lost the top-level HAS_TRAFFIC_MANAGEMENT guard that PR1.5 introduced, leaving the #else/#endif tail orphaned and causing compile errors on non-TMM builds. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Ben Meadors <benmmeadors@gmail.com>
122 lines
6.0 KiB
C++
122 lines
6.0 KiB
C++
#pragma once
|
||
#include <MeshRadio.h>
|
||
#include <NodeDB.h>
|
||
#include <RadioInterface.h>
|
||
#include <cmath>
|
||
#include <cstdint>
|
||
#include <meshUtils.h>
|
||
#define ONE_DAY 24 * 60 * 60
|
||
#define ONE_MINUTE_MS 60 * 1000
|
||
#define THIRTY_SECONDS_MS 30 * 1000
|
||
#define TWO_SECONDS_MS 2 * 1000
|
||
#define FIVE_SECONDS_MS 5 * 1000
|
||
#define TEN_SECONDS_MS 10 * 1000
|
||
#define MAX_INTERVAL INT32_MAX // FIXME: INT32_MAX to avoid overflow issues with Apple clients but should be UINT32_MAX
|
||
|
||
#define min_default_telemetry_interval_secs IF_ROUTER(ONE_DAY / 2, 30 * 60)
|
||
#define default_gps_update_interval IF_ROUTER(ONE_DAY, 2 * 60)
|
||
#define default_telemetry_broadcast_interval_secs IF_ROUTER(ONE_DAY / 2, 60 * 60)
|
||
#define default_broadcast_interval_secs IF_ROUTER(ONE_DAY / 2, 60 * 60)
|
||
#define default_broadcast_smart_minimum_interval_secs 5 * 60
|
||
#define min_default_broadcast_interval_secs IF_ROUTER(ONE_DAY / 2, 60 * 60)
|
||
#define min_default_broadcast_smart_minimum_interval_secs 5 * 60
|
||
#define default_wait_bluetooth_secs IF_ROUTER(1, 60)
|
||
#define default_sds_secs IF_ROUTER(ONE_DAY, UINT32_MAX) // Default to forever super deep sleep
|
||
#define default_ls_secs IF_ROUTER(ONE_DAY, 5 * 60)
|
||
#define default_min_wake_secs 10
|
||
#define default_screen_on_secs IF_ROUTER(1, 60 * 10)
|
||
#define default_node_info_broadcast_secs 3 * 60 * 60
|
||
#define default_neighbor_info_broadcast_secs 6 * 60 * 60
|
||
#define min_node_info_broadcast_secs 60 * 60 // No regular broadcasts of more than once an hour
|
||
#define min_neighbor_info_broadcast_secs 4 * 60 * 60
|
||
#define default_map_publish_interval_secs 60 * 60
|
||
|
||
enum class TrafficType { POSITION, TELEMETRY };
|
||
|
||
// Traffic management defaults
|
||
#define default_traffic_mgmt_position_precision_bits 19 // ~90m grid cells (±45m)
|
||
#define default_traffic_mgmt_position_min_interval_secs (ONE_DAY / 2) // 12 hours between identical positions
|
||
|
||
// Hop scaling defaults
|
||
#define default_hop_scaling_min_target_nodes 40 // walk threshold: first hop reaching this cumulative count
|
||
#define default_hop_scaling_max_target_nodes 80 // generous extension ceiling (2 × min)
|
||
#define default_hop_scaling_min_target_nodes_floor 5 // minimum allowed min_target_nodes
|
||
#define default_hop_scaling_max_target_nodes_ceiling 512 // maximum allowed max_target_nodes
|
||
|
||
#ifdef USERPREFS_RINGTONE_NAG_SECS
|
||
#define default_ringtone_nag_secs USERPREFS_RINGTONE_NAG_SECS
|
||
#else
|
||
#define default_ringtone_nag_secs 15
|
||
#endif
|
||
#define default_network_ipv6_enabled false
|
||
|
||
#define default_mqtt_address "mqtt.meshtastic.org"
|
||
#define default_mqtt_username "meshdev"
|
||
#define default_mqtt_password "large4cats"
|
||
#define default_mqtt_root "msh"
|
||
#define default_mqtt_encryption_enabled true
|
||
#define default_mqtt_tls_enabled false
|
||
|
||
#define IF_ROUTER(routerVal, normalVal) \
|
||
((config.device.role == meshtastic_Config_DeviceConfig_Role_ROUTER || \
|
||
config.device.role == meshtastic_Config_DeviceConfig_Role_ROUTER_LATE) \
|
||
? (routerVal) \
|
||
: (normalVal))
|
||
|
||
class Default
|
||
{
|
||
public:
|
||
static uint32_t getConfiguredOrDefaultMs(uint32_t configuredInterval);
|
||
static uint32_t getConfiguredOrDefaultMs(uint32_t configuredInterval, uint32_t defaultInterval);
|
||
static uint32_t getConfiguredOrDefault(uint32_t configured, uint32_t defaultValue);
|
||
// Note: numOnlineNodes uses uint32_t to match the public API and allow flexibility,
|
||
// even though internal node counts use uint16_t (max 65535 nodes)
|
||
static uint32_t getConfiguredOrDefaultMsScaled(uint32_t configured, uint32_t defaultValue, uint32_t numOnlineNodes);
|
||
static uint32_t getConfiguredOrDefaultMsScaled(uint32_t configured, uint32_t defaultValue, uint32_t numOnlineNodes,
|
||
TrafficType type);
|
||
static uint8_t getConfiguredOrDefaultHopLimit(uint8_t configured);
|
||
static uint32_t getConfiguredOrMinimumValue(uint32_t configured, uint32_t minValue);
|
||
|
||
private:
|
||
// Note: Kept as uint32_t to match the public API parameter type
|
||
static float congestionScalingCoefficient(uint32_t numOnlineNodes)
|
||
{
|
||
if (numOnlineNodes <= 40) {
|
||
return 1.0;
|
||
} else {
|
||
// Resolve SF and BW from preset or manual config
|
||
// When use_preset is true, config.lora.spread_factor and bandwidth may be 0
|
||
// because applyModemConfig() sets them on RadioInterface, not on config.lora
|
||
float bwKHz;
|
||
uint8_t sf;
|
||
uint8_t cr;
|
||
if (config.lora.use_preset) {
|
||
modemPresetToParams(config.lora.modem_preset, false, bwKHz, sf, cr);
|
||
} else {
|
||
sf = config.lora.spread_factor;
|
||
bwKHz = bwCodeToKHz(config.lora.bandwidth);
|
||
}
|
||
|
||
// Guard against invalid values
|
||
sf = clampSpreadFactor(sf);
|
||
bwKHz = clampBandwidthKHz(bwKHz);
|
||
|
||
// throttlingFactor = 2^SF / (BW_in_kHz * scaling_divisor)
|
||
// With scaling_divisor=100:
|
||
// In SF11 and BW=250khz (longfast), this gives 0.08192 rather than the original 0.075
|
||
// In SF10 and BW=250khz (mediumslow), this gives 0.04096 rather than the original 0.04
|
||
// In SF9 and BW=250khz (mediumfast), this gives 0.02048 rather than the original 0.02
|
||
// In SF7 and BW=250khz (shortfast), this gives 0.00512 rather than the original 0.01
|
||
float throttlingFactor = static_cast<float>(pow_of_2(sf)) / (bwKHz * 100.0f);
|
||
|
||
#if USERPREFS_EVENT_MODE
|
||
// If we are in event mode, scale down the throttling factor by 4
|
||
throttlingFactor = static_cast<float>(pow_of_2(sf)) / (bwKHz * 25.0f);
|
||
#endif
|
||
|
||
// Scaling up traffic based on number of nodes over 40
|
||
int nodesOverForty = (numOnlineNodes - 40);
|
||
return 1.0 + (nodesOverForty * throttlingFactor); // Each number of online node scales by throttle factor
|
||
}
|
||
}
|
||
}; |