Meshtastic Repeater Network Patterns

Designing for Reliability: N+1 Redundancy

Designing for Reliability: N+1 Redundancy

In plain terms: make sure that no single node failure can split your network into two halves that can no longer reach each other. A mesh network is only as reliable as its weakest single point of failure. In graph theory terms, a node whose removal splits a connected graph into two or more disconnected components is called a cut vertex. Real-world Meshtastic deployments often develop cut vertices without operators realizing it - especially as networks grow organically. N+1 redundancy means ensuring that for every critical backbone node, at least one backup path exists so that the loss of any single node does not partition the network.

Important caveat: a second node at the SAME site protects only against single-device failure (firmware crash, board death). It does NOT protect against the failures most likely in an emergency - site power loss, lightning, tower collapse, fire, flooding - because both nodes share that fate. For incident-grade redundancy, the backup MUST be at a separate site on independent power.

Identifying Critical Backbone Nodes

Start by drawing your network on paper or a whiteboard. Place each node as a dot and draw lines between nodes that can hear each other directly. Now ask: if I erase this dot, does the network split into two disconnected groups? Any node where the answer is yes is a cut vertex and a single point of failure.

For larger networks, the most reliable analysis is manual: run meshtastic --traceroute between nodes on opposite sides of a suspect backbone node - if the only path routes through that one node, it is critical. Third-party visualizers such as meshmap.net can help you eyeball topology, but they only show nodes that report to the public MQTT server (a gateway is required) and are not authoritative; do not rely on them as a complete picture of your mesh.

Providing Backup Paths

Once critical nodes are identified, the fix is straightforward in concept: ensure each one has a backup. Two approaches work well in practice:

Testing Redundancy

Testing is non-negotiable. Make sure no single node failure splits your network: a backup path that exists on paper but has never been verified may fail in practice due to marginal signal, wrong node roles, or misconfigured hop limits. The test procedure is simple:

  1. Identify two nodes on opposite sides of the critical backbone node you want to test.
  2. Run a traceroute between them and record the path.
  3. Briefly power the backbone node down to simulate failure. (Power it off - do NOT simply unplug its antenna while it is transmitting; transmitting into a disconnected antenna can damage the radio.)
  4. Run the traceroute again. The path should route around the powered-down node.
  5. If routing fails, the backup path is insufficient and must be improved before the critical node is trusted.

A passing reroute test proves the backup path exists, not that it has capacity. Under incident load the backup path carries the failed node's traffic on top of its own and may saturate. Verify the backup path also has channel-utilization headroom for surge traffic, not just connectivity.

Test redundancy at least quarterly, and ALWAYS as part of pre-incident readiness checks or before any planned event you expect to rely on the network. Annual testing is insufficient for networks used in emergencies - paths degrade silently between tests, and any time the physical environment changes significantly you should retest.

Network Mapping Tools

Third-party visualizers such as meshmap.net provide a visual overlay of node positions for nodes that report to the public MQTT server, which can help you spot obvious topological bottlenecks - but they are not authoritative and will not show nodes that lack an internet gateway. The ground truth for redundancy verification is the meshtastic --traceroute <destination_id> CLI command, which reveals the actual hop-by-hop path packets take in real time.

Planned Maintenance Procedures for Live Networks

Planned Maintenance Procedures for Live Networks

Taking a backbone node offline for maintenance - whether for firmware updates, hardware replacement, or antenna adjustments - affects the users routing through it. With proper planning, that impact can be reduced to a brief interruption rather than a prolonged outage. This page describes a repeatable maintenance procedure for Meshtastic backbone nodes in active community networks.

Pre-Maintenance Checklist

Complete these steps before any planned maintenance window:

During Maintenance

Where possible, power down gracefully rather than performing a hard power-off. A graceful shutdown lets the node stop transmitting cleanly and avoids cutting off a packet mid-transmission or corrupting the node's stored configuration during a flash or write. For solar-powered nodes, the practical approach is to disconnect the load output of the charge controller rather than physically unplugging the node in the dark.

Perform the maintenance task - firmware flash, hardware swap, antenna replacement - as quickly as practical. Every additional minute of downtime increases the chance of a user encountering a failed message delivery.

Post-Maintenance Verification

Before declaring the node returned to service, verify it from multiple directions:

  1. Send a test message from a node that previously routed through the maintained node and confirm delivery.
  2. Run traceroutes from multiple directions to confirm the node is routing normally.
  3. If you run a monitoring system (for example MQTT/Grafana or a third-party map such as meshmap.net), confirm it shows the node online AND recently active. Note that seeing a node online is necessary but NOT sufficient - a node can appear up while failing to relay (for example, wrong role or rebroadcast mode). Always also complete step 1 (a real test message routed through the node) before declaring it back in service.
  4. Update your network maintenance log with the date, work performed, and any configuration changes.

Emergency Rollback Procedure

If a replacement node does not work as expected - wrong firmware, hardware fault, or configuration error - act quickly. Restore the original node if it is still functional, or connect a known-good spare configured to match the original settings. If neither is possible, notify the community immediately that the outage is extended and provide an estimated restoration time. For any network used in emergencies, keep at least one spare node per critical site, PRE-CONFIGURED to match that site (role, channel/PSK, hop limit, fixed position). An unconfigured spare is not a substitute - configuring on-site under incident pressure wastes the very time the spare was meant to save.

After any unplanned extended outage, conduct a brief post-incident review: what failed, why, and what process change would prevent recurrence. Even a one-paragraph note in the maintenance log is valuable for future operators.

Channel Utilization Management

Channel Utilization Management

Channel utilization (CU) is one of the most important health metrics for a Meshtastic network, yet it is frequently misunderstood or ignored until problems become severe. Understanding what CU measures, what causes it to rise, and how to bring it back down is essential knowledge for any network operator running more than a handful of nodes.

What Is Channel Utilization?

Channel utilization is the percentage of time the radio channel is occupied by transmissions, measured over a rolling 1-minute window. Meshtastic breaks the minute into six 10-second sub-windows and sums the milliseconds of airtime used in each. A CU of 10% therefore means the channel was busy for roughly 6 seconds out of the last minute (not 90 seconds out of 15 minutes). Meshtastic calculates CU locally on each node by monitoring the time its own radio is keyed up, plus time spent sensing channel activity from other nodes (via Channel Activity Detection). Note that CU counts the airtime of all LoRa traffic heard on the same modem settings, including non-Meshtastic LoRa. The reported CU figure is visible in the Meshtastic app under the node detail view and in the device telemetry metrics channel. (Do not confuse this 1-minute window with the EU 10% duty-cycle limit, which is measured over a 1-hour window.)

Healthy, Warning, and Critical Thresholds

These bands match the color coding used by the Meshtastic apps:

If you prefer a stricter operational target for infrastructure (for example, keeping a busy backbone channel below ~20-25% to preserve headroom for incident traffic), treat that as a conservative operator guideline, not the system threshold. CU is also a lagging rolling average: a channel can briefly saturate while the reported figure still looks healthy.

Common Causes of High Channel Utilization

Several factors drive CU up. Understanding which ones apply to your network guides the correct fix:

Remediation Strategies

These remediations are proactive. They must be applied before an incident. Once a channel saturates during an event, Meshtastic has no automatic congestion recovery; the only mid-incident fix is to reduce traffic (fewer senders, shorter messages, suppress telemetry), which requires discipline from every operator rather than a single config change - and reconfiguring deployed nodes mid-incident is often impossible. Provision capacity headroom in advance.