Skip to content

fix(firmware): stop nodes falling back to 1 Mbps 802.11b, which can fill the channel - #2081

Open
clonea1 wants to merge 1 commit into
ruvnet:mainfrom
clonea1:contrib/wifi-no-11b
Open

clonea1 wants to merge 1 commit into
ruvnet:mainfrom
clonea1:contrib/wifi-no-11b

Conversation

@clonea1

@clonea1 clonea1 commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

What this fixes

A node whose WiFi link degrades can fall back to 1-2 Mbps 802.11b and stay there for hours. Each CSI datagram then costs tens of times the airtime, so one or two slow nodes fill the shared 2.4 GHz channel and every node's frame rate slides. The TX queue also backs up behind the slow link, which looks like a memory leak (deep free-heap dips, send_fail bursts) when it is not.

This PR stops the node's own rate control from using 802.11b rates, so it cannot go below 6 Mbps OFDM, whatever rates the AP advertises. It may explain part of #978 (WiFi disruption when the full fleet powers on) and #1183 (ENOMEM backoff under a weak uplink).

Changes (4 files, +60)

  • main.c: esp_wifi_config_11b_rate(WIFI_IF_STA, true) after esp_wifi_init() and before esp_wifi_start(). Non-fatal: if the driver refuses, the node logs it and runs as before.
  • c6_sync_espnow.c: the ESP-NOW broadcast peer is set to 11g 6 Mbps OFDM right after esp_now_add_peer(). ESP-NOW defaults to 1 Mbps 802.11b, which is not sendable once 11b is off; OFDM also carries the LTF that CSI is computed from.
  • nvs_config.{c,h}: NVS u8 key allow_11b, default 0. Setting it to 1 restores the previous behaviour on the next boot, a rollback that needs no reflash.

This changes default behaviour for every node. That is deliberate (the old default is the failure mode), and allow_11b=1 is the way back for a deployment that needs 802.11b.

Evidence (9 x ESP32-C6, IDF v5.4, one AP, channel 11; measured from the AP's per-client stats and AP-side captures)

8 days before after
channel busy, daily median (p90) 80-85% (85-87%) 44% (56%)
node-minutes per day at <= 2 Mbps uplink 988-1968 89
fleet accepted CSI frames/s ~200 ~253 (+26%)
node-to-node CSI frames/s 38.7 50.8 (+31%)

Overnight since: the first night still had 134 node-minutes at <= 2 Mbps (dips that recovered on their own, see below); the second had none and the channel never went above 56% busy, although no rejoin trigger fired that night. No rollbacks or crash reboots on either.

Caveats, stated plainly:

  • This patch and the ESP-NOW rate change shipped together, so the gain cannot be split between them.
  • The block is partial on the C6. AP-side captures still show RTS at 2 Mbps and some retries at 5.5 Mbps from the nodes. The C6 driver also refuses any 2.4 GHz protocol set without 11B (esp_wifi_set_protocol returns ESP_ERR_INVALID_ARG on IDF 5.4.0, 5.4.4 and 5.5.5), so this API is the strongest lever firmware has.
  • 6 Mbps OFDM needs more SNR than 1 Mbps DSSS: the weakest far node-to-node ESP-NOW links got worse (in our fleet, one long pair dropped from ~2 to ~0.1 frames/s), while most links improved.
  • Rate dips still happen (here, legacy IoT clients rejoining the AP set them off), but they now recover by themselves in under an hour instead of sticking.

Tested

  • Upstream CI's three variants, built in espressif/idf:v5.4 exactly as firmware-ci.yml does, against unmodified main for comparison:

    variant size vs main limit / warn (KB) new warnings
    esp32s3 8mb 1,193,920 B +2,480 B 1200 / 1100 0
    esp32s3 4mb 945,984 B +2,416 B 1152 / 1100 0
    esp32c6 4mb 1,045,424 B +3,296 B 1152 / 1100 0

    The 8mb variant is already over the 1100 KB soft warning on main (1163.5 KB); this adds 2.4 KB and stays under the hard limit.

  • Host tests (test/Makefile host_tests): identical results on main and on this branch (adr110 21/21, vitals 30/30, mmwave 8/8, thermal, csi_sanitize, c6_antenna, serial_onboarding, delivery_contract 3/3).

  • Every API used (esp_wifi_config_11b_rate, esp_now_set_peer_rate_config, esp_now_rate_config_t, WIFI_PHY_RATE_6M) checked in the IDF v5.4 headers; none is target-gated.

  • On hardware: the same change has run on all 9 C6 nodes since 2026-09-30. Not run on S3 hardware.

If your fleet's frame rate decays with uptime

Worth checking before blaming firmware:

  • Look at the AP's per-client receive rate for each node. A node sitting at 1-2 Mbps costs roughly 10% of the channel by itself.
  • A minimum data rate on the AP may not be enough. UniFi, for example, can enforce a 12 Mbps minimum while still advertising the 802.11b rates (minrate_ng_advertising_rates: false), and clients keep using them.
  • Another BSS on the same radio that still allows 802.11b can keep protection mode on for everyone.
  • Free-heap dips that deepen with uptime on a node with a slow link are usually the TX queue backing up, not a leak; the minimum-heap watermark only ever goes down, so it reads like one.

🤖 Generated with Claude Code

…ill the channel

When a node's link degrades, the WiFi driver's rate control can settle on
1-2 Mbps 802.11b and stay there for hours. Every CSI datagram then costs
tens of times the airtime, the shared channel fills up, and every node's
frame rate slides. The TX queue also backs up behind the slow link, which
shows as deep free-heap dips and send_fail bursts that look like a leak.

- main.c: esp_wifi_config_11b_rate(WIFI_IF_STA, true) after esp_wifi_init()
  and before esp_wifi_start(), so rate control cannot drop below 6 Mbps
  OFDM, whatever rates the AP advertises.
- c6_sync_espnow.c: set the ESP-NOW broadcast peer to 11g 6 Mbps OFDM right
  after esp_now_add_peer(). ESP-NOW defaults to 1 Mbps 802.11b, which is not
  sendable once 11b is disabled on the STA; OFDM also carries the LTF that
  CSI is computed from.
- nvs_config.{c,h}: NVS u8 key "allow_11b" (default 0). Setting it to 1
  restores the previous behaviour on the next boot, as a no-reflash
  rollback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@clonea1
clonea1 force-pushed the contrib/wifi-no-11b branch from 04784d6 to 5fd4be4 Compare October 10, 2026 21:15

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant