Files
claudetools/clients/cascades-tucson/reports/2026-06-16-unifi-full-audit.md
Howard Enos 8f72178d8a sync: auto-sync from HOWARD-HOME at 2026-06-17 09:35:47
Author: Howard Enos
Machine: HOWARD-HOME
Timestamp: 2026-06-17 09:35:47
2026-06-17 09:35:58 -07:00

7.4 KiB
Raw Blame History

Cascades of Tucson — UniFi Full Audit (2026-06-16)

Generated by the unifi-wifi skill (read-only). Fleet: 77 U7-Pro APs, 12 switches, ~587 clients, no UniFi gateway (pfSense firewall). All collectors ran clean.

Issues — prioritized

# Sev Issue Detail / fix
1 HIGH 2.4 GHz saturation (the "bad for some users") 75 radios at auto/full power, ch6/1/11 at 16k33k neighbor BSSIDs; 123 clients on 2.4 at ~10% retry. Fix: power-down 65 radios to Low, then disable 9 redundant (plan below).
2 HIGH ~25 switch ports linked at 100 M but gig-capable FastEthernet/cabling-or-NIC issue capping those APs/devices at 100 Mbps (1st/2nd/3rd-floor switches). Physical: re-terminate/replace cable or check NIC.
3 MED 6 GHz essentially unused — 1 client of 587 75 6E radios live, nearly empty. Enable band-steering (bandsteer/bands) to offload 5 GHz onto the clean band.
4 MED 5 GHz on 80 MHz (76/77) + 55 on DFS 80 MHz kills spatial reuse in density → 40 MHz. DFS empirically low-risk here (0 radar events) but move to non-DFS for resilience (near Davis-Monthan).
5 MED 6 APs with 2.4 min-RSSI OFF; 4 APs off the 1/6/11 plan min-RSSI OFF: 615, 608, 505, 517, 622, salon. Off-plan (auto): 128, 108, 108U7-Pro, salon.
6 LOW 3 offline switches + 2 disconnected APs + 1 firmware update Offline: Switch 2nd Floor #2, Switch 4th Floor #2, USW Pro Max 16. APs: 108 ×2 (cable run pending, known). 1 device upgradable.
7 LOW p38 (1st Floor USW) 4.0% tx-drop rate Correlates with the underspeed/heavy-traffic ports; investigate after #2.

WiFi detail

  • 2.4 GHz: 77 radios, all 20 MHz (good); power auto×75 (want Low); channels 1:20 / 6:28 / 11:25 / auto:4. min-RSSI OFF on 6. Neighbor density: ch6 33,376 · ch1 19,355 · ch11 16,598 BSSIDs. Live retry avg 10.2%.
  • 5 GHz: 77 radios, 80 MHz ×76 / 40 MHz ×1; rogue density biased to ch149/157 (busy upper). 463 clients. retry avg 8.0%.
  • 6 GHz: 75 radios active, 1 client. Wide open.
  • AP satisfaction (live): min 90 / median 98 / max 100 → healthy in aggregate; the pain is the 2.4 GHz client tail.

Data-backed radio plan (optimize-radios + /proc/ui_neighbor SNR matrix)

  • Phase A — power-down 65 2.4 radios to Low (smaller cells cut mutual interference; coverage-safe).
  • Phase C — disable 9 redundant 2.4 radios after re-measure (each heard by ≥2 strong neighbors): 127→128, 229→128, 248→348, 330→128, 445→347/348/247, 428→128, 622→505/615/608, Kitchen→Memcare TV room, Dining Room→memcare piano. Est. interference-airtime removed: ~619.
  • Channel plan available: 2.4 GHz 1/6/11 graph-color (co-channel pairs 92→35); 5 GHz non-DFS (20→0 and all off DFS).

Switch / PoE (12 switches, 29 flags)

  • ~25 ports at 100 M but gig-capable (see #2). PoE budgets healthy (e.g. 1st-floor 160/600 W).
  • 3 offline switches (above).

Gateway / WAN

  • No UniFi gateway (pfSense) → WAN/internet not measurable via UniFi. Adoption: APs 77 (2 disc), switches 12 (3 disc), 587 clients.
  1. Physical: fix the ~25 underspeed ports (#2) + chase the 3 offline switches / AP 108 cable.
  2. WiFi Phase A: power-down 2.4 to Low per zone, validate with watch-ap (live before/after).
  3. Enable 6 GHz band-steering + 5 GHz 80→40 MHz non-DFS channel plan.
  4. Set 2.4 min-RSSI on the 6 OFF APs; pin the 4 off-plan APs to 1/6/11.
  5. Phase C: disable the 9 redundant 2.4 radios after re-measure.

(All changes via the gated apply-radio/apply-wlan/channel-plan scripts — per zone, with rollback + live validation. Nothing applied in this audit.)


pfSense health check (2026-06-16) — ruling out the gateway as a WiFi factor

Investigated the Cascades pfSense (192.168.0.1, pfSense Plus 25.07-RELEASE, Netgate) over the site VPN via SSH, to confirm whether any gateway-side issue contributes to the "WiFi bad for some users" symptom. Verdict: pfSense is healthy and is NOT a contributor — the problem is RF-side (2.4 GHz).

Area Finding WiFi impact
DHCP exhaustion 0 "no free leases" events in dhcpd.log. WiFi/AP pool 192.168.0.0/22 (range 192.168.2.23.254, cap ~507) only 270 active (~53%); per-unit /28s + 10.0.20/.50 all have headroom Ruled out (was the top suspect)
DNS unbound resolver running Fine
WAN Dual Cox — WAN1 184.191.143.62/30, WAN2 72.211.21.217/27, both active full-duplex, WAN_Group gateway group, no loss/down events logged Fine
Firewall states 28,368 / 790,000 limit Fine
CPU / mbuf / uptime load 0.6, mbufs nominal, 10-day uptime Healthy

Architecture: per-unit design — 199 DHCP subnets, mostly 10.x.y.0/28 per apartment (assisted- living L2 isolation) + the 192.168.0.0/22 staff/AP network (APs + most WiFi clients). Active DHCP backend is ISC (Kea config present but dormant).

Minor (not WiFi-related): igc3/WAN2 logged 1707 input-errors + 1707 "collisions", but the link is 2.5GbE full-duplex/active with zero gateway loss — consistent with the known Intel I225/I226 2.5G counter quirk, not a real fault. No action needed unless WAN2 misbehaves.

Conclusion: gateway/DHCP/DNS/WAN are not bottlenecking the wireless. The 2.4 GHz remediation (power-down + coverage-redundancy disables) remains the correct and sole fix for the client-experience tail.

Daytime re-check (2026-06-17, ~09:33, loaded network)

Post-remediation loaded-network measurement (overnight 6/17: 24 of 76 2.4 radios disabled, 42 set to Low ~6 dBm; Floors 5/6 + mesh untouched). Fleet healthy: 77 adopted / 2 disconnected (same known 108 + 1) -- no AP went offline.

Active-radio-only snapshot (09:33, excludes the 24 disabled): cu_total 67%, cu_interf 48%, clients 105 (vs pre-change active cu_total 77% / cu_interf 64%).

Time-of-day-controlled hourly (this AM vs same hours yesterday @ full power):

2.4 site-wide AM Yesterday (full) Today (post)
cu_total 79% 44%
cu_interf 66% 32%
retry% 17.0% 23.4%
satisfaction 39 30
sta/radio 2.32 2.21

Mixed result -- read honestly:

  • WIN: 2.4 interference/airtime is clearly down (fewer contending transmitters + smaller cells). On the active radios, cu_interf ~64% -> ~48% (snapshot) / 66% -> 32% (morning hourly avg, fewer radios).
  • CONCERN: retry rose (17.0 -> 23.4%) and satisfaction fell (39 -> 30) time-of-day-controlled. Likely OVER-THINNING: Low = ~6 dBm is a 17 dB cut from auto (~23), and with 24 radios also disabled some edge 2.4 clients now reach a farther/weaker AP -> more retransmits, lower satisfaction. (Partly composition: fewer 2.4 clients today, so the remaining tail may be the stubborn legacy/far devices.)

RECOMMENDATION (needs Howard's go -- this re-check is read-only):

  1. Bump the KEPT 2.4 radios from Low -> Medium (~12-15 dBm): restores client signal while keeping cells smaller than full power. apply-radio cascades ng power medium --zone ... per floor, then re-measure.
  2. Do NOT expand disables further; consider re-enabling a specific radio if a dead zone/complaint appears.
  3. Re-measure retry/satisfaction after the Low->Medium bump (same time-of-day) to confirm recovery. The interference goal is met; the lever now is finding the power floor that keeps cells tight WITHOUT starving edge clients. Medium is the likely sweet spot vs the aggressive Low.