Files
claudetools/clients/cascades-tucson/reports/2026-06-16-unifi-full-audit.md
Howard Enos 8f72178d8a sync: auto-sync from HOWARD-HOME at 2026-06-17 09:35:47
Author: Howard Enos
Machine: HOWARD-HOME
Timestamp: 2026-06-17 09:35:47
2026-06-17 09:35:58 -07:00

111 lines
7.4 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Cascades of Tucson — UniFi Full Audit (2026-06-16)
Generated by the `unifi-wifi` skill (read-only). Fleet: **77 U7-Pro APs, 12 switches, ~587 clients**,
no UniFi gateway (pfSense firewall). All collectors ran clean.
## Issues — prioritized
| # | Sev | Issue | Detail / fix |
|---|---|---|---|
| 1 | **HIGH** | 2.4 GHz saturation (the "bad for some users") | 75 radios at auto/full power, ch6/1/11 at 16k33k neighbor BSSIDs; 123 clients on 2.4 at ~10% retry. Fix: power-down 65 radios to Low, then disable 9 redundant (plan below). |
| 2 | **HIGH** | ~25 switch ports linked at 100 M but gig-capable | FastEthernet/cabling-or-NIC issue capping those APs/devices at 100 Mbps (1st/2nd/3rd-floor switches). Physical: re-terminate/replace cable or check NIC. |
| 3 | **MED** | 6 GHz essentially unused — **1 client** of 587 | 75 6E radios live, nearly empty. Enable band-steering (`bandsteer`/`bands`) to offload 5 GHz onto the clean band. |
| 4 | **MED** | 5 GHz on 80 MHz (76/77) + 55 on DFS | 80 MHz kills spatial reuse in density → 40 MHz. DFS empirically low-risk here (0 radar events) but move to non-DFS for resilience (near Davis-Monthan). |
| 5 | **MED** | 6 APs with 2.4 min-RSSI OFF; 4 APs off the 1/6/11 plan | min-RSSI OFF: 615, 608, 505, 517, 622, salon. Off-plan (auto): 128, 108, 108U7-Pro, salon. |
| 6 | **LOW** | 3 offline switches + 2 disconnected APs + 1 firmware update | Offline: Switch 2nd Floor #2, Switch 4th Floor #2, USW Pro Max 16. APs: 108 ×2 (cable run pending, known). 1 device upgradable. |
| 7 | **LOW** | p38 (1st Floor USW) 4.0% tx-drop rate | Correlates with the underspeed/heavy-traffic ports; investigate after #2. |
## WiFi detail
- **2.4 GHz:** 77 radios, all 20 MHz (good); power auto×75 (want Low); channels 1:20 / 6:28 / 11:25 / auto:4.
min-RSSI OFF on 6. Neighbor density: ch6 33,376 · ch1 19,355 · ch11 16,598 BSSIDs. Live retry avg **10.2%**.
- **5 GHz:** 77 radios, 80 MHz ×76 / 40 MHz ×1; rogue density biased to ch149/157 (busy upper). 463 clients. retry avg **8.0%**.
- **6 GHz:** 75 radios active, **1 client**. Wide open.
- **AP satisfaction (live):** min 90 / median 98 / max 100 → healthy in aggregate; the pain is the 2.4 GHz client tail.
## Data-backed radio plan (optimize-radios + /proc/ui_neighbor SNR matrix)
- **Phase A — power-down 65** 2.4 radios to Low (smaller cells cut mutual interference; coverage-safe).
- **Phase C — disable 9** redundant 2.4 radios after re-measure (each heard by ≥2 strong neighbors):
127→128, 229→128, 248→348, 330→128, 445→347/348/247, 428→128, 622→505/615/608, Kitchen→Memcare TV room,
Dining Room→memcare piano. Est. interference-airtime removed: ~619.
- **Channel plan available:** 2.4 GHz 1/6/11 graph-color (co-channel pairs **92→35**); 5 GHz non-DFS
(**20→0** and all off DFS).
## Switch / PoE (12 switches, 29 flags)
- ~25 ports at 100 M but gig-capable (see #2). PoE budgets healthy (e.g. 1st-floor 160/600 W).
- 3 offline switches (above).
## Gateway / WAN
- No UniFi gateway (pfSense) → WAN/internet not measurable via UniFi. Adoption: APs 77 (2 disc), switches
12 (3 disc), 587 clients.
## Recommended sequence
1. Physical: fix the ~25 underspeed ports (#2) + chase the 3 offline switches / AP 108 cable.
2. WiFi Phase A: power-down 2.4 to Low per zone, validate with watch-ap (live before/after).
3. Enable 6 GHz band-steering + 5 GHz 80→40 MHz non-DFS channel plan.
4. Set 2.4 min-RSSI on the 6 OFF APs; pin the 4 off-plan APs to 1/6/11.
5. Phase C: disable the 9 redundant 2.4 radios after re-measure.
(All changes via the gated `apply-radio`/`apply-wlan`/`channel-plan` scripts — per zone, with rollback +
live validation. Nothing applied in this audit.)
---
## pfSense health check (2026-06-16) — ruling out the gateway as a WiFi factor
Investigated the Cascades pfSense (`192.168.0.1`, **pfSense Plus 25.07-RELEASE**, Netgate) over the site
VPN via SSH, to confirm whether any gateway-side issue contributes to the "WiFi bad for some users"
symptom. **Verdict: pfSense is healthy and is NOT a contributor — the problem is RF-side (2.4 GHz).**
| Area | Finding | WiFi impact |
|---|---|---|
| **DHCP exhaustion** | **0** "no free leases" events in dhcpd.log. WiFi/AP pool `192.168.0.0/22` (range 192.168.2.23.254, cap ~507) only **270 active (~53%)**; per-unit /28s + `10.0.20/.50` all have headroom | **Ruled out** (was the top suspect) |
| **DNS** | unbound resolver running | Fine |
| **WAN** | Dual Cox — WAN1 `184.191.143.62/30`, WAN2 `72.211.21.217/27`, both active **full-duplex**, `WAN_Group` gateway group, **no loss/down events** logged | Fine |
| **Firewall states** | 28,368 / 790,000 limit | Fine |
| **CPU / mbuf / uptime** | load 0.6, mbufs nominal, 10-day uptime | Healthy |
**Architecture:** per-unit design — **199 DHCP subnets**, mostly `10.x.y.0/28` per apartment (assisted-
living L2 isolation) + the `192.168.0.0/22` staff/AP network (APs + most WiFi clients). Active DHCP
backend is **ISC** (Kea config present but dormant).
**Minor (not WiFi-related):** `igc3`/WAN2 logged 1707 input-errors + 1707 "collisions", but the link is
2.5GbE full-duplex/active with zero gateway loss — consistent with the known Intel I225/I226 2.5G counter
quirk, not a real fault. No action needed unless WAN2 misbehaves.
**Conclusion:** gateway/DHCP/DNS/WAN are not bottlenecking the wireless. The 2.4 GHz remediation
(power-down + coverage-redundancy disables) remains the correct and sole fix for the client-experience tail.
## Daytime re-check (2026-06-17, ~09:33, loaded network)
Post-remediation loaded-network measurement (overnight 6/17: 24 of 76 2.4 radios disabled, 42 set to Low
~6 dBm; Floors 5/6 + mesh untouched). Fleet healthy: 77 adopted / 2 disconnected (same known 108 + 1) --
**no AP went offline.**
**Active-radio-only snapshot (09:33, excludes the 24 disabled):** cu_total 67%, cu_interf 48%, clients 105
(vs pre-change active cu_total 77% / cu_interf 64%).
**Time-of-day-controlled hourly (this AM vs same hours yesterday @ full power):**
| 2.4 site-wide AM | Yesterday (full) | Today (post) |
|---|---|---|
| cu_total | 79% | 44% |
| cu_interf | 66% | 32% |
| retry% | 17.0% | 23.4% |
| satisfaction | 39 | 30 |
| sta/radio | 2.32 | 2.21 |
**Mixed result -- read honestly:**
- WIN: 2.4 interference/airtime is clearly down (fewer contending transmitters + smaller cells). On the
active radios, cu_interf ~64% -> ~48% (snapshot) / 66% -> 32% (morning hourly avg, fewer radios).
- CONCERN: **retry rose (17.0 -> 23.4%) and satisfaction fell (39 -> 30)** time-of-day-controlled. Likely
OVER-THINNING: Low = ~6 dBm is a 17 dB cut from auto (~23), and with 24 radios also disabled some edge
2.4 clients now reach a farther/weaker AP -> more retransmits, lower satisfaction. (Partly composition:
fewer 2.4 clients today, so the remaining tail may be the stubborn legacy/far devices.)
**RECOMMENDATION (needs Howard's go -- this re-check is read-only):**
1. Bump the KEPT 2.4 radios from Low -> **Medium** (~12-15 dBm): restores client signal while keeping cells
smaller than full power. `apply-radio cascades ng power medium --zone ...` per floor, then re-measure.
2. Do NOT expand disables further; consider re-enabling a specific radio if a dead zone/complaint appears.
3. Re-measure retry/satisfaction after the Low->Medium bump (same time-of-day) to confirm recovery.
The interference goal is met; the lever now is finding the power floor that keeps cells tight WITHOUT
starving edge clients. Medium is the likely sweet spot vs the aggressive Low.