111 lines
7.4 KiB
Markdown
111 lines
7.4 KiB
Markdown
# Cascades of Tucson — UniFi Full Audit (2026-06-16)
|
||
|
||
Generated by the `unifi-wifi` skill (read-only). Fleet: **77 U7-Pro APs, 12 switches, ~587 clients**,
|
||
no UniFi gateway (pfSense firewall). All collectors ran clean.
|
||
|
||
## Issues — prioritized
|
||
|
||
| # | Sev | Issue | Detail / fix |
|
||
|---|---|---|---|
|
||
| 1 | **HIGH** | 2.4 GHz saturation (the "bad for some users") | 75 radios at auto/full power, ch6/1/11 at 16k–33k neighbor BSSIDs; 123 clients on 2.4 at ~10% retry. Fix: power-down 65 radios to Low, then disable 9 redundant (plan below). |
|
||
| 2 | **HIGH** | ~25 switch ports linked at 100 M but gig-capable | FastEthernet/cabling-or-NIC issue capping those APs/devices at 100 Mbps (1st/2nd/3rd-floor switches). Physical: re-terminate/replace cable or check NIC. |
|
||
| 3 | **MED** | 6 GHz essentially unused — **1 client** of 587 | 75 6E radios live, nearly empty. Enable band-steering (`bandsteer`/`bands`) to offload 5 GHz onto the clean band. |
|
||
| 4 | **MED** | 5 GHz on 80 MHz (76/77) + 55 on DFS | 80 MHz kills spatial reuse in density → 40 MHz. DFS empirically low-risk here (0 radar events) but move to non-DFS for resilience (near Davis-Monthan). |
|
||
| 5 | **MED** | 6 APs with 2.4 min-RSSI OFF; 4 APs off the 1/6/11 plan | min-RSSI OFF: 615, 608, 505, 517, 622, salon. Off-plan (auto): 128, 108, 108U7-Pro, salon. |
|
||
| 6 | **LOW** | 3 offline switches + 2 disconnected APs + 1 firmware update | Offline: Switch 2nd Floor #2, Switch 4th Floor #2, USW Pro Max 16. APs: 108 ×2 (cable run pending, known). 1 device upgradable. |
|
||
| 7 | **LOW** | p38 (1st Floor USW) 4.0% tx-drop rate | Correlates with the underspeed/heavy-traffic ports; investigate after #2. |
|
||
|
||
## WiFi detail
|
||
- **2.4 GHz:** 77 radios, all 20 MHz (good); power auto×75 (want Low); channels 1:20 / 6:28 / 11:25 / auto:4.
|
||
min-RSSI OFF on 6. Neighbor density: ch6 33,376 · ch1 19,355 · ch11 16,598 BSSIDs. Live retry avg **10.2%**.
|
||
- **5 GHz:** 77 radios, 80 MHz ×76 / 40 MHz ×1; rogue density biased to ch149/157 (busy upper). 463 clients. retry avg **8.0%**.
|
||
- **6 GHz:** 75 radios active, **1 client**. Wide open.
|
||
- **AP satisfaction (live):** min 90 / median 98 / max 100 → healthy in aggregate; the pain is the 2.4 GHz client tail.
|
||
|
||
## Data-backed radio plan (optimize-radios + /proc/ui_neighbor SNR matrix)
|
||
- **Phase A — power-down 65** 2.4 radios to Low (smaller cells cut mutual interference; coverage-safe).
|
||
- **Phase C — disable 9** redundant 2.4 radios after re-measure (each heard by ≥2 strong neighbors):
|
||
127→128, 229→128, 248→348, 330→128, 445→347/348/247, 428→128, 622→505/615/608, Kitchen→Memcare TV room,
|
||
Dining Room→memcare piano. Est. interference-airtime removed: ~619.
|
||
- **Channel plan available:** 2.4 GHz 1/6/11 graph-color (co-channel pairs **92→35**); 5 GHz non-DFS
|
||
(**20→0** and all off DFS).
|
||
|
||
## Switch / PoE (12 switches, 29 flags)
|
||
- ~25 ports at 100 M but gig-capable (see #2). PoE budgets healthy (e.g. 1st-floor 160/600 W).
|
||
- 3 offline switches (above).
|
||
|
||
## Gateway / WAN
|
||
- No UniFi gateway (pfSense) → WAN/internet not measurable via UniFi. Adoption: APs 77 (2 disc), switches
|
||
12 (3 disc), 587 clients.
|
||
|
||
## Recommended sequence
|
||
1. Physical: fix the ~25 underspeed ports (#2) + chase the 3 offline switches / AP 108 cable.
|
||
2. WiFi Phase A: power-down 2.4 to Low per zone, validate with watch-ap (live before/after).
|
||
3. Enable 6 GHz band-steering + 5 GHz 80→40 MHz non-DFS channel plan.
|
||
4. Set 2.4 min-RSSI on the 6 OFF APs; pin the 4 off-plan APs to 1/6/11.
|
||
5. Phase C: disable the 9 redundant 2.4 radios after re-measure.
|
||
|
||
(All changes via the gated `apply-radio`/`apply-wlan`/`channel-plan` scripts — per zone, with rollback +
|
||
live validation. Nothing applied in this audit.)
|
||
|
||
---
|
||
|
||
## pfSense health check (2026-06-16) — ruling out the gateway as a WiFi factor
|
||
|
||
Investigated the Cascades pfSense (`192.168.0.1`, **pfSense Plus 25.07-RELEASE**, Netgate) over the site
|
||
VPN via SSH, to confirm whether any gateway-side issue contributes to the "WiFi bad for some users"
|
||
symptom. **Verdict: pfSense is healthy and is NOT a contributor — the problem is RF-side (2.4 GHz).**
|
||
|
||
| Area | Finding | WiFi impact |
|
||
|---|---|---|
|
||
| **DHCP exhaustion** | **0** "no free leases" events in dhcpd.log. WiFi/AP pool `192.168.0.0/22` (range 192.168.2.2–3.254, cap ~507) only **270 active (~53%)**; per-unit /28s + `10.0.20/.50` all have headroom | **Ruled out** (was the top suspect) |
|
||
| **DNS** | unbound resolver running | Fine |
|
||
| **WAN** | Dual Cox — WAN1 `184.191.143.62/30`, WAN2 `72.211.21.217/27`, both active **full-duplex**, `WAN_Group` gateway group, **no loss/down events** logged | Fine |
|
||
| **Firewall states** | 28,368 / 790,000 limit | Fine |
|
||
| **CPU / mbuf / uptime** | load 0.6, mbufs nominal, 10-day uptime | Healthy |
|
||
|
||
**Architecture:** per-unit design — **199 DHCP subnets**, mostly `10.x.y.0/28` per apartment (assisted-
|
||
living L2 isolation) + the `192.168.0.0/22` staff/AP network (APs + most WiFi clients). Active DHCP
|
||
backend is **ISC** (Kea config present but dormant).
|
||
|
||
**Minor (not WiFi-related):** `igc3`/WAN2 logged 1707 input-errors + 1707 "collisions", but the link is
|
||
2.5GbE full-duplex/active with zero gateway loss — consistent with the known Intel I225/I226 2.5G counter
|
||
quirk, not a real fault. No action needed unless WAN2 misbehaves.
|
||
|
||
**Conclusion:** gateway/DHCP/DNS/WAN are not bottlenecking the wireless. The 2.4 GHz remediation
|
||
(power-down + coverage-redundancy disables) remains the correct and sole fix for the client-experience tail.
|
||
|
||
## Daytime re-check (2026-06-17, ~09:33, loaded network)
|
||
|
||
Post-remediation loaded-network measurement (overnight 6/17: 24 of 76 2.4 radios disabled, 42 set to Low
|
||
~6 dBm; Floors 5/6 + mesh untouched). Fleet healthy: 77 adopted / 2 disconnected (same known 108 + 1) --
|
||
**no AP went offline.**
|
||
|
||
**Active-radio-only snapshot (09:33, excludes the 24 disabled):** cu_total 67%, cu_interf 48%, clients 105
|
||
(vs pre-change active cu_total 77% / cu_interf 64%).
|
||
|
||
**Time-of-day-controlled hourly (this AM vs same hours yesterday @ full power):**
|
||
| 2.4 site-wide AM | Yesterday (full) | Today (post) |
|
||
|---|---|---|
|
||
| cu_total | 79% | 44% |
|
||
| cu_interf | 66% | 32% |
|
||
| retry% | 17.0% | 23.4% |
|
||
| satisfaction | 39 | 30 |
|
||
| sta/radio | 2.32 | 2.21 |
|
||
|
||
**Mixed result -- read honestly:**
|
||
- WIN: 2.4 interference/airtime is clearly down (fewer contending transmitters + smaller cells). On the
|
||
active radios, cu_interf ~64% -> ~48% (snapshot) / 66% -> 32% (morning hourly avg, fewer radios).
|
||
- CONCERN: **retry rose (17.0 -> 23.4%) and satisfaction fell (39 -> 30)** time-of-day-controlled. Likely
|
||
OVER-THINNING: Low = ~6 dBm is a 17 dB cut from auto (~23), and with 24 radios also disabled some edge
|
||
2.4 clients now reach a farther/weaker AP -> more retransmits, lower satisfaction. (Partly composition:
|
||
fewer 2.4 clients today, so the remaining tail may be the stubborn legacy/far devices.)
|
||
|
||
**RECOMMENDATION (needs Howard's go -- this re-check is read-only):**
|
||
1. Bump the KEPT 2.4 radios from Low -> **Medium** (~12-15 dBm): restores client signal while keeping cells
|
||
smaller than full power. `apply-radio cascades ng power medium --zone ...` per floor, then re-measure.
|
||
2. Do NOT expand disables further; consider re-enabling a specific radio if a dead zone/complaint appears.
|
||
3. Re-measure retry/satisfaction after the Low->Medium bump (same time-of-day) to confirm recovery.
|
||
The interference goal is met; the lever now is finding the power floor that keeps cells tight WITHOUT
|
||
starving edge clients. Medium is the likely sweet spot vs the aggressive Low.
|