Files
claudetools/clients/cascades-tucson/reports/2026-06-16-unifi-full-audit.md
Howard Enos 58ecc5ad40 unifi-wifi: pfSense gateway access via SSH (pfSense-ssh.sh) + pfSense health section; layer OFF HOLD
DECISION (Mike, 2026-06-16): drop the RESTAPI package — VPN + SSH shell reads the same data and makes
changes. Confirmed Cascades pfSense is Plus 25.07-RELEASE (current; the "too old" premise was wrong) and
admin SSH = real shell (no menu). The upgrade/package blocker is moot; compat layer is off hold.

- NEW scripts/pfsense-ssh.sh: audit (version/WAN-media/gateway-events/DHCP-exhaustion/states/DNS/load/NIC),
  dhcp (pool utilization + no-free-leases), run "<cmd>" (arbitrary, incl changes; operator-gated). Cred
  from clients/<slug>/pfsense-firewall; system OpenSSH via askpass. Validated live on Cascades.
- audit report: added "pfSense health check (2026-06-16)" — DHCP NOT exhausted (192.168.0.0/22 pool 270/507,
  0 no-free-leases), DNS up, dual-WAN stable (no gateway flaps), states/load healthy => gateway is NOT a
  WiFi factor; the 2.4 GHz RF work is the sole fix. (Minor: igc3/WAN2 I225 2.5G counter quirk, not a fault.)
- ROADMAP §E + SKILL.md updated to the SSH backend decision; REST pfsense-backend.sh kept dormant/optional.
- Remaining: named gated CONTROL verbs over SSH (easyrule block-ips, pf/fw toggles) + optional gw-* dispatch.
- Closed obsolete coord todo (upgrade-pfSense-for-RESTAPI).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 18:42:54 -07:00

5.4 KiB
Raw Blame History

Cascades of Tucson — UniFi Full Audit (2026-06-16)

Generated by the unifi-wifi skill (read-only). Fleet: 77 U7-Pro APs, 12 switches, ~587 clients, no UniFi gateway (pfSense firewall). All collectors ran clean.

Issues — prioritized

# Sev Issue Detail / fix
1 HIGH 2.4 GHz saturation (the "bad for some users") 75 radios at auto/full power, ch6/1/11 at 16k33k neighbor BSSIDs; 123 clients on 2.4 at ~10% retry. Fix: power-down 65 radios to Low, then disable 9 redundant (plan below).
2 HIGH ~25 switch ports linked at 100 M but gig-capable FastEthernet/cabling-or-NIC issue capping those APs/devices at 100 Mbps (1st/2nd/3rd-floor switches). Physical: re-terminate/replace cable or check NIC.
3 MED 6 GHz essentially unused — 1 client of 587 75 6E radios live, nearly empty. Enable band-steering (bandsteer/bands) to offload 5 GHz onto the clean band.
4 MED 5 GHz on 80 MHz (76/77) + 55 on DFS 80 MHz kills spatial reuse in density → 40 MHz. DFS empirically low-risk here (0 radar events) but move to non-DFS for resilience (near Davis-Monthan).
5 MED 6 APs with 2.4 min-RSSI OFF; 4 APs off the 1/6/11 plan min-RSSI OFF: 615, 608, 505, 517, 622, salon. Off-plan (auto): 128, 108, 108U7-Pro, salon.
6 LOW 3 offline switches + 2 disconnected APs + 1 firmware update Offline: Switch 2nd Floor #2, Switch 4th Floor #2, USW Pro Max 16. APs: 108 ×2 (cable run pending, known). 1 device upgradable.
7 LOW p38 (1st Floor USW) 4.0% tx-drop rate Correlates with the underspeed/heavy-traffic ports; investigate after #2.

WiFi detail

  • 2.4 GHz: 77 radios, all 20 MHz (good); power auto×75 (want Low); channels 1:20 / 6:28 / 11:25 / auto:4. min-RSSI OFF on 6. Neighbor density: ch6 33,376 · ch1 19,355 · ch11 16,598 BSSIDs. Live retry avg 10.2%.
  • 5 GHz: 77 radios, 80 MHz ×76 / 40 MHz ×1; rogue density biased to ch149/157 (busy upper). 463 clients. retry avg 8.0%.
  • 6 GHz: 75 radios active, 1 client. Wide open.
  • AP satisfaction (live): min 90 / median 98 / max 100 → healthy in aggregate; the pain is the 2.4 GHz client tail.

Data-backed radio plan (optimize-radios + /proc/ui_neighbor SNR matrix)

  • Phase A — power-down 65 2.4 radios to Low (smaller cells cut mutual interference; coverage-safe).
  • Phase C — disable 9 redundant 2.4 radios after re-measure (each heard by ≥2 strong neighbors): 127→128, 229→128, 248→348, 330→128, 445→347/348/247, 428→128, 622→505/615/608, Kitchen→Memcare TV room, Dining Room→memcare piano. Est. interference-airtime removed: ~619.
  • Channel plan available: 2.4 GHz 1/6/11 graph-color (co-channel pairs 92→35); 5 GHz non-DFS (20→0 and all off DFS).

Switch / PoE (12 switches, 29 flags)

  • ~25 ports at 100 M but gig-capable (see #2). PoE budgets healthy (e.g. 1st-floor 160/600 W).
  • 3 offline switches (above).

Gateway / WAN

  • No UniFi gateway (pfSense) → WAN/internet not measurable via UniFi. Adoption: APs 77 (2 disc), switches 12 (3 disc), 587 clients.
  1. Physical: fix the ~25 underspeed ports (#2) + chase the 3 offline switches / AP 108 cable.
  2. WiFi Phase A: power-down 2.4 to Low per zone, validate with watch-ap (live before/after).
  3. Enable 6 GHz band-steering + 5 GHz 80→40 MHz non-DFS channel plan.
  4. Set 2.4 min-RSSI on the 6 OFF APs; pin the 4 off-plan APs to 1/6/11.
  5. Phase C: disable the 9 redundant 2.4 radios after re-measure.

(All changes via the gated apply-radio/apply-wlan/channel-plan scripts — per zone, with rollback + live validation. Nothing applied in this audit.)


pfSense health check (2026-06-16) — ruling out the gateway as a WiFi factor

Investigated the Cascades pfSense (192.168.0.1, pfSense Plus 25.07-RELEASE, Netgate) over the site VPN via SSH, to confirm whether any gateway-side issue contributes to the "WiFi bad for some users" symptom. Verdict: pfSense is healthy and is NOT a contributor — the problem is RF-side (2.4 GHz).

Area Finding WiFi impact
DHCP exhaustion 0 "no free leases" events in dhcpd.log. WiFi/AP pool 192.168.0.0/22 (range 192.168.2.23.254, cap ~507) only 270 active (~53%); per-unit /28s + 10.0.20/.50 all have headroom Ruled out (was the top suspect)
DNS unbound resolver running Fine
WAN Dual Cox — WAN1 184.191.143.62/30, WAN2 72.211.21.217/27, both active full-duplex, WAN_Group gateway group, no loss/down events logged Fine
Firewall states 28,368 / 790,000 limit Fine
CPU / mbuf / uptime load 0.6, mbufs nominal, 10-day uptime Healthy

Architecture: per-unit design — 199 DHCP subnets, mostly 10.x.y.0/28 per apartment (assisted- living L2 isolation) + the 192.168.0.0/22 staff/AP network (APs + most WiFi clients). Active DHCP backend is ISC (Kea config present but dormant).

Minor (not WiFi-related): igc3/WAN2 logged 1707 input-errors + 1707 "collisions", but the link is 2.5GbE full-duplex/active with zero gateway loss — consistent with the known Intel I225/I226 2.5G counter quirk, not a real fault. No action needed unless WAN2 misbehaves.

Conclusion: gateway/DHCP/DNS/WAN are not bottlenecking the wireless. The 2.4 GHz remediation (power-down + coverage-redundancy disables) remains the correct and sole fix for the client-experience tail.