PREVENTATIVE MAINTENANCE
Network Infrastructure Health Check Playbook
A repeatable quarterly or annual technical inspection for switches, racks, cabling, fibre, wireless, storage and power.
Switching
- Check critical interface errors and flaps.
- Check stack/ring state.
- Check STP topology and unexpected root changes.
- Check LACP members.
- Check PoE budget.
- Back up configurations.
Cabling and fibre
- Inspect rack patching and labels.
- Check critical fibre receive levels against optic specifications.
- Inspect/clean connectors before intrusive tests.
- Review unresolved failed certification results.
WiFi
- Review AP health and uplinks.
- Review channel utilisation/retries.
- Compare busy-hour coverage complaints with current design.
Storage
- Check pool/RAID state.
- Review drive alerts.
- Review free capacity and growth.
- Confirm backup jobs and perform sampled restore tests.
Power
- Review UPS alarms and battery age.
- Check protected load.
- Review critical PDU/power mapping.
Deliverable
Produce a short exception report: critical, recommended and informational findings. A health check is most useful when it identifies changes since the previous baseline.
Compare, do not just inspect
A health check should compare against the previous review: capacity growth, new errors, firmware age, battery age, failed backups, changed topology and added unmanaged devices.
Risk ranking
Rank findings by business impact and evidence. "Stack ring is half" is more actionable than "switch is old". Recommend corrective work with the reason, not simply a product replacement.
Do not create an outage during inspection
Keep intrusive tests, firmware upgrades and failover tests for approved maintenance windows. A health check can identify the need for those tests without performing them immediately.
Firmware and support state
Record current software/firmware and compare with vendor-supported or recommended releases. Do not upgrade automatically during the assessment. Flag unsupported or security-relevant versions for planned remediation.
Configuration drift
Compare current configuration with the documented baseline. Look for temporary VLANs, open trunk lists, disabled security controls, old static routes and ports that no longer match the patch schedule.
Capacity review
Check uplink utilisation, switch port availability, PoE headroom, storage free space, UPS load, rack U-space and cable-management capacity. Capacity problems are easier to solve before the next project arrives.
Environmental review
Check rack temperature alarms, blocked vents, dust, water risk, loose power leads and unsupported equipment resting on other hardware. Physical issues often precede electronic failure.
Backup and restore evidence
Verify configuration backups for network devices and sampled restore capability for business data. A successful scheduled job without a restore test is incomplete evidence.
Health report structure
- Critical: active service or data-loss risk.
- High: redundancy/security/control not working as designed.
- Medium: capacity or maintainability problem.
- Low: documentation or housekeeping improvement.