KNOWLEDGE CENTRE / TECHNICAL GUIDE

Hot-Swap Drives and Power Supplies Explained

Hot-swap means a supported component can be replaced while the system remains powered. It does not mean every component can be removed safely at any time.

Hot-swap means a supported component can be replaced while the system remains powered. It does not mean every component can be removed safely at any time.

Practical example: small business virtualisation server

Hot-swap maintenance workflow
Hot-swap maintenance workflowHot-pluggable does not mean safe to remove without checking redundancy. IdentifyConfirm failed componentVerifyCheck redundancyReplaceFollow model procedureConfirmRebuild and health Hot-pluggable does not mean safe to remove without checking redundancy.

Hot-pluggable does not mean safe to remove without checking redundancy.

Scenario: A 35-user business wants one physical server for directory services, an accounting application, file services and three small virtual machines.

The design starts with the workload rather than a model number. Memory is budgeted for the hypervisor, each virtual machine and growth. Storage is separated into performance and capacity requirements. RAID is selected for availability, while a separate backup target is retained for recovery. Dual power supplies and multiple network interfaces are considered if the required server supports them.

Decision: The final server should be chosen only after checking processor capacity, DIMM population rules, drive bays and backplane, RAID controller, network interfaces, PSU configuration and the supported expansion path.

Common mistakes and selection checklist

Common mistakes

  • Removing a component because it is labelled hot-plug without confirming the platform state.
  • Replacing the wrong drive in a degraded RAID set.
  • Assuming redundancy has been restored immediately after inserting the replacement.

What happens if you get it wrong?

A maintenance action intended to avoid downtime can instead cause an outage or data loss. A replacement drive may also require a lengthy rebuild before redundancy is restored.

Selection checklist

  • Identify the failed component using the server's management and status indicators.
  • Confirm the exact bay, PSU or module before removal.
  • Verify that the remaining redundant component can carry the workload.
  • Follow the model-specific replacement procedure.
  • Monitor rebuild or recovery until the system reports a healthy redundant state.

Hot-swap requires platform support

The server chassis, backplane, controller, firmware and operating configuration must support the component and the replacement procedure.

Hot-swap drives

A drive in a redundant RAID array may be replaceable while the server remains online. The array must still have sufficient redundancy, and the correct failed drive must be identified before removal.

Hot-swap power supplies

Dell, HPE and Lenovo documentation commonly requires compatible power supplies with matching output capacity or part specifications for redundant configurations. One functioning PSU must be able to support the active system load during replacement.

Replacement is not the end of the job

After replacement, verify that the system recognises the component, redundancy is restored, alerts are cleared and any RAID rebuild completes successfully.

Hot-swap is not risk-free.

Use the exact service procedure for the server model and confirm redundancy before removing a component.

Technical references
  1. Dell PowerEdge PSU Replacement
  2. HPE DL380 Gen10 User Guide
  3. Lenovo ThinkSystem Technical Specifications