How to Check SSD Health Before It Fails

How to Check SSD Health Before It Fails

Kevin Wu
September 6, 2026· 9 min read

To check SSD health, first make sure irreplaceable data exists on another device, then identify the exact physical drive and read the operating system’s health status plus any model-supported SMART, NVMe, or manufacturer diagnostics. Treat those values as evidence, not a promise: a drive can report healthy immediately before an abrupt controller, power, firmware, or NAND failure.

Key Takeaways

  • Backup comes before benchmarking, repair, firmware work, or repeated scans.
  • Record the exact model, serial suffix, capacity, connection, and volume before reading a status.
  • Combine health fields with I/O errors, disconnects, read-only changes, temperature, and trend.
  • USB enclosures may hide or translate SMART and NVMe telemetry.
  • A healthy indicator reduces uncertainty; it never replaces a tested backup.

Use the device and app troubleshooting guide to identify the failing layer. This guide checks an SSD that the system can still enumerate. A disk that never appears belongs in the external-drive visibility guide.

What should you do before checking the drive?

If the SSD contains the only copy of important files, do not begin with a full-surface scan, sustained benchmark, firmware update, repair utility, or repeated power cycle. Copy the highest-value readable data to separate storage and open a representative sample there. When reads are unstable or the data is irreplaceable, stop and consider professional recovery instead of consuming the remaining stable time.

Identify the physical device rather than trusting a familiar volume name. Record:

  • manufacturer and exact model;
  • capacity and, when safe to reveal, the last few serial characters;
  • internal NVMe/SATA or external USB/Thunderbolt connection;
  • enclosure or adapter model for an external SSD;
  • operating-system and firmware versions;
  • every volume currently hosted by that physical disk;
  • the first symptom and its time.

This prevents a common mistake: reading the health of the computer’s internal SSD while troubleshooting an external disk, or running a command against a similarly sized backup drive.

Which built-in tools check SSD health?

Start with read-only information already exposed by the operating system. On macOS, Disk Utility’s information panel can show a device’s SMART status when the storage path supplies it. Apple says a “failing” SMART status means the disk should be backed up and replaced.[1] A status may be absent for an external enclosure, and absence does not mean failure or health.

On Windows, the Storage PowerShell module can list physical disks and fields such as health status, operational status, media type, and size.[2] Use an elevated shell only if the supported command requires it, and begin with listing commands rather than repair or reset operations. Graphical storage settings or the device manufacturer’s official utility may provide the same information more safely for a nontechnical user.

Linux tools vary by distribution, kernel, controller, and transport. Use the distribution’s maintained package and the drive or enclosure maker’s documentation. Do not paste a destructive command from a forum because it contains the word “SMART.” Confirm that the command is read-only and targets the correct device.

For all platforms, save the complete output with a date, model, connection path, and workload state. A single green badge is less useful than a comparable record taken again after a symptom.

Which SSD SMART or NVMe fields are useful?

Field names and thresholds are vendor-specific. NVMe defines a SMART / Health Information log with critical warnings and values such as temperature, available spare, percentage used, data units, power-on hours, unsafe shutdowns, media and data integrity errors, and error-log entry counts.[3] Not every number is a failure prediction.

EvidenceWhat it can tell youWhat it cannot prove
Critical warningController has raised a defined health conditionExact remaining life
Available spareSpare capacity relative to the vendor thresholdThat all existing data is readable
Percentage usedVendor estimate of endurance consumedA precise failure date
Media/data-integrity errorsController recorded uncorrected data eventsWhich file is affected without further evidence
TemperatureCurrent or recorded thermal conditionThat heat caused every observed error
Unsafe shutdownsPower was lost without a clean shutdownThat the SSD itself caused the loss
Error-log entriesCommands generated logged errorsSeverity without the command and status context

For SATA SSDs, attributes such as reallocated or uncorrectable errors, wear indicators, CRC errors, and temperature may be useful, but raw values and normalized scales differ. Interpret them with the exact manufacturer’s documentation. A generic program that labels every nonzero raw value “bad” can create false alarms.

USB-to-SATA and USB-to-NVMe bridges may block, translate, or incompletely expose telemetry. Compare the enclosure’s supported features before concluding that an empty SMART page proves the drive is healthy. Do not remove an SSD from a sealed or hardware-encrypted enclosure unless the vendor supports it and data is backed up.

How do symptoms and trends change the assessment?

Health telemetry matters most when tied to an observable event. Create a short timeline with the action, file direction, elapsed time, temperature, connection path, system message, and whether the device or only the volume disappeared.

Escalate the situation when you see any of these patterns:

  • new media or data-integrity errors;
  • repeated I/O errors while reading ordinary files;
  • a drive that becomes read-only without an intentional policy;
  • capacity, model, or namespace information changing unexpectedly;
  • critical warning or failing status;
  • repeated disconnects across known-good cables and hosts;
  • files that no longer open or compare with a trusted copy;
  • abnormal heat, odor, discoloration, or damaged connectors.

A steadily increasing error count is more concerning than a stable historical count, but only compare equivalent readings from the same model, tool version, and connection. Resetting logs, changing enclosures, or switching drivers can break the trend.

If an external SSD disconnects under load, follow the transfer-disconnection guide. It separates cable, port, power, enclosure, heat, sleep, filesystem, and media evidence. Health data alone cannot isolate those layers.

Should you benchmark or scan the whole SSD?

Not as the first health check. A full write benchmark changes data and adds wear. A sustained read scan adds heat and load and may be the wrong choice when the device is already unstable. Filesystem repair changes metadata but does not test every flash cell or controller path.

After data is safe, a short read-only manufacturer self-test or supported diagnostic can add evidence. Read its documentation first: confirm whether it writes data, how long it runs, whether stable external power is required, and whether it supports the exact model and enclosure.

If performance is the only concern, establish a small, disposable workload and compare it with the manufacturer’s conditions. Nearly full capacity, thermal throttling, background indexing, encryption, small random files, a slow USB bridge, and operating-system caching can all change speed without proving imminent failure.

Do not use repeated benchmarks to “see whether it gets worse.” One reproducible failure plus protected data is more useful than ten uncontrolled stress runs.

When should you replace, monitor, or seek recovery?

Replace the SSD after backing up when the system or manufacturer reports a critical failure, integrity errors rise, ordinary reads fail, the disk becomes unexpectedly read-only, or the same device fails through independent known-good paths. Preserve purchase details and a minimal non-sensitive reproduction for warranty support.

Monitor rather than immediately replace when the only observation is an unchanged historical counter, the manufacturer documents it as normal, and files, self-tests, system logs, temperature, and connections show no related symptoms. Record a new baseline and shorten the interval until the trend is understood.

Seek recovery help before further testing when the only copy is valuable and reads are unstable, the device disappears repeatedly, capacity changes, hardware is damaged, or encryption depends on the original controller or enclosure. Do not initialize, secure-erase, update firmware, or open sealed hardware first.

The data-corruption guide explains why redundancy, verification, and recovery planning matter beyond one health reading. A backup is only useful if its files and credentials can actually be restored.

Summary

  • Protect the data and identify the exact physical SSD before reading health information.
  • Start with built-in, read-only status and the manufacturer’s supported diagnostic path.
  • Interpret SMART and NVMe values using model-specific definitions.
  • Combine telemetry with errors, temperature, disconnects, read-only changes, and trends.
  • Avoid destructive benchmarks, repair, firmware changes, and secure erase before backup.
  • Replace or seek recovery when critical warnings or repeatable cross-path faults appear.

FAQ

Can an SSD fail while SMART says it is healthy?

Yes. Telemetry can reveal known conditions, but sudden controller, firmware, power, connector, or flash failures may occur without a useful warning. Keep a tested backup.

What SSD health percentage is bad?

There is no universal percentage. Use the exact manufacturer’s meaning, threshold, warranty terms, and associated warning fields rather than applying one number to every model.

Does percentage used mean my files are about to disappear?

No. It is generally an endurance estimate, not a countdown to failure. Consider critical warnings, integrity errors, spare capacity, symptoms, and backup status together.

Why is SMART unavailable for my external SSD?

The USB or Thunderbolt enclosure, bridge, driver, or operating system may not pass the command through. Check the enclosure and SSD maker’s supported diagnostic method.

Is a slow SSD necessarily failing?

No. Free space, heat, interface speed, encryption, background work, file size, and cache behavior affect performance. Errors and repeatable abnormal changes provide stronger evidence.

Should I run a full surface scan?

Only after important data is safe and the exact supported tool and risk are understood. A full scan adds load and may be harmful when the device is unstable.

Can filesystem repair fix SSD health?

Filesystem repair can address logical metadata inconsistencies; it cannot repair worn flash, a failing controller, unstable power, or a damaged connector. It also writes metadata.

How often should I check SSD health?

Check after a new warning or symptom and periodically according to the device’s importance and vendor guidance. Automated alerts and verified backups are more useful than obsessive manual checks.

Sources

  1. Apple Support - Check if a Mac disk is about to fail — https://support.apple.com/guide/mac-help/check-if-a-mac-disk-is-about-to-fail-mchlp2548/mac
  2. Microsoft Learn - Get-PhysicalDisk — https://learn.microsoft.com/en-us/powershell/module/storage/get-physicaldisk?view=windowsserver2025-ps
  3. NVM Express - NVM Express Base Specification Revision 2.2 — https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-Revision-2.2-2025.03.11-Ratified.pdf

Sources checked 6 September 2026.

Related Articles

Start your 3-day free trial

Sign up to experience all premium features at no cost.

*Available only to new users. Each user is limited to one trial.

How to Check SSD Health Before It Fails | AethoVPN