October 3, 2026

DSM Condemned a Perfectly Good Drive: Here's How I Got It Back

A field story from a Synology DS1813+ running RAID 6 (SHR-2) on DSM 7.1.1: one bumped SATA link, a drive with zero media faults, and a repair wizard that refused to admit it.

The setup

I've got a DS1813+ running as my backup repository target. The pool was mid-build-out (volume empty, still staging the final drive layout) when somebody bumped the unit. One hard bump.

Within a minute the box was beeping and the Storage Manager page went red: disk 4 dropped off the array. Bay 4 is a WD Red 4TB with about 47,700 power-on hours on it, the oldest drive in the box, and the one I'd been watching anyway.

First: what actually broke

Before touching anything, I pulled the logs over SSH. This matters, because the screen tells you what DSM decided, and the kernel log tells you what actually happened. Those were two very different stories.

The kernel log showed the drive falling off the SATA link at 09:57:54, a plain connection-status event (PHYRdyChg DevExch), the link going down, two failed hard resets, then the drive re-training at 3.0 Gbps a few seconds later. Then it dropped again, DSM's disk watchdog condemned it, pulled it from all three arrays it belonged to, and marked it crashed in the pool database.

A bump. The drive got its link severed twice mid-transaction, and DSM did what DSM is designed to do: assume the worst, protect the pool, scream about it.

Meanwhile SMART on that drive read, and still reads, all zeros. Reallocated sectors: 0. Pending sectors: 0. Offline uncorrectable: 0. CRC errors: 0. Spin retries: 0. Nothing. The drive never logged a single media fault, before, during, or after the event. The pool was empty, so there was no data risk at all. This was a drive I wanted back, not replaced.

The reseat that didn't reseat

I pulled the tray and reseated it. The drive spun up, trained at 3.0 Gbps, and DSM read its model and serial just fine. And then… solid yellow.

DSM's own disk page said the drive was healthy, while also saying deactivated. The storage pool still showed degraded, and clicking Repair said I needed to replace the failed drive. Both messages at once. Very helpful.

So I did what any sensible person does: rebooted the whole unit. Result: nothing. The boot re-assembled the array exactly as before, minus my drive, and re-flagged it from the pool's own database. If you take one lesson from this whole story, take this one: rebooting does not clear a crashed-disk flag on DSM. The verdict lives in persistent state (/etc/space), not in whatever the kernel assembled this morning.

The catch-22

Here's where it gets stupid. DSM's repair wizard will only accept a drive it considers Uninitialized, a blank disk, never part of anything. But my drive was a known former member of this exact pool, and DSM had it tagged Deactivated, still registered as a crashed member, keyed by serial number.

  • Still a member? → Wizard says "replace the failed drive."
  • Cleared of membership? → Wizard has nothing to repair.
  • A formerly-crashed disk that's actually fine? → There is no third option. There is no "retest" button. The drive passed every test; there was nothing to retest. What was stuck was bookkeeping.

Un-condemning it

DSM ships CLI tools that do the same things the GUI does, without the hand-holding. The disk's crashed verdict lives in the pool's space table as a JSON blob. I could see it plain as day: faulty_disks: ['WD-WCC7K1LKR8HH']. That's the whole "counter" everyone asks about. It's not a SMART counter at all. It's a line in a database.

synostgdisk --disk-deactivate /dev/sdd cleared it: the faulty list emptied, the runtime flag file disappeared. Progress. But now the repair backend itself (synostgpool --auto-repair) said "No space can be repaired," because that flow looks for a failed member still attached to the pool, and I'd just un-registered it. Catch-22 confirmed from both directions.

The actual fix

So I stopped asking DSM's permission and ran the same mechanics its own wizard would have run:

  1. Partition the drive to match its twin. The pool's other 4TB drive gave me the reference geometry (8 GB system partition, 2 GB swap, remainder data). I cloned the partition table head and tail (1 MB at each end of the disk) and re-read the table. Three partitions, exact same offsets.
  2. Re-add the drive to all three arrays it belongs in: the two system mirrors and the big RAID-6, with mdadm --add. The kernel accepted all three and started rebuilding immediately. The system partitions finished in minutes. The 18.5 TiB RAID-6 rebuild is the long one.
  3. Let it run. ~40 MB/s on this chassis; that's not a throttle problem, I checked the kernel's rebuild-rate knobs, that's just what a single-missing-member RAID-6 rebuild costs on 5,400 to 5,900 rpm desktop drives. About 26 hours. The drives all blink together during a rebuild because every surviving member is being read to reconstruct the missing one. That's normal, not a fault.

The beeping stopped the moment the crashed flag cleared, which tells you where the alarm was aimed: not at the drive, at the verdict.

What I'd tell someone else

  • The LED and the verdict are about the bookkeeping, not the drive. Solid yellow days after a reseat just means DSM still believes its database. Trust the SMART counters and the kernel log over both.
  • A condemned drive has to become "Uninitialized" before the wizard will touch it, and DSM gives you no GUI path to get there from Deactivated. On an empty pool, wiping the stale member signatures and re-adding via mdadm is the honest version of what Repair does anyway.
  • Rebooting is not a fix for state that lives in /etc/space. I proved it the lazy way so you don't have to.
  • Counters don't move on a link drop. If your "failed" drive shows zero defect counters before and after the event, you didn't lose a drive; you lost a connection, and something (a bump, a tray, a backplane) owes the array an apology.
  • Time it while the pool is empty. The only reason this was a 26-hour inconvenience instead of a scary decision is that I did it before the backup repository went back on the array. Do your rebuild windows when you have nothing to lose, and they stop being emergencies.

The drive is back in the array rebuilding as I write this, SMART counters still zeros across the board, temperature normal. When it finishes, the pool goes back to full width and DSM finally gets to be right about something.

Test environment: Synology DS1813+, DSM 7.1.1-42962, SHR-2 (RAID 6), WD Red 4TB (WD40EFRX), October 2026. Pool volume empty throughout: zero data exposure.

DSM Condemned a Perfectly Good Drive: Here's How I Got It Back | Tech Connect Arizona