How to diagnose a setup in which each time the same slot disk of a NAS extension is fried up

I have the following setup :

  • Synology DS920+ NAS hosting two 6TB disks (all disk in this post are Western DIgital Red Pro) in RAID1, two 4TB in RAID 1

  • DX517 extension connected to the NAS, hosting two 22TB in RAID1

  • APC Back UPS – BX2200MI-FR powering the whole stuff

Two times (after a power failure in april 25 and after the death of the UPS’ battery and its replacement this august) after rebooting the NAS following an incident the first slot disk of the NAS extension was fried up (at least I thought) and not seen anymore by the NAS and the volumes concerned by it marked as degraded.

The first time I was in a gigantic hurry for several reasons and bought a new 22TB disk, this time the current prices (went times 3 where I am) the anger and having time made me investigate : I removed the disk, plugged it with a SATA to USB-C to a laptop, rand tests (HDDScan, read, write, erase) during 4 days : the disk is in perfect health.

I put it back in the NAS’ extension, powered up the NAS and the disk wasn’t still seen by the NAS, but its degraded volume was now “repairable” (it wasn’t at reboot after UPC replacement). It is now being repaired and I am pretty confident that the repair will finish with a success, as now more thant 50% of the repair is done and the disk appears again and it’s flagged as healthy.

I would like to know what happened to be able to avoid this happening again. (This costed me 3 days plus an old laptop up since 5 days and it is also a liability).

Remarks :

  1. I don’t think it is a coincidence that each time it was the first slot disk of the extension that went bananas. (After, after reboot, that disk is the first to powered up in the extension.)

  2. This second time it wasn’t the disk. I will exhume the disk from april 25 when I’ll have time and I am convinced that it’ll turn out that the disk is healthy. My guess is that either there’s and issue with the extension and perhaps with its 1st slot or with Synology OS which for some unknown reason marks disks “off” after incidents, marking that a full erasure of the disk as I did is enough to remove.

  3. I must precise that each time all precautions where taken (everything, NAS, UPS unplugged/powered off/on in the right order, setup completely cleaned before power up etc)