A RAID 5 is built to shrug off one dead disk, which is exactly why so many of them run degraded for months with nobody the wiser. Then a second member goes, or a rebuild sticks at forty per cent, and a volume nobody backed up because it was thought to be safe simply stops existing. Here is what one calculated block per stripe can and cannot do, why the repair is riskier than the failure, and how five drives are put back together on copies.
RAID 5 opens at £500 + VAT and climbs with the number of members, confirmed in writing when the free examination closes on the second working day after your disks are logged in. Member drives only, each marked with the bay it came out of — no chassis, no controller card.
Parity sets do not arrive here in a hundred varieties. There are roughly six, and yours decides both what the work involves and how much of the damage is still reversible. Read until something sounds familiar, then take your hands off the array.
One member is out and parity is standing in for it. Nothing has been lost, everything is still reachable, and this is the hour worth acting in — which is why it gets thrown away. No one feels an emergency while the files still open. Get today working copy onto separate storage before a replacement disk goes anywhere near a bay.
The sums ran out. A single gap in a stripe can be calculated; two cannot, so the set is dropped rather than served up wrong. On the console that reads as final and it very seldom is, because second failures are usually partial. Image that disk, borrow parity from the rest, and the volume assembles away from your hardware.
Rebuilding asks each surviving member for every sector it owns, front to back, which is more reading than those drives have done since the day they went into the cabinet. A weak one gives up part way. If the bar has sat at 40 per cent since last night, cut the power rather than pressing start, because start begins the identical punishing sweep again from zero.
Swap a controller, reflash one, clear one or move it to another chassis and it comes up not recognising what is in front of it, then offers to import, to clear or to create. Every one of those three is a write. The description it has forgotten is still on the members, so nothing has to be guessed at unless somebody clicks.
Late, with the business stopped, rebuilding the array from the same disks is the last idea in the room. Occasionally it costs nothing whatever. Occasionally it lays a fresh parity pattern straight across live data. The whole difference is which initialisation option was on screen, and a diagnostic establishes that where a guess cannot.
Not a hardware fault anywhere in the chassis. A LUN dropped during a migration, a partition formatted by somebody working from the wrong list, a file system left inconsistent by a power cut off the Foleshill Road, malware working through the shares while the building was empty. Nothing is opened for any of it, the work still runs on images, and the band is still the array band.
Take a volume, cut it into stripes and lay those stripes across three disks or more. Alongside each stripe RAID 5 stores one extra block, and that block is not a spare copy of anything. It is the exclusive-or of everything else in the same stripe: one value from which any single missing piece of that stripe can be worked out again. Which disk holds the calculated block shifts as you travel down the volume, so no member becomes the bottleneck for every write, and the cost of that scheme is one disk out of the total. Five 4TB drives give you 16TB rather than 20TB. The missing 4TB is the premium.
What the premium buys is one absent member. Lose it and the controller works every read out on the fly from what remains, so the array keeps serving files, more slowly, with nothing at all held back. Lose a second and each stripe now presents two unknowns against a single equation. Nothing resolves that from parity — not a better controller, not a laboratory, not money — which is why the set is taken offline rather than allowed to hand out plausible rubbish.
Worth stating bluntly, because the opposite belief starts most of the calls that come in. A parity set has a view on precisely one event: a disk dying. It has no view whatever on a folder dragged into the bin at twenty past four, a database left inconsistent by a dirty shutdown, a volume formatted during a migration weekend, a fire in a comms cupboard, or ransomware grinding through the shares over a bank holiday. Every one of those is written faithfully and instantly to all the members together, calculated blocks included.
Uptime is what an array sells. A backup is a second copy somewhere the same accident cannot reach it. They are separate purchases, whatever the box implied, and the firms that come out of an array failure in good order are invariably the ones that bought both.
RAID 6 carries two calculated blocks per stripe rather than one. It therefore tolerates two absences and, far more usefully, survives a second failure occurring during a rebuild — which is the failure that actually turns up. RAID 10 puts mirrored pairs underneath a stripe, so it can lose a disk out of each pair and continue, and it rebuilds by copying one drive rather than reading all of them, which is a much gentler afternoon for the hardware. A plain mirror gives the kindest recovery in the trade and has a page to itself at RAID 1 recovery.
At the far end is RAID 0, which holds nothing back at all and is the array fault Coventry searches for hardest. If your volume disappeared the instant one drive stopped, with no amber-light spell in which things still worked, the set is striped rather than parity-protected and that page is the one you want.
Slotting a replacement in and pressing rebuild feels like the grown-up response, and it is the usual route by which one dead disk becomes a dead array. The reason is arithmetic rather than misfortune. To reconstruct a stripe the controller must read every sector of every remaining member without once meeting an error it cannot get past. Across five large modern drives that means terabytes of uninterrupted reading, held up for hours or days, from disks bought in a single batch, racked on a single afternoon, and turning identical hours in the same warm cabinet ever since.
Manufacturers publish a rate for sectors that will not read. The number is small and it is not zero, and small numbers multiplied by terabytes are how a rebuild discovers a fault nobody knew about. Age is the part that gets underestimated. When one drive from a matched set reaches the end of its useful life, the rest are usually weeks behind it, and the workload that proves the point is the rebuild somebody just kicked off.
Unless the array is deliberately taken out of service, the survivors do that heavy reading on top of the ordinary working day — users opening files, a backup window, a database committing. Should a second member drop mid-pass, what is left is a half-written replacement, a set that will not assemble, and members that no longer agree about what the volume was. Recoverable in the great majority of cases, and a longer, more delicate job than it ever needed to be.
Copy before anything else. A degraded array is still handing over every byte you own, so put the working data onto whatever has room for it, today. A plain external disk bought at lunchtime is a better decision than anything you will consider later. Once that copy exists, a rebuild becomes a reasonable idea.
Where the material has serious value and no restore has ever been proved, have the members imaged before the rebuild rather than after it. An image lifted from a healthy degraded set is a clean starting point; one lifted after a failed rebuild is a starting point with damage baked in.
If the array is already down, everything useful is a subtraction. Stop forcing members online, because each forced assembly writes metadata and can move the sequence numbers recording which drive holds stale content — and those numbers are often exactly what tells a bench the right order. Refuse every dialogue offering to initialise, clear or create; a quick initialise generally touches metadata alone and can be lived with, a full one writes zeros through the data region and cannot. Leave the disks in their bays instead of pulling the lot out to read labels. And keep repair utilities away from a volume that will not mount, because CHKDSK, fsck and the commercial equivalents all assume sound storage beneath them and write their verdicts to every member at once — what CHKDSK really does goes through that in detail.
Your drives are not assembled, mounted or mended. Each member is read across to laboratory storage on its own, on equipment designed around disks that stall, vanish from the bus mid-transfer or need thirty seconds to produce a single sector. Healthy members go over quickly. A member with weak ground is taken in stages — the sound territory banked at the first attempt, the difficult territory revisited later with tighter timeouts and one head at a time — so that a drive with few hours left in it spends those hours on ground nobody has covered yet.
After that, your disks sit on a shelf untouched for the rest of the job. Everything downstream happens against the copies, which is what allows a wrong turn to cost an afternoon of computing instead of your files, and what makes a second attempt available whenever the first proves unsatisfying. It is also why a laboratory can be patient where an administrator at midnight cannot.
An array is a rule for showing several disks as one volume, and the rule has to be established before any filename exists. Block size. The sequence the members sit in. Which way the calculated block rotates as you descend the volume, and whether data follows it in step. And the offset, meaning how far into each drive the array region starts, because controllers keep space at the front for themselves. One of those four wrong and the volume either declines to mount or — considerably worse — mounts and issues files quietly padded with somebody else stripes. That second result is the one to be afraid of, since it wears the appearance of success.
Metadata gets read and never believed on its own. Dell PERC, HP Smart Array, LSI and Adaptec all keep a description in reserved space on the members; Linux software arrays keep an md superblock; several manufacturers share one on-disk format between them. All of that is a hypothesis. It is then tested against the contents themselves: do file system structures fall exactly where this geometry says they should, does parity validate across whole stripes, do directory records point at files that open. Where description and content disagree, the content is what counts.
Here is parity used as it was designed to be used — once, slowly, against copies, rather than under load on live hardware. With the geometry settled, blocks missing from a damaged member are calculated back from the survivors, so a drive that reads only part of the way through still contributes everything it managed and the holes are filled by arithmetic. That is frequently the difference between a partial result and a complete one, and it stays available only because nobody rebuilt onto the originals first.
After a clean two-member failure the normal outcome is the whole volume with its folder tree and its filenames exactly as they were, because the file system was never harmed at all — the array merely could not be presented. Deleted material, a formatted volume or an encryption run vary far more, and the diagnostic identifies which of those you have before any figure is settled.
Three endings really are bad, and putting them here is fairer than letting you find them on an invoice. A rebuild that ran all the way through with the members in the wrong sequence overwrites as it travels. A full initialise, as distinct from a quick one, puts zeros through the data region. And on a thin-provisioned volume, space that was never written to holds nothing to find, because nothing was ever there. Everything else is a question of hours rather than of possibility.
Ransomware on an array is charged as array work from £500 + VAT. It is a recovery, and dressing it up as an investigation would only be a way of charging more for identical bench time. Each band is set out again on the cost page.
Looking costs nothing. The examination finishes on the second working day after your disks are logged in, and it hands over one fixed figure in writing next to a realistic date. Neither is provisional and nothing at all starts until you have accepted both. A lone drive is usually done two to four working days past authorisation. A set runs longer, for the plain reason that five members means five imaging passes before anybody can start testing geometries.
No fix, no fee applies to logical work — a set that will not assemble although the disks are sound, a deleted volume, a broken file system, an encryption run. Standing outside it are chip-level work, DVR jobs, forensic jobs and any failure that is mechanical or electronic; where a member has died physically, half the agreed figure falls due before invasive work begins.
Only the members travel. Chassis, rails, controller card and power supplies all stay in your cabinet, since none of them get used at this end and every one of them turns a modest parcel into freight. Mark each drive with the bay it occupied, reading across the front of the unit, and take a photograph of that front panel before anything is withdrawn so the marks can be checked afterwards. Now and again that photograph is worth a day.
Wrap them individually and pack until nothing can shift inside one rigid box. Put the booking-in form in with them, plus the controller model, the level if it is known to you, and the order in which members dropped. Two accurate sentences of history regularly save hours at the far end. Most firms use Special Delivery, which is tracked and covered. If handing them over suits you better, Oxford takes drop-offs on weekdays from 9:00am to 5:30pm and the run from Coventry is around fifty-five miles of M40, an hour or so. No Coventry counter exists, and no one in this network collects.
They are dull and business-critical in equal measure. Four bays holding twenty years of drawings in an engineering shop between Nuneaton and Bedworth. The accounts server in a practice off the ring road. A studio near the cathedral quarter whose artwork volume doubles as its archive. Warehouse systems on the estates ringing the M6 and M69, running the pick lists on hardware nobody has restarted in three years. Test output from the motor trade at Whitley and Gaydon. In nearly every case the array had behaved so well for so long that somebody stopped opening the backup report, and the restore turned out to be the second thing that failed that week.
There is a longer account of the failure itself in what goes wrong during a RAID 5 rebuild, and the broader service pages are RAID recovery and server data recovery.
A degraded array is still handing over every byte you own. Switching it off costs an afternoon of access; rebuilding onto a tired member can cost the volume.
Multi-disk jobs travel as a set of bare drives sharing one padded box, and that parcel wants tracking and covering for what is written on it rather than what it weighs. Booked over a Coventry post office counter in the afternoon, it usually reaches the Oxford bench the next working morning — faster than most firms manage to spare somebody for the run down the M40.
As a rule the storage comes out and the machine stays where it is. That applies to a laptop, a tower, an iMac and to the recorder sitting under a counter. Taking equipment apart is not something this bench does, and a repair shop will free a drive in a few minutes. Two things go the other way: an external drive stays sealed inside its own case, and a NAS travels as a complete unit with its disks still in their bays. A Fusion Mac is a third case — both of its drives come out and travel together, each one labelled. The single situation nobody can work around is memory soldered flat onto a mainboard, which is how Apple Silicon Macs and a good many slim laptops are built: if the storage will not unbolt, there is no parcel to send.
↓ Print the shipping & booking-in form (PDF)
Put Oxford Data Recovery on the label. From Coventry it is roughly fifty-five miles straight down the M40, about an hour if you would rather drive it in than post it. Either way you are told the moment it is logged, and the free diagnostic finishes two working days later.
Unsure what ought to go in the box? Ring 0800 689 0668 before you seal it, or work through the free online diagnostic and let it do the asking.