Home / Devices / SAN

SAN Data Recovery Coventry

Shared storage fails in a way that is hard to miss: the moment it stops, everything sitting on top of it stops in the same instant. Pool metadata that will not parse, LUNs refusing to come online, a datastore full of virtual machines the hypervisor declines to mount. All of that is rebuilt here for companies and IT providers across Coventry and Warwickshire, from one shelf to a full rack, as routine weekly work rather than something squeezed in as a favour.

Looking at a SAN volume costs nothing. What comes out of that examination is one written figure, agreed between us before anybody picks up a screwdriver: from £500 + VAT on a SAN, the figure following the number of disks in the pool.

Logical work is no fix, no fee. Four things sit outside that: electronic and mechanical failures, chip-level work, DVR jobs and forensic work, and anything physical is half up front. All five bands are laid out on the data recovery cost page, while data recovery services follows a postal job from the parcel to the answer.

// thirty faults, roughly by how often they turn up

Thirty ways it fails, and what lies behind each one

Tracing a symptom back to whatever caused it is where the job genuinely starts, and these thirty cover very nearly every box opened on the Oxford bench. A fault that is not on the list is not an unfamiliar one: describe it over the telephone and you will get a straight reading of the odds before you have spent a penny on postage.

A LUN that stopped presenting itself to every host at once

Half a dozen servers report the same fault within a minute of one another and none of them can see the volume they boot from. That is what a lost LUN looks like from above. Underneath it, the cause is far more often a mapping or a metadata fault than a shelf of dead drives, and the drives are usually in good order.

Pool metadata that no longer parses

Enterprise arrays keep a description of themselves — which extents belong to which volume, where each one starts, what has been thin-provisioned and what has not. Damage there makes perfectly healthy disks present nothing at all. Rebuilding that description from the members is the substance of most of the work on this bench.

Four disks out of one shelf inside the same minute

Backplanes, expanders and power supplies take groups of drives out together, and on the console that reads as a catastrophe. It usually is not. The disks themselves are commonly undamaged, and every one of them is imaged and the pool put back from the copies regardless of what the array said about them.

A firmware update that killed the surviving controller

Dual-controller shelves are updated one head at a time so service continues, and an update that goes badly on the second head leaves the whole thing dark. None of the recovery depends on either controller working, because the layout is derived from the disks. The urge to fit a third head from somewhere is the risk here.

A datastore the hypervisor refuses to mount

VMFS has its own structures and its own ways of breaking, and a datastore that will not mount is very often intact underneath the error. The virtual disks are pulled straight out of an image of the LUN. The layer above that, inside each guest, has a page to itself on this site.

A volume unmapped during a provisioning change

Somebody was reclaiming space, or tidying, and a LUN that turned out to be in use went with it. What that writes is configuration rather than content, so a shelf taken out of service quickly is usually recoverable whole. What costs is the days it keeps serving other traffic afterwards.

A thin pool that ran out of real blocks

Thin volumes advertise more capacity than physically exists, which works until the pool behind them fills. Writes then fail in ways the hosts were never written to expect, and file systems on top go inconsistent within minutes. It is recoverable, and it is one of the commoner reasons shared storage stops without a single disk having failed.

A snapshot chain with a link missing

Snapshots refer back to blocks in earlier states, so a broken chain can make the current view unreadable while every block involved still exists somewhere on the disks. Putting the chain back together from the metadata is the job, and where it succeeds it usually returns the lot.

Deduplication tables damaged, and every file affected

Where one copy of a block is referenced from a thousand places, damage to the index takes out far more than its own size suggests. Rebuilding those reference tables is the difference between a shelf full of readable disks and a shelf full of readable nonsense, and on a dedupe appliance it is the whole of the work.

The fast tier gone, and a piece missing from nearly every file

Automatic tiering keeps busy blocks on flash and idle ones on spinning disks, so a volume lives on both at once. Lose the flash and what goes missing is a slice out of nearly every file rather than a list of files. Send both tiers, and mark which drives came from which.

Targets that vanished from the network overnight

A shelf that has disappeared from every host is not necessarily a shelf with a fault. Switch failures, address clashes and an undocumented network change all look identical from the server end. Ruling that out before anything is unplugged costs nothing and now and then ends the matter there.

A zoning or masking change that hid the storage

Fabric changes remove a host's view of a volume without touching a byte of it. From the host it is indistinguishable from a dead array. Confirming the configuration first is worth the ten minutes, because it is easily the cheapest possible outcome and it is not a rare one.

An expansion shelf unplugged while the pool was live

Pulling a shelf out of a running pool leaves the array with half its members and a volume that no longer adds up. Plugging it back in afterwards does not always restore matters, because the array may have written in the interval. Both shelves travel, both labelled.

An array built on volumes presented by another array

Layered arrangements need unpicking one storey at a time, and each storey has its own mapping to reconstruct. It is slower rather than harder. What it does need is a description on the booking form, because guessing at how the layers were stacked costs a day that need not be spent.

Unreadable regions spread across most of the members

Years of service without a media scan leave every drive in a shelf with its own scatter of dead patches. Not one reads the whole way through, so the surviving ground on all of them is stitched into a single complete picture. That is a normal week here.

A rebuild running against a pool that could not take it

Reconstruction reads every remaining member at full rate for hours, and that is precisely the load a second marginal disk fails under. Once one has stalled, neither the array nor the person who owns it can give an accurate account of what state the pool is in. Halt it, then pull the power.

Guests that will not start once the storage is back

The LUN mounts, the datastore appears, and the virtual machines are broken. That is a second problem with its own tools. Virtual disk files are repaired one at a time after the volume underneath them is readable, and a guest whose file system was mid-write when the storage went is repaired inside as well.

A database that will not open after the outage

Databases are held open and written to constantly, which makes them the first thing a storage failure ruins and one of the more repairable. Transaction logs frequently carry enough to bring the file forward, and older consistent copies survive in unallocated space more often than anybody expects.

A platform whose vendor no longer exists

Shelves from firms that were bought, merged or wound up turn up here regularly, along with kit that went out of support a decade ago. None of the reconstruction depends on a live contract or on the manufacturer still trading, because it works from the disks. Age is rarely the obstacle it looks like.

Array-level encryption with the key inside a dead controller

Self-encrypting drives and array-level encryption both put a key between the platters and the data. Get the key back off a head unit or out of a key manager and the volume opens as normal. Where it has truly gone, you hear so at the examination and not after an invoice.

Ransomware that travelled down a mounted volume

A compromised host with a LUN attached passes the attack straight through to the storage, and the array itself is physically untouched afterwards. Where the compromised account was not privileged enough to delete snapshots, those come through untouched. This is ordinary array work at £500 + VAT and upward, not an investigation.

A shelf that was moved and never came back up

Disks that have spun without interruption for six or seven years often carry wear that only declares itself when they are stopped and started again. A relocation, or a long power cut, is regularly the last event before several members give up at once. The restart exposed it; the journey did not cause it.

Drives reseated into the wrong slots

Some platforms record their own position and cope. Others do not, and a shelf that has been reordered can end up rewriting its metadata to match. Take a photograph of every enclosure front before a drive moves, and write the slot on each one. The order can be reasoned out from parity afterwards, and there is no sense paying for that when a phone camera settles it.

A backup appliance that failed with the backups inside it

Purpose-built backup targets are arrays with a deduplication layer over them, so a failure costs an organisation its storage and its safety net in the same event. Those are rebuilt here, and the reference tables usually are the job rather than a detail of it.

A pool whose first parity build never finished

Put into service before the initial synchronisation completed, or interrupted during it, an array can carry parity that does not match its own data. Everything behaves until a member fails, at which point reconstruction produces convincing rubbish. This is identified during the assessment rather than found out halfway through.

A SAS expander that removed twenty-four disks at a stroke

One expander failing takes an entire enclosure out of view, which on the management screen is indistinguishable from total loss. The drives are typically fine. This is among the more encouraging findings on the page, and it is a common one on older shelves.

No diagram, no records and a rack full of disks

The person who built it left, the documentation left with them, and what remains is hardware nobody can describe. That is a routine arrival rather than a hopeless one. The arrangement comes out of the disks. Records make it quicker; their absence does not make it impossible.

One volume presented to two servers that both wrote to it

Without a clustered file system, two hosts writing the same blocks damage each other's work quickly and thoroughly. Some of it is repairable from the structures that survive. How much depends entirely on how long both machines were live and how busy they were.

An IT provider who has already had a go

Second opinions on enterprise storage are ordinary here and a good share of them come good. Set out what was tried, in what order, and how each attempt ended, because that is what decides where this bench begins. Looking costs nothing on a fourth attempt just as it does on a first.

A loading bay that has stopped moving

Not a fault, a priority. Warehousing and distribution across the parks off the M6 and M69 run on systems nobody dares restart, and when the storage under one goes, the bay stops with the paperwork. Mention it when you ring and the drives go to the front. Two working days for the free examination applies regardless.

Underneath it is an array; above it is where the trouble lives

Strip a SAN back and you find disks in RAID groups, reconstructed the way any set is: every member imaged read-only, the parameters worked out from what is on the images, nothing written back to your hardware. The complication sits above that. Enterprise storage inserts a virtualisation layer between the physical groups and the volumes it presents to hosts, and that layer maintains its own map of which physical extents belong to which LUN. Thin provisioning lives there. So do automatic tiering between flash and spinning media, deduplication tables and snapshot chains. When the map is damaged, a shelf of entirely healthy disks presents nothing anybody can use, which is why a large share of these jobs turn out to be metadata problems wearing the costume of a hardware failure, and rebuilding that map is the bulk of SAN work. Higher still there is usually a file system, most often VMFS carrying virtual machines, and a datastore ESXi will not mount is generally sound underneath and releases its virtual disks readily once the LUN below it has been put back together.

What goes in the crate

Member disks only. Label each one with its shelf and its slot, and photograph the front of every enclosure before the first drive comes out, because five seconds with a phone camera settles what would otherwise have to be deduced from parity. Controllers, chassis, caddies where they are awkward to remove, and rails all stay in your rack, since the layout is recoverable from the members and no replacement hardware is wanted at this end. One thing does have to travel that people forget: any cache device or flash tier that formed part of the arrangement, labelled as such, because a volume spread across two tiers cannot be read from one of them. Pricing follows the array band, from £500 + VAT and rising with the member count, and the free examination closes two working days after the disks are booked in. Logical work carries no fix, no fee; members that failed mechanically or electronically do not, and half is payable up front on that part. Two instructions before you start packing: if a rebuild is running, stop it, and if a management console is offering to create a new pool, decline.

// the tools it takes

What stands on the bench, and why any of it matters

Shared storage is reconstructed from read-only copies of its members, exactly as any other array is. What makes it slower is the number of storeys standing above the disks, and most of the equipment below exists to take those storeys apart one at a time without disturbing what is underneath them.

Imaging a whole shelf at once

Twelve, twenty-four or forty-eight drives copied in parallel with retry limits enforced by the imager rather than by software. Nothing about reading a shelf one drive after another is difficult. It is simply slow, and running the channels together is what keeps a full rack inside a sensible number of days.

Fibre channel, SAS, SATA and NVMe, all write-blocked

Every interface an enterprise shelf has used in the past twenty years, including the older parallel SCSI and fibre channel drives that plenty of racks still contain. The shelves that fail are seldom the newest hardware in the building, so reading the old interfaces matters more than reading the new ones.

The mapping layer between disks and volumes, rebuilt

Extents, pool descriptors, thin-provisioning tables and the records that say which physical ground belongs to which LUN. Take that storey away and what is left is an ordinary array; damage it and a rack of perfectly good drives offers a host nothing it can mount.

VMFS and the other file systems that sit on a volume

Once the LUN is back the file system on it is separate work. VMFS in particular keeps its own structures and its own resource files, and a datastore that a hypervisor rejects is usually sound underneath and gives up its virtual disks without much argument.

Virtual disks lifted out and repaired one at a time

VMDK, VHDX and QCOW2 files extracted from a recovered datastore, then mended individually where one of them spans a region that could not be read. A guest that refuses to boot after the storage returns is a distinct problem and is treated as one.

Snapshot chains and deduplication indexes put back

Reference tables reconstructed so blocks written once and pointed at from a thousand places can be found again, and chains re-linked so an earlier consistent view becomes reachable. On a deduplicating appliance these two are not details, they are the recovery.

// badges that turn up in the post

Storage platforms taken in

HPE 3PAR, Nimble or MSANetApp, FAS or E-SeriesDell PowerVault or CompellentDell EMC VNX or UnityIBM DS series or StorwizePromise or InfortrendEternus, or another Fujitsu shelfA JBOD chain, Supermicro or genericAn iSCSI or fibre channel targetHyperconverged or software-defined nodes

What comes out of a rack and into a parcel

Shared storage is priced as an array: from £500 + VAT, rising with the member count, with the scope settled after the free examination because enterprise sets differ from one another far more than desktop ones do. Looking is free and the answer arrives on the second working day after the drives are logged in. Logical work carries no fix, no fee. A member that failed electronically or mechanically sits outside that, and the physical part takes 50% up front. Enterprise jobs are ordinary weekly work on this bench rather than something fitted in around the rest, and they arrive from manufacturers and their suppliers around Coventry, from logistics operators on the parks off the M6 and M69, from the two universities and from IT providers looking after a dozen Warwickshire clients on one platform. Two things, said plainly. A rebuild that is running should be stopped. An offer from a management tool to build a fresh pool, or to initialise a volume, should be declined. Both positions stay recoverable right up until somebody says yes.

// getting it ready for the post

Before the box is taped shut — take the drive out if it comes out

Drives only, each marked with the enclosure and the slot it came out of. Photograph every shelf front before the first drive comes out, because slot order is the one thing here that a phone camera can record perfectly and a bench has to reconstruct. Leave the head units, the chassis, the caddies that fight you and the rails where they are: none of them are needed, since the arrangement is derived from the members themselves. Wrap each drive so it cannot knock against its neighbours and use a carton stiff enough to hold its shape under the weight. Send it tracked and insured to Oxford Data Recovery, John Eccles House, Oxford Science Park, Robert Robinson Avenue, Littlemore, Oxford, OX4 4GP. Coventry to that door is about fifty-five miles down the M40 and takes roughly an hour, and drop-offs are taken in person Monday to Friday, 9:00am to 5:30pm; a courier you book yourself is equally welcome. There is no counter in Coventry and nothing is collected. Where the parcel is unusual — several shelves, a tiering pair that has to stay together, a chain of expansion units — ring 0800 689 0668 before you pack it, and say on the same call if the business has stopped.

// getting your media to Oxford

Posting a device in — what goes in the box

Almost everything worked on here arrived in the post. A drive that is already in trouble has an easier time boxed, padded and insured than it does being carried round in a bag for an afternoon, and a parcel handed over in Coventry today is normally booked in at Oxford tomorrow morning.

As a rule the storage comes out and the machine stays where it is. That applies to a laptop, a tower, an iMac and to the recorder sitting under a counter. Taking equipment apart is not something this bench does, and a repair shop will free a drive in a few minutes. Two things go the other way: an external drive stays sealed inside its own case, and a NAS travels as a complete unit with its disks still in their bays. A Fusion Mac is a third case — both of its drives come out and travel together, each one labelled. The single situation nobody can work around is memory soldered flat onto a mainboard, which is how Apple Silicon Macs and a good many slim laptops are built: if the storage will not unbolt, there is no parcel to send.

  • Use a box or padded mailer with some rigidity to it, and pack around the drive until nothing shifts when the parcel is tilted. Mains adaptors, docks and leads are not wanted at this end.
  • Sending a RAID or a server? Only the member disks travel — not the chassis, not the controller — and each one wants its bay number written on it. Take a photograph of the front of the unit before anything is pulled; it costs nothing and now and again it saves a day.
  • Print the shipping and booking-in form (PDF), add a name, a number you will answer and a sentence on how the trouble began, and drop it in beside the media.
  • Most people use Special Delivery, which is tracked and covered; a courier of your own does the same job. You can also bring it: the Oxford reception takes devices over the counter, Mon–Fri 9:00am–5:30pm. Neither a Coventry counter nor a collection round exists.
// write this on the label

Oxford Data Recovery

John Eccles House
Oxford Science Park
Robert Robinson Avenue
Littlemore, Oxford, OX4 4GP

↓ Print the shipping & booking-in form (PDF)

Put Oxford Data Recovery on the label. From Coventry it is roughly fifty-five miles straight down the M40, about an hour if you would rather drive it in than post it. Either way you are told the moment it is logged, and the free diagnostic finishes two working days later.

Unsure what ought to go in the box? Ring 0800 689 0668 before you seal it, or work through the free online diagnostic and let it do the asking.

// SAN recovery questions

Common questions

The array band applies: from £500 + VAT, climbing with the number of members, exactly as for any other set. The examination is free, finishes two working days after the disks arrive, and produces one written figure before work begins. Logical work carries no fix, no fee. Members that failed mechanically or electronically sit outside that and take half up front, because donor parts are ordered for those particular disks.
That is generally recoverable. A datastore the hypervisor refuses is usually intact underneath, and the virtual disks are pulled straight out of an image of the LUN rather than through ESXi. If a virtual disk turns out to span a damaged region, repairing it is a separate step afterwards. There is a page here dealing with VMware and Hyper-V in more detail.
Look at the backplane, the expander or the power path before blaming the drives. A single expander failure can take twelve or twenty-four disks out of view simultaneously, which reads on the console like a total loss and very seldom is. In most of these the disks themselves are undamaged and the pool is reassembled from images of them.
It makes no difference. Nothing in the reconstruction depends on an active support contract or on the manufacturer still trading, because the work is done from the disks. Fibre channel and SAS shelves of some age turn up here regularly and the interfaces are already on the shelf. Age is far less often the obstacle than it looks from outside.
// the rest of the week's work

What else reaches this bench

// where to read next

Pages that go further than this one

The bench is ready whenever you are.

Looking at it costs nothing, one written figure follows, and the band governing this page is from £500 + VAT on a SAN, the figure following the number of disks in the pool.