A gloss-black storage array bezel, immaculate and still carrying protective film on its badge, surrounded by dusty units in the same rack.

Specimen VIII

The Alibi

Threat rating
Endemic
Domain
Diagnosis

First contact

A Lync 2010 deployment coming apart. Spinning wheels on messages that failed a minute later, robotic voices on calls, calls that never connected at all. Weeks of support, no progress. Reading the case on the plane the shape of it was obvious enough — every symptom was a performance symptom — and support had already said so, more than once. “It can’t be the SAN,” the customer had told them. They had bought a very expensive one.

Behaviour

It attaches to whatever the organisation paid the most for. The price is the entire defence: a component that cost more than the project it supports acquires the standing of a witness rather than a suspect, and investigation routes politely around it. Weeks of effort go into every other layer first, in good faith, by competent people, because the expensive thing has an alibi and nothing else in the estate does.

In its most common form the shelter is a shared pool. Everything is placed on the flagship because the flagship is the best thing available, so latency-sensitive workloads end up sharing spindles with whatever else is hungry. The array does precisely what it was bought to do, for too many things at once, and the invoice is what stops anyone checking.

Signs of infestation

Before investigating anything, write down which components have already been excluded from suspicion, and the reason beside each. Any exclusion resting on cost, vendor tier, or how recently it was bought is an assumption wearing the word, and it is where to start.

Then measure the excluded thing first, with the cheapest instrument available and no theory attached. Basic counters settle in an afternoon what architecture debate does not settle in a month. In a modern estate the shelter is the premium tier rather than the array: check whether latency-sensitive workloads share a storage account, an App Service Plan, an elastic pool or a node pool with something noisy, and read the throttling and queue depth rather than the SKU name.

Containment

Moving the workload off the shared pool fixes the outage, and fixes nothing else. The condition is the exclusion, and it survives the fix: once the symptoms stop, the expensive thing is available again and the reasoning that excluded it has not changed. What holds is a written record — the counters, the vendor’s own supportability documentation, the date — filed where somebody who was not in the room can find it. The report outlives the decision, and that is the only part of this anyone controls.

Recovered note

They ran it on a smaller, cheaper SAN as a test and the problems disappeared overnight. A few weeks later they moved it back, and everything came back with it. At least the account manager had my report to point at.