8 min

Do you need an on-site spare parts stock for a 24/7 SLA?

Learn when an on-site spare parts stock for a 24/7 SLA beats a supplier reserve and how to verify availability, delivery, and ownership.

Do you need an on-site spare parts stock for a 24/7 SLA?

A 24/7 SLA does not become dependable because a contract says "a reserve is available." It becomes dependable when a specific compatible part is assigned to a specific environment, can arrive within the allowed downtime, and no one can pass responsibility for an error among the customer, supplier, and carrier.

An on-site stock gives you physical control, but it ties up cash and quickly turns into a museum of incompatible parts without disciplined records. A supplier reserve saves capital and transfers some of the upkeep, but only when it is allocated to you, separated from saleable inventory, and confirmed regularly. For most round-the-clock systems, the most dependable choice is not either option in pure form, but a tiered arrangement: the most consequential items near the equipment, expensive and rarely used assemblies at the supplier, and commodity items in the general logistics network.

Reliability is measured by more than the storage location

The storage location alone tells you nothing about recovery. I have seen on-site stores where the required power supply sat behind a locked door while the night engineer had no access. I have also seen supplier warehouses deliver an allocated part faster than the customer's staff could approve an internal issue. You need to compare the entire path from the monitoring alert to the service returning to operation.

That path has four separate intervals: diagnosis, confirmation of the required item, physical delivery, and replacement followed by testing. A promise of "delivery within four hours" covers only one interval. If the contract does not define when the clock starts, the supplier may start it after its own diagnosis while the customer counts from the failure. Both may be formally correct while the system remains down.

Procurement teams often blur three concepts that need to remain separate:

  • availability means the part physically exists at the stated location;
  • reservation means nobody may sell or issue it to another customer;
  • readiness to install means compatibility, completeness, firmware, and access to the operating site have been checked.

Getting this distinction wrong is expensive. Ordinary saleable stock can disappear a minute before the incident. A reserved controller can have the wrong revision. A fully ready item has a serial number, a link to the supported configuration, inspected packaging, and a named recipient at night or over a weekend.

Count tied-up capital together with downtime cost

An on-site spare is justified when the expected damage from a delay exceeds the stock's total cost of ownership. Comparing purchase prices alone almost always produces the wrong choice. Cash sits on a shelf, space and records cost money, parts age, batteries lose capacity, and uncommon boards sometimes get written off without ever being installed.

For each item, I use a simple annual calculation:

Стоимость своего ЗИП = цена капитала + хранение + проверки + страхование + ожидаемое списание
Стоимость резерва поставщика = плата за резерв + доставки + проверки договора + остаточный риск задержки
Ожидаемый риск простоя = вероятность отказа × дополнительное время восстановления × ущерб за час

This is not an attempt to forecast the exact amount. The formula forces system owners to state their assumptions. If nobody knows the loss per hour, the reliability discussion has no economic basis yet. For a hospital, you should separately consider access to clinical systems. For a bank, consider the time to restore operations and obligations to customers. For a factory, include a stopped line and a safe restart.

Do not multiply the price of a part by the number of identical servers. Failures are not always independent, but keeping one assembly for every machine usually makes no sense either. Start with the size of the installed fleet, actual replacement history, the permitted time without redundant capacity, and whether you can temporarily move load away from the failed node. One replacement power supply can cover ten similar systems if losing one supply does not shut a server down. That approach is dangerous for a single storage controller with no redundancy.

The popular advice to "buy one of every part number" appeals to procurement because it produces a ready-made bill of materials. It is poor advice because it treats a cheap cable and an expensive system board alike, ignores common causes of failure, and consumes the budget on items that could be delivered without putting the SLA at risk. Build the stock around consequences and replacement time, not the equipment catalog.

To keep ranking from depending on who argues loudest in a meeting, assign every item a class. Class A means a failure stops a critical service or immediately removes its permitted protection, and outside delivery cannot arrive in time. Class B means the service continues on redundancy, but the part must arrive before the elevated-risk window closes. Class C allows ordinary delivery. Beside the part number, record the symptom that tells the on-duty engineer this is the required assembly. Otherwise, the warehouse may open quickly at night and still issue the wrong part.

Check correlated failures as well. Drives from the same batch, a shared firmware defect, or a power event can disable several units close together. A history of individual replacements will not reveal that exposure. These scenarios need either more stock or another recovery method, such as moving load to a site with a different hardware base. Do not turn a remote theoretical risk into an unlimited warehouse: document the scenario, estimate likelihood without invented precision, and state the loss the organization accepts.

It also helps to separate the cost of readiness from the cost of actual use. The readiness charge covers allocation, storage, and testing even if nobody needs the part. Use covers delivery, engineering work, and replacement of the issued item. This split lets you compare on-site and supplier stock honestly instead of calling an unused reserve fee "wasted money." You paid for availability on a bad day, just as you pay for any other standby capacity.

Delivery time starts at the incident, not the order

An SLA needs a recovery deadline, not an attractive courier time. The clock should start at the recorded event or at a call from an authorized on-duty employee. Within the deadline, the supplier diagnoses the fault, confirms a compatible part number, authorizes its issue, clears access control, delivers the part, and hands it to the engineer. If an activity is excluded, the contract must name it and set a separate limit.

Break the route into actual minutes and owners. Suppose an incident happens on Saturday at 02:10. Monitoring creates an incident at 02:12, the on-duty engineer confirms a hardware symptom at 02:25, and the supplier accepts the case at 02:31. The storekeeper does not arrive until 04:00, then the team discovers that the courier's vehicle is not cleared to enter the site. The part reaches the rack at 06:20. The formal "two-hour delivery" may have started at 04:10 after the issue was processed even though the service had already breached its target.

This review exposes hidden queues:

  1. Who can declare a hardware incident without a manager's approval?
  2. Who can open the warehouse and sign an issue around the clock?
  3. Which documents are required for removal and site entry?
  4. Is a remote site inside the promised delivery area?
  5. Who carries the risk if a road closes or a flight is cancelled?

Geography matters especially for distributed organizations. "Nationwide" is not a substitute for a table of deadlines by site. Each site needs a normal route, a backup route, a maximum time, and a rule for unavailable transport. A modest local kit at a remote site is sometimes more dependable than a large central warehouse hundreds of kilometers away.

Someone must own parts obsolescence

An obsolete part remains an asset only in the accounts. In practice, it may not support the installed firmware, may require a discontinued adapter, may lose capacity in storage, or may fail its incoming test after years on a shelf. The contract must define in advance who tracks compatibility and who pays to replace the reserve when the main equipment changes.

With an on-site warehouse, responsibility naturally stays with the infrastructure owner. That does not mean it can remain unnamed. Assign an owner for the parts list, not only an employee responsible for physical inventory. The first person owns applicability and lifecycle, while the second usually owns physical custody. If you merge those roles, a stocktake can report a perfect quantity of useless parts.

A supplier reserve needs rotation rules. The supplier should replace an item when the manufacturer ends support, when a new revision of the installed equipment is no longer compatible, or when a control test finds a defect. The customer, in turn, must report configuration changes. You cannot demand compatibility with a system the supplier has never seen since its upgrade.

I would not accept a clause allowing an "equivalent or improved replacement" without an approval process. A board that looks improved in a catalog may require a different driver, change failover behavior, or break a certified configuration. Confirm equivalence on a test system or in a written compatibility matrix, not through a manager's opinion during an incident.

The inspection interval depends on the part type. You can test drives and fans regularly without an elaborate lab. System boards and controllers need a compatible chassis, a matching firmware version, and sometimes a license. Battery modules need their own storage conditions and residual-capacity measurements. Giving every item the same annual visual inspection produces a report, not confidence.

Proof of availability must leave an auditable record

Servers and support together
GSE supplies S200 Series servers and provides round-the-clock technical support.
Select a solution

A manager's statement in an email does not prove an allocated reserve exists. You need a register showing which physical unit has been assigned, where it is stored, what it supports, when it was checked, and when the reservation expires. Not every minor item needs a serial number, but without one it is difficult to distinguish expensive, critical assemblies from saleable stock.

A minimum record can look like this:

{
  "reserve_id": "R-0241",
  "customer_asset_group": "DB-CLUSTER-A",
  "part_number": "PN-EXAMPLE-01",
  "serial_number": "SN-008731",
  "quantity": 1,
  "location_code": "WH-AST-01",
  "status": "reserved",
  "compatibility_checked_at": "2026-06-15",
  "next_check_at": "2026-09-15"
}

The part number and quantity answer "what." The serial number and warehouse code answer "does this exact unit exist." Status and dates answer "can we trust it now." A working register should also contain supported models and revisions, the last test result, storage conditions, the person responsible for the inspection, and a movement history.

It is useful to confirm the reserve at three levels. A monthly export shows the reserve contents and exceptions. A quarterly sample check compares records with labels and a photograph of the packaging. A periodic test issue verifies that somebody can actually collect the part at night, clear it through security, and deliver it to the site. Set the frequency by criticality and rate of change, not by a uniform calendar.

The customer needs access to evidence, but does not necessarily need direct access to the supplier's warehouse system. A signed report, an export with an immutable record identifier, and a right to conduct sample inspections usually provide enough control. If the supplier will not show at least the part number, quantity, location, and inspection date, it is offering a promise rather than a reserve.

A shared pool can work, but it needs a different calculation. If one physical part covers several customers, the supplier should disclose the priority rule and the simultaneous demand the pool can withstand. "Available to all customers" means that during a widespread defect, the item goes to whoever filed first or had the most persistent manager. For systems with a strict SLA, treat that stock as an extra source rather than a guaranteed reserve.

A photograph of a box helps only when it includes an identifier and date. A picture without a serial number can support several reports, and an old photograph says nothing about the part's current location. During a sample check, ask for the customer-selected unit to be photographed beside a one-time verification code, then compare its label with the register. Do not do this for every fastener. Use the method where cost or criticality justifies proof of physical presence.

Compatibility also needs a snapshot of conditions rather than a permanent "checked" mark. Retain the main equipment model, hardware revision, firmware version, and test date. When any of those parameters changes, the former result becomes conditional until a new review. If a full test requires downtime or an expensive lab, define the acceptable indirect evidence and the risk that remains with the customer in advance.

Do not let a report hide exceptions. Its opening rows should show items with expired inspections, temporary substitutions, breached storage conditions, or incomplete kits. A single exceptions row matters more than a hundred green rows. Every exception needs an owner, correction deadline, and temporary protection, or monthly confirmation becomes a ritual. The SLA owner, not the employee who merely counts boxes, should decide whether to accept the report.

A hybrid reserve usually survives incidents better

A hybrid arrangement places parts according to how quickly their absence becomes unacceptable. Tier-one items stay on site: power supplies, fans, drives, cables, and other assemblies that fail more often and restore service quickly. The exact contents depend on the architecture, so do not copy this list without checking your system.

Tier two sits in a regional warehouse or with a service partner. It holds expensive boards, controllers, and complete nodes that are rarely needed but must arrive within hours. Tier three is the supplier's shared stock for parts that can take one or several days to arrive. It works only where a cluster, recovery site, or capacity reserve can support the wait.

Use two axes to make the decision: the maximum time without the part and the difficulty of replacement. If the system loses its permitted redundancy immediately after the first failure, you may need the part on site even while service continues. This is a separate case: the SLA has not failed yet, but the next failure will stop the system. The contract should treat that state as a priority incident rather than a routine replenishment request.

A kit is sometimes safer than an individual part. After a system-board failure, an engineer may discover a damaged connector, missing fastener, or incompatible cable. If the replacement procedure calls for thermal material, fasteners, cables, and boot media, keep them in one sealed kit with a contents sheet. A complete kit lowers the risk of a second trip, which SLA calculations rarely include.

A hybrid scheme needs a replenishment rule. After an on-site spare is issued, the supplier should receive a signal immediately, and the stock restoration deadline must differ from the service recovery deadline. Otherwise, the first incident succeeds and a second incident a week later finds an empty bin. For critical parts, define a temporary replacement until the permanent item arrives.

The contract must describe a managed reserve

Spares matched to the configuration
GSE agrees reserve contents together with the server configuration and support terms.
Discuss the project

A good SLA appendix reads like an operating procedure. It does not leave the sentence "the supplier ensures the availability of required components" unexplained. It includes a parts list or a rule for building one, storage points, access arrangements, deadlines by site, rules for starting and stopping the clock, a compatibility-confirmation process, and consequences for non-performance.

I check at least these terms:

  • the reserve is allocated to the customer and cannot be used for other cases without written permission;
  • a part-number substitution requires approval and a compatibility check;
  • the supplier reports shortages, movements, and failed tests before an incident;
  • the customer can inspect the register and sample physical items;
  • after an issue, a separate replenishment deadline and a temporary protection measure apply.

Responsibility should follow control of the risk. The customer owns configuration accuracy, access to the site, and timely notice of changes. The supplier owns custody, allocation, testing, and delivery of its reserve. A carrier may operate the route, but the supplier should not use the carrier to avoid the obligation if the supplier selected the logistics arrangement.

A service credit does not restore service. It is useful as a price signal for a breach, but actual protection comes from a backup route, the right to collect the part with the customer's transport, a temporary local item, and a clear escalation path. If the only response to an undelivered board is a discount on the next invoice, the customer still carries all operational risk.

ISO 22301 requires organizations to identify the resources needed for continuity and test their arrangements through exercises. The standard does not say every resource must be stored on site. That is a sensible position: an organization may outsource storage, but it cannot outsource its duty to prove that the recovery arrangement works.

State ownership and the point at which risk transfers separately. A part may belong to the customer while physically remaining at the supplier, or it may remain the supplier's property until issue. The first model needs rules for stocktakes, insurance, and return when the contract ends. In the second, the supplier must not quietly replace an allocated reserve with a promise to find an equivalent on the market. For incident response, the ban on disposal and clear liability for damage matter more than the accounting label on the asset.

The contract also needs an exit process. If the customer changes suppliers, reserved parts cannot remain in limbo. State which items the customer buys, which return to shared stock, who pays for transport, and how long protection continues during transition. Otherwise, the organization may lose its reserve on the day the former support agreement ends, before the new supplier has accepted the fleet.

Finally, separate a failure of availability from a failure of delivery. If a monthly check finds a missing part, the service already has a defect even though no incident has occurred. The supplier should restore the reserve quickly and report how it covers the risk in the meantime. Waiting for a real failure before recognizing a breach rewards an attractive register with no physical units behind it.

A test incident beats a perfect warehouse report

Round-the-clock support
GSE technical support operates 24/7 through a nationwide service network.
Select a solution

You cannot accept a reserve by certificate and forget it until a failure. Run a test issue without warning the supplier's operating shift, but within an agreed window so you do not create a false emergency. Select one critical item, use the ordinary support channel, and measure every handoff between people and systems.

The record should capture the event time, case acceptance, issue decision, warehouse opening, courier handoff, site arrival, and engineer access. Record discrepancies separately: a wrong telephone number, missing pass, damaged seal, different serial number, or incomplete kit. A result such as "completed in three hours" hides causes that may combine differently in a real incident.

You do not have to replace a healthy part in a production system after delivery. You can inspect labels, completeness, and compatibility on a test system, then return the item under a new seal. But for each complicated class of assembly, it is worth completing a replacement on a test configuration at least once. That is how you find forgotten adapters, incompatible firmware, and instructions available only to an employee who is on leave.

Track several measures rather than one mean time. The share of confirmed items shows register quality. The age of the latest check shows how current the evidence is. Time from the case to the issue decision exposes bureaucracy. The share of test deliveries completed on time shows whether the promise can be performed. An average without the worst result is useless for a 24/7 SLA because the agreement gets tested on the bad day.

Every fleet change should trigger a reserve review. A new server revision, firmware update, system relocation, or access-control change can break the existing arrangement without any hardware failure. Connect configuration management to the spare-parts register so a change cannot close without an answer on compatibility and logistics.

The register needs controls of its own. Restrict manual changes to the "reserved" status, retain the author and time of every edit, and replace deletion with record closure and a reason. Reconcile register quantities with warehouse movements and service-system cases. If the same serial-numbered unit appears in two customer reports, the check should raise an exception before the next monthly export.

A simple evidence-age signal helps. Green means the physical reconciliation and test remain current, yellow warns that expiry is close, and red forbids counting the item as SLA-ready. A color does not replace the date and result, but it keeps an on-duty engineer from reading the full record history during an incident. Define a temporary route for every red item in advance: another unit, a compatible assembly, or a load transfer.

After every real failure, compare plan with fact. Which part was needed, was the diagnosis correct, where did time disappear, did the team need a second trip, and when was the reserve replenished? Do not turn the review into a search for blame. Its purpose is to change the stock level, route, kit, or instruction before the next event. If three issues in a row end without installation, that is not a warehouse success. It is a reason to inspect remote diagnosis quality.

Assign a correction deadline to every gap and test it in the next exercise. Otherwise, reports will repeat the same locked night entrance or obsolete telephone number for years. Reserve maturity shows in whether known obstacles disappear from the next measurement, not in the thickness of the procedure.

Choose by failure scenario, not procurement model

An on-site stock is more dependable when the part is needed sooner than the supplier can physically deliver it, the site is remote, access is difficult, or the organization will not entrust evidence of availability to another party. A supplier reserve makes more sense for expensive, rarely needed assemblies that can cover a large uniform fleet, provided that the supplier holds allocated units, confirms their condition, and accepts responsibility for logistics.

Before deciding, take the ten most damaging hardware failures and fill one row for each: affected service, permitted time, required assembly, diagnostic method, reserve location, route, issue authority, replacement time, temporary measure, and obsolescence owner. If a row contains "decide during the incident," the SLA still rests on hope.

For infrastructure projects, GSE can agree the division between local and supplier-held stock together with the equipment configuration, delivery, and round-the-clock support across Kazakhstan. Even then, the customer needs an auditable register and a test issue: the supplier's name does not prove readiness.

Do not choose between two warehouses as a matter of principle. Put every critical part number on the timeline of a specific failure. The part belongs at the point from which it can reliably pass diagnosis, issue, transport, and installation before the permitted time expires. Everything else can stay where capital and upkeep cost less.

FAQ

What belongs in a server spare-parts kit?

It includes compatible assemblies, consumables, and accessories needed to recover a specific configuration. Choose the contents by failure consequences and delivery time rather than copying a general catalog.

Can a supplier's saleable stock count as a reserve?

No, not if the supplier can sell it to another customer. A reserve must be allocated, have a status and reservation term, and leave an auditable record of its storage location.

Who should pay to replace an obsolete reserved part?

The contract should state this directly. The supplier usually owns parts at its warehouse while the declared configuration remains unchanged, while the customer bears the cost if it changes equipment without notice.

How often should spare-parts stock be checked?

The interval depends on criticality, the type of part, and the pace of fleet changes. Review the register regularly, and periodically issue and test complicated or expensive assemblies on compatible equipment.

Does every spare part need a serial number?

Part number, batch, and quantity are enough for cables and minor consumables. For expensive critical assemblies, a serial number proves the supplier is showing the same physical unit and not counting shared stock several times.

When should the SLA delivery clock start?

It is best to count from the hardware incident record or a call from an authorized on-duty employee. If the clock starts after the supplier's diagnosis, set a separate maximum for diagnosis.

Which approach is safer for a remote site?

Local stock is usually safer for a part with a short permitted delay. Expensive and rare assemblies can stay regionally if a tested route fits the full recovery deadline in poor weather and at night.

How do you set the quantity of identical spare parts?

Consider fleet size, common failure causes, replacement history, built-in redundancy, and replenishment time. The rule "one part per server" is almost always either too expensive or too imprecise.

Is a service credit enough for a missed delivery?

No. A credit compensates only part of the financial loss and does not restore the system. You need a backup route, a right to collect the part, a temporary replacement, and a working escalation path.

How can you test a supplier reserve before a real incident?

Request a register with part numbers, identifiers, locations, and inspection dates, then run a sample reconciliation. A test night issue reveals whether the warehouse opens, documents clear, and delivery finishes within the promised time.