8 min

Which branch network SLA costs less?

Compare a branch network SLA by downtime cost, recovery time, geography, and spare pools to avoid overpaying for engineer dispatch.

Which branch network SLA costs less?

A four-hour engineer dispatch sounds stronger than next-business-day equipment replacement. For a branch network, that is often the wrong benchmark. The right SLA for a branch network depends on when the workstation or server returns to service and the full cost of that recovery, not on how quickly a specialist reaches the door.

I have seen contracts where the provider reliably arrived within four hours, logged the fault, and ordered a part. The branch then remained down for another two days. The SLA was formally met, but the business bought expensive activity instead of recovery. The opposite extreme is just as common: a next-day replacement promise looks sensible until a Friday failure in a remote town is not resolved until Tuesday.

Choose support by equipment class and the consequences of failure. One service level cannot fit a point-of-sale workstation, a server node, an accountant's computer, and a terminal in a training room. Calculate the cost of unavailability first, then test the logistics and spare-device pool, and only then buy the recovery time you need.

Four hours to an engineer is not four hours to recovery

A four-hour dispatch answers only when a specialist must arrive. It says nothing about when work will resume. A contract usually calls this response time or dispatch time. The business needs service recovery time: the period from registering a confirmed incident until the user can do the job again at an acceptable level of performance.

The difference is clear when a system board fails. The engineer arrives after three hours and twenty minutes, runs diagnostics, confirms the defect, and meets the dispatch target. The required board is not available in the region. It leaves the central warehouse on an evening flight and then goes to a courier. The workstation returns to service after thirty hours, while the provider's report shows green status.

Next-business-day replacement commits to an outcome only when the contract defines "replacement" precisely. The provider might replace a component, the entire device, or merely accept the faulty unit and issue a receipt. The SLA should say that the replacement device is installed, boots, connects to the agreed network, and gives the employee access to the required workflow. Otherwise, a new box at the security desk solves the provider's task but not the branch's.

Separate four metrics:

  • time to acknowledge the ticket;
  • time to begin remote diagnostics;
  • engineer arrival time;
  • time to restore operation.

Combining them into one measure helps only the person preparing an attractive monthly report. Use the last metric to compare options and retain the others as control points. A four-hour dispatch has value when the engineer brings the right part, can replace it on site, and also meets a separate recovery target.

Downtime cost starts with lost productivity

The hourly cost of downtime is the sum of direct loss, paid inactivity, deferred work, and extra expense, not just an employee's salary. For a branch, build the calculation around the business process: how many transactions stop, how many people depend on the failed unit, whether work can move elsewhere, and how much must be caught up later.

A simple model works well:

Incident cost =
  hours unavailable × (direct loss per hour
  + cost of idle employees per hour
  + extra expenses per hour)
  + one-time recovery expenses

Direct loss is clear for a sales outlet or paid service: use the margin that was not earned, not total revenue. For an internal department, count the paid time of people who genuinely cannot continue. If employees switch to paper, another computer, or queued tasks, use a productivity-loss factor below one. Do not claim the whole payroll as damage when half of the work continues.

Extra expenses include urgent shipping, overtime, a local administrator's trip, rented equipment, and repeated data entry. Deferred work costs money too, though it is easy to hide. If the department handles the backlog in the evening, downtime becomes overtime and more errors. If the queue simply shifts, the damage appears as a delay in the next process.

Calculate the effect on dependent departments separately. A broken computer rarely stops the office next door, while a failed authentication node or local server may block dozens of workstations. Do not count the same loss twice: lost revenue may already include idle operators. Finance should see each component in the formula and remove overlaps, or the most expensive SLA will win because of an arithmetic error.

The calculation does not need perfect accounting precision. It needs a range that finance can challenge and still accept. Calculate conservative, expected, and severe cases. If the more expensive SLA pays back only in the severe case, do not apply it to every device. If it pays back under conservative assumptions, the argument about support price is already settled.

Do not add reputational damage to every hour without a defensible link. It may be real for a customer-facing branch, but double counting quickly inflates the model. Identify events with special consequences separately: missing a regulatory deadline, stopping a medical appointment, being unable to execute a payment, or failing to close a reporting period.

Expected annual loss is the number to compare

Comparing the price of an SLA with one frightening outage is wrong. Make the decision on expected annual loss: failure probability, the difference in recovery time, and hourly cost for each equipment class.

The minimum calculation looks like this:

Annual benefit of the faster SLA =
  number of devices
  × failures per device per year
  × downtime hours avoided
  × hourly downtime cost

Net effect =
  benefit of the faster SLA
  - annual premium for the faster SLA
  - cost of internal readiness

Suppose the network has 300 identical workstations. As an example, not a forecast, internal data shows 0.08 hardware failures per device per year. A four-hour dispatch restores a device in 10 hours on average, while next-business-day replacement takes 18 hours. The difference is 8 hours. At a downtime cost of 12,000 tenge per hour, the expected annual benefit of the faster option is 300 × 0.08 × 8 × 12,000 = 2,304,000 tenge.

Compare that figure with the premium for four-hour coverage on all 300 devices. If the premium is 4 million tenge, one fast SLA for everything fails the economic test. But the result changes if 40 workstations serve customers and cost 45,000 tenge per hour while the rest have a manual workaround. The first group can keep the fast level and the others can receive next-day replacement.

Test how sensitive the answer is to the two most disputed inputs: failure frequency and the hours the faster option actually saves. Build a table with low and high estimates for both. If a small change in either number reverses the decision, do not sign a long contract based on one estimate. Request a pilot or a short first term that lets you reassign devices between levels without a penalty.

The formula often omits the "cost of internal readiness." Fast dispatch is useless if the engineer waits forty minutes for a pass, the responsible employee does not answer, and the firmware password belongs to someone on leave. The organization spends money maintaining contacts, remote access, permissions, configuration backups, and staff availability. That cost belongs to the chosen service level.

Collect at least twelve months of data: device type, location, ticket opening, diagnostic, arrival and recovery times, cause, part used, and process impact. If you lack history, begin with ranges and label the assumptions. Recalculate after a quarter using facts instead of defending the original guess.

Branch geography changes what the promise means

One SLA across locations with different transport access almost always conceals exceptions. Four hours in a major city and four hours in a district center require different infrastructure: local engineers, parts depots, transport, access procedures, and an alternative route for bad weather or a canceled flight.

Ask the provider for a coverage matrix for every address, not a colored map by region. Each branch row should state the service city, estimated travel time, ticket hours, the cutoff for same-day service, available spare parts, and weekend rules. The phrase "subject to transport availability" turns a commitment into a preference unless the contract explains who determines availability and how.

Compare the coverage matrix with the actual asset register. Contracts routinely retain old addresses after a branch moves, and the new premises remain unsupported until an appendix is signed. Set a deadline for the provider to accept an address change and agree on an interim level for the move. For mobile sites and temporary offices, define a service territory instead of one building.

Read the definition of a business day with particular care. A ticket accepted after the cutoff on Friday may be treated as a Monday ticket, with replacement on Tuesday. For a branch that works Saturdays, that is already four calendar days. Compare offers on the same scale of calendar hours for typical failure times: Tuesday morning, Friday evening, and just before a holiday period.

The contract also needs a time zone. A central service desk may register an incident in Astana time while the branch opens and closes on a local schedule. State the time zone that starts the clock, each site's hours, and the channel that creates the official timestamp. A message to a personal account manager should not sit outside the ticket system.

Ask for facts that support the coverage model: where engineers and common parts are located, who replaces a sick specialist, and how equipment travels when the airport is closed. That does not demand commercial secrets. It tests whether the provider can deliver the promise you are buying.

A spare pool often costs less than urgent dispatch

A fleet without random configurations
The L200, M200, and S200 lines provide a basis for a controlled branch fleet.
Contact GSE

A local stock of compatible devices reduces downtime more reliably than a distant engineer's trip when the branch can perform a simple swap from instructions. This option spends money on spare equipment, storage, testing, updates, and return logistics instead of permanent field-staff readiness.

A spare pool is not a pile of old computers in a cupboard. A device needs a compatible configuration, a working drive or approved image, the required cables, an asset number, and a clear issue procedure. Someone must power it on and update it periodically. A spare that needs two hours of updates and cannot join the domain after a year in storage does not provide fast recovery.

Pool size depends on concurrent failures and replenishment time. For ordinary workstations, a reorder-point calculation is useful:

Reorder point =
  average spare use during replenishment lead time
  + safety stock for a failure spike

Do not replace safety-stock analysis with an arbitrary percentage. Use history by model and region, including seasonal delays and a common batch defect. When all devices came from one purchase lot, failures may correlate and the average understates the risk. Keeping a complete spare server at every branch is expensive, but a set of critical replaceable parts or one compatible unit per region sometimes performs better.

Next-business-day replacement works when the process can tolerate an overnight interruption, delivery is predictable, and the configuration can be restored quickly. It performs worse than four-hour dispatch when the device is unique, data cannot move safely, the branch cannot carry out a switchover, or each stopped hour is expensive. A box alone does not settle system-image, licensing, encryption, and network-identity issues.

Design replacement security before the incident. A failed drive may contain data, and the spare device may lack current policies and protective controls. Record who removes and seals the drive, where it is stored, who approves destruction or warranty transfer, and how the temporary device is confirmed clean after return. These actions must not hold up recovery, so prepare the clean image, accounts, and encryption before anything fails.

Assign ownership of the spare. The provider can hold a pool at a regional depot, the customer can keep it at the branch, or both sides can use a mixed arrangement. The contract should state who tests readiness, who pays to replenish the pool after an issue, and how many hours it takes to return the spare capacity to availability.

Servers and workstations need different levels

Do not automatically assign a fast SLA to a server and a slow one to a workstation. Business impact and the available workaround set the priority. A clustered server with tested failover may wait a day for a part, while the only dispatch clerk's computer can stop all shipping at a branch.

Classify from the business function back to the technical object. First set the maximum interruption the process can tolerate. Then identify the devices, networks, power, applications, and people it depends on. Assign a recovery time to every link only after that. This order reveals awkward gaps: an expensive server SLA does not help when the only network switch has a five-day service term.

A practical matrix has four classes. Class A stops a regulated or revenue process without a workaround. Class B materially reduces capacity, but a temporary workaround exists. Class C affects one employee who can move to another desk. Class D covers training, spare, and noncritical devices. The names do not matter, but the criteria must be observable.

Class A usually needs a recovery target in calendar hours, remote diagnostics immediately after registration, and prepositioned parts or ready failover. Class B may work with four-hour dispatch during business hours. Classes C and D often need only next-business-day replacement or centralized repair. Assign an exception to a specific asset with a reason, not to the owner's job title.

Do not confuse service availability with device health. If an employee continues on a spare computer, service has recovered even though the failed device is still under repair. That is acceptable to the business. If a backup server starts but runs at half speed while the queue grows, formal availability does not mean the agreed quality has returned.

The contract must measure results without loopholes

One owner for the fleet
In-house manufacturing and system integration connect supply, configuration, and ongoing support.
Scope the project

A good SLA defines hours, the events that start and stop the clock, priority levels, the recovery criterion, exclusions, and evidence of performance. A penalty cannot repair a bad definition: a provider can honestly meet a meaningless target.

ITIL 4 service level management guidance calls for clear targets tied to consumer utility, warranty, and experience. I agree with linking targets to the business, but those words are not enough for a contract. Convert each target into an observable event: a user completed a test transaction, monitoring received a normal signal, or the branch owner accepted recovery.

ISO/IEC 20000-1:2018 requires organizations to plan, design, transition, deliver, and improve services through a service management system. That leads to a useful test: you cannot evaluate four-hour dispatch separately from incident records, inventory, suppliers, changes, and performance review. A provider's certification may confirm that processes exist, but it does not prove coverage at your address or the presence of the board you need.

Before signing, check eight fields condensed into these five points:

  1. The clock starts when the ticket is automatically registered, not when the provider manually acknowledges it.
  2. Impact and urgency determine priority, with a short escalation period for disputes.
  3. Pauses are allowed only for listed reasons, and the customer can see the start and end of each pause.
  4. An agreed test confirms recovery; part delivery or engineer arrival does not count as the result.
  5. Reports show the median, 90th percentile, breaches, and excluded tickets separately by region and class.

A compliance percentage without a time distribution deceives. Nine quick replacements can hide one two-day outage of a critical node. The median shows the normal case, the 90th percentile shows the long tail, and the breach list lets you investigate causes. With a small sample, inspect each ticket instead of pretending to have statistical precision.

Service credits or penalties should offset part of the fee and create an incentive, but they do not replace continuity planning. Liability caps are often far below the actual loss. If downtime is unacceptable, redundancy and ready replacement protect the operation better than a right to receive a discount next month.

Define repeat incidents too. If a device fails again for the same reason after repair, a new ticket number must not reset the quality analysis. Set an observation period after recovery and a problem-incident rule for repeats. When a defect may affect a batch, the provider should show the root cause, the change made, and the affected serial numbers. That control reduces future downtime, while a speed measure only records how quickly the team came back.

A hybrid plan beats one service tier

Servers with local support
S200 servers are made in Kazakhstan and backed by GSE's 24/7 technical support.
Request a consultation

Most branch networks benefit from a mix: rapid recovery for critical nodes, next-business-day replacement for ordinary workstations, and an internal spare where logistics are unreliable. One maximum tier overcharges for noncritical devices. One cheap tier transfers too much risk to a few important processes.

Review segmentation every quarter and after a material network change. Devices move between classes when redundancy appears, branch hours change, or a workaround disappears. The configuration register must match the contract appendix: a serial number without an address, role, and service class is nearly useless.

Do not buy four-hour dispatch for the entire fleet merely to simplify procurement. The idea is popular for an understandable reason: one item is easier to approve and one compliance percentage is easier to report. It is expensive because you pay for readiness where replacement is cheaper and still lack a recovery target where one is needed.

Compare offers with the same total-cost model. Add the spare pool, shipping, image maintenance, internal staff time, access arrangements, parts, and expected residual downtime to the subscription fee. Assess the risk of nonperformance at remote sites separately. A cheap offer with broad exclusions can have the highest expected cost.

When selecting a manufacturer and integrator, one party responsible for configuration, supply, and support can help if the contract retains measurable times and an address-level coverage matrix. GSE.kz manufactures computers and servers in Kazakhstan and has 24/7 technical support and a nationwide service network, which should be tested at the specific addresses and equipment classes in your project.

The decision must survive a bad Friday

Test the final choice against several difficult failure moments, not the provider's presentation. Use a critical server in a remote branch, an operator's workstation on Friday evening, and a batch defect. For each case, trace the route from registration to the test transaction and record who restores service, where, and with what equipment.

If the four-hour engineer arrives without the part, that tier does not reduce the downtime you care about. If "next-day" replacement excludes weekends, recalculate the wait in calendar hours. If an internal spare cannot be issued without an absent custodian, it exists only in a spreadsheet.

Decide using three figures for each class: the annual tier premium, expected downtime hours prevented, and cost of the remaining risk. Add a limit for the maximum tolerable single interruption. Favorable averages must not justify one failure the branch cannot survive.

Put a business owner in charge of the model instead of leaving it only to procurement or IT. The business confirms downtime cost and tolerable interruption, IT verifies the technical recovery path, and procurement records measurable commitments. Each quarter, the owner compares actual incidents with the calculation and explains deviations. Without that role, the fleet changes while the expensive tier renews automatically.

The owner also decides which breaches require a process change and which ones reveal the wrong service class.

Run a pilot at several different addresses and define the criteria in advance. A test incident must pass through official channels, remote diagnostics, site access, delivery or dispatch, replacement, configuration recovery, and acceptance testing. Do not tell the operations team the exact test time, or you will measure a rehearsal.

An SLA pays when its metric matches the moment the business returns to work. Four hours to a person at the door and the next business day to a box can be compared only after both are converted into the same recovery time. When that conversion is impossible, you are buying a promise whose price cannot be calculated.

FAQ

Which costs less: four-hour engineer dispatch or next-day replacement?

It depends on the difference in actual recovery time and the hourly cost of downtime. If the engineer arrives without the part, replacement may restore work sooner; for an expensive critical process, four-hour dispatch pays only when it includes a recovery commitment.

How does response time differ from recovery time?

Response time ends when the provider accepts the ticket or starts work. Recovery time ends when the user can complete the agreed operation again, so it is the second measure that matters to the economics.

How do I calculate a branch's hourly downtime cost?

Add lost margin, the cost of employees who are genuinely idle, and extra hourly expenses. Account for workarounds, and do not count all revenue or the entire payroll as damage without a sound reason.

Does every computer need a four-hour SLA?

Usually not. Ordinary workstations where employees can move often need only a replacement, while the fast tier belongs on devices that stop a revenue or regulated process.

How do I size a spare equipment pool?

Use average spare consumption during replenishment lead time and add stock for a spike in concurrent failures. History by model and region is more reliable than an arbitrary percentage of the fleet.

What should count as equipment replacement under an SLA?

Replacement is complete when a compatible device is installed, configured, and passes the agreed work test. Delivering a box to storage or replacing a part without testing does not restore the service.

How should weekends count for next-business-day replacement?

Convert the promise into calendar hours for several registration times, including Friday evening and holidays. The contract should fix every branch's schedule, time zone, and ticket cutoff.

Which reports are needed to monitor an SLA?

Require timestamps for every control point, the median, 90th percentile, breach list, and exclusions by region and class. One overall compliance percentage hides rare but expensive delays.

Can penalties replace an SLA for downtime?

No. A penalty returns only part of the fee after the event and is often capped by the contract. If the process cannot stop, it needs redundancy, ready replacement, and a tested switchover procedure.

How can I test a provider's promise before signing?

Run a pilot at both an urban and remote address through official support channels. Test diagnostics, site access, part availability, configuration recovery, and acceptance, not just arrival time.