8 min

Choose your own data center or an integrator by outage cost

Compare your own data center or an integrator across three load scenarios, CAPEX, power, staffing, redundancy, and launch times.

Choose your own data center or an integrator by outage cost

The choice between an in-house site and an integrator's infrastructure does not come down to the price of a rack. It breaks down when a team compares a building with rented capacity, forgets the cost of delay, and treats the maximum forecast as a constant load. One project then pays for idle UPS capacity and cooling for years, while another unexpectedly hits a contractual limit during a seasonal peak.

I make this choice using three scenarios, one time horizon, and equal availability. First, I separate the IT equipment required under either option. I then calculate the facility, electricity, people, redundancy, and commissioning time. This order quickly shows where an in-house data center provides control and where an organization is buying expensive capacity it does not yet use.

Define the unit of comparison first

You must compare working computing capacity under the same division of responsibilities, not one rack with another. With an in-house data center, the customer pays for the room, utility feed, UPS units, generators, distribution, cooling, fire protection, physical security, network, monitoring, spare parts, and on-call coverage. An integrator's offer may already include some of these costs in its fee, but migration, connectivity, remote hands, backups, and extra capacity may be separate.

Put servers, storage systems, and licenses in a separate layer. If both options require the same cluster, its purchase price does not help you choose a site. A difference appears only when the integrator offers another consumption model, such as dedicated hardware, a shared pool, or resources on demand. In that case, compare guaranteed processors, memory, IOPS, bandwidth, and permitted resource overcommitment rather than server prices.

Set five boundaries before requesting commercial proposals:

  • average and peak IT load in kW, not the sum of power supply nameplate ratings;
  • required recovery time and acceptable data loss for each service;
  • redundancy mode for power, cooling, network, and the computing platform itself;
  • work included in the monthly fee;
  • the date by which capacity must accept production traffic.

The last item often changes the decision more than CAPEX. If a new system must run in four months, an in-house facility scheduled to take two years is not an alternative, even if its ten-year cost looks lower.

Three scenarios separate growth from panic

Three scenarios are enough for the decision: committed, planned, and stress. They should not be percentages that finance mechanically added to last year's figures. Tie each scenario to a business event, such as the number of branches, an analytics launch, a business-system migration, image retention, or model training.

Consider an illustrative organization with a five-year horizon. In scenario A, average IT load reaches 60 kW and peaks at 90 kW. In scenario B, the average is 180 kW and the peak is 270 kW. In scenario C, the average reaches 450 kW and the peak reaches 675 kW. These figures do not describe the market or replace measurements. They demonstrate the calculation, and the reader substitutes their own data.

You cannot design an in-house facility exactly for the average. For this example, assume installed IT capacities of 120, 360, and 900 kW respectively. A twofold margin here covers the peak, one failed component, and the next expansion step, but it does not mean that a factor of two suits every organization. If load grows in 200 kW blocks, the module should match that block. If growth is gradual, a large reserve merely freezes capital.

Scenario summary:

  • Scenario A: average IT load 60 kW, peak 90 kW, installed in-house capacity 120 kW; starting position - integrator.
  • Scenario B: average IT load 180 kW, peak 270 kW, installed in-house capacity 360 kW; starting position - test both options.
  • Scenario C: average IT load 450 kW, peak 675 kW, installed in-house capacity 900 kW; starting position - own data center or hybrid.

The final row is not the answer. At 60 kW, an in-house facility may be mandatory because of data controls or a lack of a suitable nearby site. At 450 kW, the integrator may win if the customer needs capacity within months or the load persists only for short windows. The table directs the investigation; it does not replace it.

Calculate electricity from IT load

Annual consumption comes from measured IT load and PUE, not the rating of the main circuit breaker. PUE is total site energy divided by IT equipment energy. The US Department of Energy Federal Energy Management Program guide makes a separate warning: PUE describes the efficiency of supporting infrastructure, not the efficiency of computing. An idle server can sit in a facility with a good PUE and still waste money.

For the model, assume PUE values of 1.60, 1.45, and 1.35 for the in-house facility as utilization grows. Use 1.40, 1.32, and 1.28 for the integrator's infrastructure. These are calculation assumptions, not promises about a particular facility. Ask the operator for its measured annual figure, measurement boundaries, and monthly profile. For the in-house design, obtain a calculation at partial load. Design PUE at full capacity says little about the first year of operation.

The formula is simple:

annual_kwh = average_it_kw * PUE * 8760
electricity_cost = annual_kwh * tariff_per_kwh
five_year_cost = capex + 5 * (electricity_cost + staff_cost + maintenance + service_fees)

Under these assumptions, scenario A consumes 840,960 kWh per year at the in-house site and 735,840 kWh at the integrator, a difference of 105,120 kWh. In scenario B, the results are 2,286,360 and 2,081,376 kWh, a difference of 204,984 kWh. In scenario C, the results are 5,321,700 and 5,045,760 kWh, a difference of 275,940 kWh.

To convert these figures into money, multiply each result by your full tariff. It should include energy, transmission, capacity, and any other applicable charges from the bill. If the tariff is T tenge per kWh, the integrator's annual energy advantage in the three scenarios is 105,120T, 204,984T, and 275,940T tenge. This form is more honest than a random price from somebody else's project and immediately shows the result's sensitivity to the tariff.

Check two more bills. The generator burns fuel during tests and outages, while batteries require replacement according to their condition and maintenance schedule. Ask the integrator whether electricity is included in the fee, whether the coefficient can be revised, and who pays for use above reserved capacity. A cheap rack rate means nothing if every kilowatt above the limit carries a different charge.

Capital spending loses to waiting

CAPEX for an in-house facility rises in steps rather than as a smooth cost per kilowatt. A second feed, another UPS, generator, chiller, or machine room comes in a large block. Between steps, some equipment sits idle even though the organization has already paid for design, delivery, installation, and commissioning.

Build the model from replaceable cells. For the in-house facility, use room and construction cost F, infrastructure cost per installed kilowatt K, network layer N, and commissioning P. For the integrator, use one-time migration M, connectivity L, and monthly fee S for committed capacity and services. Add IT equipment H to both sides only where it differs.

own_capex_A = F + 120*K + N + P + H_own_A
own_capex_B = F + 360*K + N + P + H_own_B
own_capex_C = F + 900*K + N + P + H_own_C

integrator_5y = M + L + H_integrator + 60*S + variable_usage
break_even_month = (own_capex - M - L) / (integrator_monthly - own_monthly)

You can paste these lines into a spreadsheet without a specialized calculator. If the denominator in the final formula is negative, the in-house site does not pay back under the chosen inputs because its monthly operation already costs more. If payback falls beyond the service life of the facility equipment or the building lease, the attractive number has no practical meaning.

The popular advice to build ten years of spare capacity at once is convenient for the designer. A large facility is easier to expand on a drawing and harder to blame for running out of power. It is usually a bad bet for the owner when demand is unpredictable. In its comparison of traditional and scalable prefabricated infrastructure, Schneider Electric attributes much of the saving to avoiding capacity built too early. Its specific percentage cannot be transferred to Kazakhstan without validation, but the reason for the saving still holds: a module that is not yet needed should not appear in the first payment stage.

Request three prices, one for each scenario, rather than a single price for the target architecture. State the price and delivery time of the next capacity step separately. Procurement will then see both the entry price and the cost of a forecasting error.

Staffing cannot hide inside support

Capacity without excess reserve
A modular server and data center configuration grows when your measured load confirms the need.
Choose a solution

An in-house data center needs people on shift, people to cover them, and specialists who do not take part in daily duty. One strong engineer does not provide a 24-hour function. Leave, illness, and a parallel incident quickly expose the difference between a name in a spreadsheet and an operating shift.

Separate facility operations, systems administration, networking, information security, and supplier management. A small site can combine roles, but it cannot remove the duties. Someone must receive a UPS alarm at night, test the generator under load, admit a contractor, update a patching diagram, and run a failure review. If an outside contractor performs that work, its cost still belongs to data center operations.

For the three scenarios, use a coverage matrix rather than an invented headcount. In each row, state hours of presence, response time, primary owner, and backup. Then calculate full employee cost, training, on-call duty, and contractors. For scenario A, test a model with remote monitoring and call-out response. Scenario B usually needs a permanent operations function. In scenario C, a general IT department no longer substitutes for a dedicated facility team.

People do not disappear under the integrator model. The customer still owns application architecture, access rights, backup, SLA control, change management, and incident acceptance. A contract with round-the-clock support does not mean the operator understands the priorities of your systems. Assign a service owner on the customer side and test the escalation path during a drill.

Compare equal service modes. If the in-house option includes 24/7 duty while the integrator's proposal provides business-hours response, the lower price buys a different risk. Conversely, do not load the external option with the largest service package when the in-house estimate assumes one employee will respond whenever possible.

Measure redundancy by failures, not the letter N

The same N+1 label does not guarantee the same resilience. Follow the power, cooling, network, and data paths from the external source to a specific application and find shared points of failure. Two UPS units are useless if both receive power through one switchboard, and two network connections help little if they enter the building along one route.

Uptime Institute defines Tier III through concurrent maintainability: every capacity component and distribution path can be taken out for planned work without stopping operations. The site still remains exposed to some equipment failures and operator errors. This distinction matters. The marketing phrase "Tier III level" without certification and without switching diagrams does not prove that maintenance can proceed without an interruption.

Run four failure tests for every scenario on paper and then during acceptance:

  1. Disconnect the main feed and trace power to the racks.
  2. Take one UPS and one cooling loop out for planned maintenance.
  3. Break each external network connection separately, then test the shared route.
  4. Lose the entire room or site and recover the service from a backup.

The fourth test separates component redundancy from disaster recovery. N+1 within one room protects against a component failure, but it does not protect against fire, flooding, prolonged site loss, or an erroneous data change. For critical systems, a second site and a tested backup often matter more than moving from one set of duplicate components to a more expensive one.

In the integrator's proposal, request single-line diagrams, responsibility boundaries, drill history, access rules, and the process for notifying customers of changes. Require the same from your own team in an in-house project. An architecture that nobody dares test by switching equipment off exists only on paper.

Launch time changes the economics

Integration instead of separate boxes
GSE connects computing, software, and data center infrastructure in one technical design.
Choose a solution

Capacity delivered after the business event has negative value regardless of its low calculated unit cost. Estimate launch time as a chain of dependencies, not as a contractor's promise to "deliver equipment in a few months."

For an in-house facility, the chain includes utility conditions, a room or site, design, review, procurement of long-lead equipment, construction and installation, connection, integrated testing, and migration. Some stages can run in parallel, but you cannot test a generator before delivery and cannot accept cooling without a load or a load bank. Add time to correct defects, not merely to sign the acceptance document.

The integrator's path is shorter only when spare capacity exists. Check available kilowatts today, reservation duration, rack or server delivery, network connections, security review, contract, and migration. The phrase "space is available" does not confirm capacity on the UPS, cooling, and generators.

Set the required date D and readiness dates O for the in-house site and I for the integrator. Delay cost equals the number of months after D multiplied by lost margin, the cost of a manual workaround, or penalties that actually apply to the project. Do not use abstract reputational damage if finance cannot explain its calculation.

If the integrator launches the workload 12 months earlier, add one year of waiting cost to the in-house option. That can completely reverse a five-year TCO. In other cases, a temporary site acts as a bridge: the organization pays the integrator until its data center opens, then moves the stable base load and keeps the ability to expand peaks quickly outside.

A hybrid contract reduces the cost of error

When growth is unpredictable, separating base and variable load works best. Stable systems with a known profile can run on in-house capacity, while new projects, seasonal peaks, and temporary computation stay with the integrator until measured history exists. This is not compromise for its own sake. It avoids buying the entire stress scenario in advance.

The contract must permit this mode. Define minimum committed volume, expansion increments, delivery time for the next block, overage price, reduction period, and exit cost. Check whether equipment and data can be removed, who handles deinstallation, the format used to hand over configurations, and how long backups remain after termination.

For organizations in Kazakhstan, GSE can design and supply server infrastructure, integrate data center solutions, and support them through a nationwide service network. That does not replace the calculation: the request should contain three load profiles, the redundancy mode, support boundaries, and prices for the next capacity increment.

Do not move the workload in one large cutover. Establish connectivity and monitoring first, then migrate a noncritical service, run a failure test, and only then move systems with strict recovery times. Every stage needs a rollback plan, or the first missing dependency turns the migration into a late-night experiment.

Test the forecast with telemetry before buying

Local servers for your data center
S200 servers are made in Kazakhstan and provide a clear supply chain for the project.
Discuss the project

A capacity forecast becomes suitable for procurement only after it is checked against measured load and the queue of approved projects. Server nameplate power is not useful for this purpose: two 1,600 W power supplies describe the power design, but they do not prove that the server constantly draws 3,200 W. Adding those ratings overstates the working load, while a monthly average hides short peaks in the other direction.

Choose the measurement period around the business cycle. A few days are not enough for a system with weekly reporting, and one ordinary month misses quarter-end processing. If you cannot wait for the full cycle, combine the measured profile with the schedule of known jobs and clearly separate measurements from calculated additions. Keep the original time series rather than only finished charts: a scenario change may require new percentiles, peak durations, and service-concurrency calculations. One maximum without a timestamp cannot show whether the load will recur or whether scheduling can separate it.

Take readings from intelligent power distribution units or rack meters. Software CPU metrics help tune applications, but they do not see power-supply losses, network equipment, or parts of the storage system. Electricity data must come from the same boundary used to size cooling and UPS capacity. Check time synchronization, or you will not be able to connect a consumption peak to backup, month-end processing, or a batch job.

For a stable workload, capture several ordinary weeks and record processing windows, seasonal campaigns, and failovers separately. Do not flatten them into one average. The model needs average power for the electricity bill, observed maximum for checking the path, and peak duration for choosing the response. A peak lasting a few seconds may be covered by the UPS, while a sustained peak must be supported by the entire power and cooling path. If the required season has not been measured, mark it as uncertainty and schedule another check instead of inventing a coefficient.

At the same time, collect processor, memory, disk capacity, IOPS, network-port, and accelerator utilization for major services. An electrical peak and a computing shortage do not always coincide. An application may be constrained by storage latency while processors are free, and model training may occupy every accelerator without filling the rack power capacity. Buying only by kW can therefore create a site that still lacks the required resource.

After measuring, separate committed growth from probable growth. A committed project has an owner, date, budget, and a clear unit of demand. A probable project has at least a range and a launch condition. The phrase "AI will grow several times over" is not an input. A valid input can be the number of accelerators, power per configuration, operating hours, and the date when equipment must be available. If the business cannot provide these data, the stress case remains an expansion option rather than the first construction stage.

Check how much growth can be removed before buying infrastructure. Shutting down forgotten test systems, limiting uncontrolled virtual-machine creation, moving archive data to a suitable storage tier, and scheduling batch work all change the shape of the graph. This is not permission to undersize the design. Include a saving in the scenario only after changing the configuration, assigning an owner, and measuring again. A promise to "optimize later" does not reduce the capacity required for acceptance.

Create a separate set of assumptions for each of the three scenarios. Beside every figure, identify its source: a meter, monitoring system, approved project specification, or decision by the business owner. Then set a date when the assumption must be confirmed again. This record helps in talks with an integrator and in defending an in-house project because the discussion concerns a specific source of load rather than whose growth percentage sounds more convincing.

One more boundary is often missed. Capacity contracted from the utility, facility-infrastructure capacity, and capacity available to the IT load are different quantities. Losses, redundancy, and PUE connect them but do not make them equal. If a designer shows 1 MW at the utility feed, ask how many kW remain for IT after one designed component fails and at the design outdoor temperature. Ask the integrator the same question at your connection point.

You can freeze the decision when the measured base is explained, committed growth fits the planned scenario, and the stress case has a contractual or designed expansion path. Before then, an exact five-year total creates false certainty. A sound model does not guess one outcome; it shows the cost of every testable outcome.

A constraint, not the average price, decides

The correct option follows the hardest constraint: timing, data controls, available power, team skills, or financing model. A five-year average price selects the winner only when both options genuinely satisfy those constraints.

An in-house site makes sense when the organization has a large, predictable base load for a long period, access to a suitable building and power, readiness to maintain a facility team, and a need for direct control of physical infrastructure. The integrator is stronger when launch must be fast, demand jumps, in-house specialists are scarce, and forecasting errors are expensive. A hybrid wins when the base is stable but the top of the graph is unknown.

For the budget review, show one page with three scenario rows and seven figures: average IT load, peak, installed capacity, annual kWh, five-year TCO, readiness date, and delay cost. Beside them, put four pieces of resilience evidence: a power diagram, a network diagram, a backup recovery result, and a record of taking one component down for maintenance. That page is enough to remove slogans about total control and rental savings from the discussion.

Do not approve the base scenario until it has survived the stress case, and do not build the full stress case until growth is confirmed. Buy the ability to expand, but pay for working kilowatts as the load appears.

FAQ

When does an in-house data center become cheaper than an integrator's infrastructure?

It happens when the stable base load is large enough, the site already has adequate power, and monthly savings can repay CAPEX and operations before the facility equipment needs renewal. Calculate payback from your own proposals and include staffing, repairs, and delay cost.

What time horizon should I use for data center TCO?

It is usually useful to calculate at least two horizons: the contract or budget cycle and the expected life of the facility infrastructure. One five-year total hides the point when batteries need replacement, power needs expansion, or the contract must end.

Should servers be included in an in-house data center's CAPEX?

Include only the difference between the options. If you buy the same servers for an in-house site and for colocation with an integrator, their common cost does not affect the choice.

How can I verify an integrator's stated PUE?

Request the measured annual figure, its accounting boundaries, and monthly data at different utilization levels. Design PUE at full capacity does not show how much energy the facility will use under your actual load.

Is N+1 redundancy enough for critical systems?

N+1 protects against the selected component failure if the design has no shared point of failure. Loss of a room, site, or data requires a separate recovery plan, backups, and regular tests.

What should I ask an integrator about spare capacity?

Ask how many kilowatts are available today on the UPS, cooling, and generators, how long they can be reserved, and how quickly the next block can be added. Empty rack positions without secured power do not solve the problem.

How many employees does an in-house data center need?

The number depends on presence hours and response time, so begin with a shift and coverage matrix. Include facility operations, networking, security, administration, leave, and outside service.

Can I use an integrator first and build my own data center later?

Yes. Temporary infrastructure can bridge the gap until your site is ready and provide real load data. Agree on portability of equipment, data, and configurations in advance so the temporary solution does not become an expensive dependency.

How should a contract account for unpredictable load growth?

Define expansion increments, capacity delivery time, overage price, minimum payment, and the right to reduce volume. Also confirm that the operator reserves power and cooling, not only physical space.

Which matters more, CAPEX or launch time?

Meeting the business date matters more. If the in-house facility will be late, add lost margin, manual workarounds, and contractual penalties for the full waiting period.