Matilda Bailey is a veteran networking specialist who has spent her career navigating the shifts from cellular expansion to the current explosion of next-generation AI infrastructure. With her deep background in high-bandwidth solutions and wireless technology, she brings a unique perspective to the modern data center dilemmwhy are we building more when so many racks sit empty? As organizations scramble to support a new era of compute, Matilda helps us understand that the disconnect between physical floor space and usable power is the most expensive hurdle in today’s landscape. We dive into the nuances of “stranded capacity,” the true cost of legacy bottlenecks, and why the old metrics of CPU utilization are no longer enough to measure success as we move through 2026 and beyond.
The conversation covers the critical distinction between installed infrastructure and productive capacity, highlighting how bottlenecks in cooling or networking can render expensive hardware useless. We explore the financial drain of underutilized facilities, the diverging infrastructure requirements for AI training versus inference, and the shift toward measuring output through metrics like tokens per kilowatt and request latency.
Many facilities currently find themselves in a paradoxical situation where they have plenty of open rack space but lack the power or cooling required for modern AI clusters. How should operators reconcile this “ghost” capacity with the constant pressure to build new infrastructure?
The tension we are seeing right now comes from a fundamental misunderstanding of what “capacity” actually means in a modern context. We have a tendency to treat installed infrastructure and usable capacity as the same thing, but as Siraj Aziz from Omdia points out, that is a costly contradiction. You can walk into a data center and see rows of empty racks, but if the facility’s electrical distribution was designed for 5 kilowatts per rack and your new AI accelerators need 40 or 50 kilowatts, those racks are essentially ghosts. It isn’t just about the physical footprint; it’s about whether the systems and workloads can actually use those resources to meet business requirements. Operators have to ask themselves why they can’t put more workloads on existing floors, and the answer is usually that the infrastructure is out of balance. Building new is often a knee-jerk reaction to a bottleneck that might be solved by smarter orchestration, but more often, it’s a sign that the legacy facility has reached a hard ceiling in power density.
When we talk about underutilization, the focus is often on the hardware, but what are the hidden operating costs that keep piling up even when these servers are sitting idle?
The financial drain of underutilized capacity is incredibly pervasive because so much of a data center’s cost is “always-on” capital and operational expenditure. When you have stranded capacity, you aren’t just wasting the server; you are paying for the UPS systems, the electrical distribution, the cooling equipment, and the facility itself that you’ve already bought and paid for. Beyond the hardware, there is the human and maintenance element—staffing, connectivity, 24/7 monitoring, and lifecycle management costs continue to tick upward regardless of whether a single token is being processed. A bottleneck in one area, like a lack of high-bandwidth networking, can effectively “freeze” the value of every other asset in the facility. It is a frustrating reality where you are bleeding money on maintenance for a high-performance cluster that can’t reach its potential because it’s waiting on a legacy router or a constrained storage pipe.
AI workloads are often discussed as a single category, but training and inference have very different needs. How does this distinction change the way you approach facility planning and upgrades?
We have to stop looking at AI as a monolith because the infrastructure dependencies for a massive training cluster are worlds apart from a regional inference node. Training clusters are these high-density, tightly coupled environments that require massive accelerators, specialized high-bandwidth networking, and substantial storage throughput to keep the GPUs fed. These setups often push rack power densities so high that you have to reconsider the entire cooling strategy, sometimes moving toward liquid cooling which introduces a nightmare of new plumbing, permitting, and facility requirements. Inference, on the other hand, is much more sensitive to latency and concurrency requirements depending on the specific model and the number of users hitting it simultaneously. You cannot assume a facility that supported traditional enterprise apps can just “pivot” to AI; if you don’t have the specific infrastructure to support the model’s throughput needs, you’re just creating more stranded capacity.
You mentioned that low CPU or GPU utilization doesn’t always mean you have excess capacity. Can you explain how a system-level view changes our understanding of performance bottlenecks?
It is a common trap to look at a dashboard, see 20% GPU utilization, and assume you have 80% of your investment just waiting for more work. In reality, that processor might be sitting there “starving” because it’s waiting on memory, storage, or a sluggish networking handshake to deliver the next packet of data. Capacity has to be evaluated across the entire system pipeline, or you risk making very expensive mistakes in your provisioning. Siraj Aziz has been very vocal about this, emphasizing that in the era of agentic AI, visibility into network and compute remains too fragmented. If you don’t have a holistic view of the workload pipeline, you might spend millions on more compute power only to find that your performance hasn’t budged because the real constraint was an under-provisioned storage tier or a lack of rack-level power.
For an operator looking at a legacy facility, the choice between a selective upgrade and a full-scale retrofit for liquid cooling is a massive financial decision. What criteria should they use to justify that kind of investment?
This is where the concept of opportunity cost becomes the deciding factor for every operator I talk to. Bruce Bateman from Omdia makes an excellent point: just because there is a massive demand for AI doesn’t mean every legacy facility should be gutted and turned into a high-density AI hub. A colocation provider might have empty racks, but if they can’t meet the power density a customer needs for a 2026-era workload, those racks will never generate revenue. You have to look at the total retrofit cost—the plumbing, the electrical upgrades, the new routers—and weigh that against the specific workload you intend to host. In many cases, it makes more sense to keep a legacy facility running conventional workloads that still generate value, rather than forcing a full conversion that might never reach a positive ROI. You only pull the trigger on a full retrofit when the expected workload and its revenue potential justify the total cost of ownership and the massive effort required to make that capacity truly usable.
As we move away from traditional metrics like “uptime” or “CPU load,” what are the new benchmarks that actually tell a story of success for modern data center leaders?
The most useful measures now are the ones that actually reflect the work being accomplished relative to the resources being burned. We are moving toward a focus on tokens per kilowatt, which Bruce Bateman highlights as a critical metric for any power-constrained facility. It forces you to ask: “How much actual AI output are we getting for every megawatt we pull from the grid?” In the past, we cared about energy consumed per unit of work, but for AI, we have to look at request latency, token throughput, and concurrency as the primary indicators of health. For a colocation operator, the metric shifts even further toward revenue or contribution margin per sellable kilowatt rather than just server utilization. The goal isn’t to hit 100% utilization on every asset—you need that headroom for spikes and maintenance—but to ensure that the “useful” workload capacity is maximized within the constraints of your energy budget and TCO.
What is your forecast for how data center operators will manage the balance between legacy infrastructure and AI demands over the next two years?
From 2026 through 2028, I expect we will see a major shift toward “precision retrofitting” rather than broad-stroke infrastructure builds. Operators are becoming much more disciplined; instead of just adding more racks, they are identifying the exact resource—be it power, cooling, or networking—that constrains a specific workload and calculating the surgical investment needed to remove that bottleneck. We will see a rise in modular AI “enclaves” within older facilities, where high-density liquid cooling is brought in specifically for inference clusters while the rest of the floor continues to hum along with traditional air-cooled enterprise apps. The successful leaders will be those who stop chasing maximum utilization for its own sake and start focusing on productive workload capacity as the ultimate measure of business value. We are entering an era where the most efficient data center isn’t necessarily the biggest or the newest, but the one that has most effectively eliminated its stranded capacity.
