AI will change data centres from the rack up

Executive Summary

  • Jean-Christophe Thery, Founder & CEO, MusaArtGallery, talks about the AI boom as it is embedded into everyday life and business; data centres are going to conduct a balancing act between latency, performance, cost, etc., because adding more general-purpose capacity will not be enough.
  • The next generation of AI infrastructure is likely to be more specialised; for many business applications, a smaller specialised model can produce the required outcome with far less compute.
  • The AI boom is often described as a race to build larger data centres. In reality, it may force the industry to adopt a broader definition of performance.

 

The AI infrastructure debate is often framed as a race for more computing power. The more important shift may be how that computing power is allocated – by workload, by geography and increasingly by access to energy.

The infrastructure question is changing

The AI infrastructure conversation is often reduced to one question: how much more computing power will we need?

That matters, but it misses the more interesting change. As AI becomes embedded in everyday business processes, data centres will have to balance performance, cost, latency and energy availability at the same time. Simply adding more general-purpose capacity will not be enough. The next generation of AI infrastructure is likely to be more specialised, more distributed and much more conscious of where both data and power are located.

Compute will become more specialised

Training a large model and serving millions of inference requests are fundamentally different jobs. They do not necessarily belong on the same infrastructure, or even in the same place.

Large-scale training will continue to favour highly concentrated environments with accelerators, high-speed networking and access to substantial power. Inference is more varied. Some workloads need the scale of centralised cloud infrastructure; others benefit from being closer to users, devices or the data source itself.

This means infrastructure planning will increasingly move away from the idea of a universal server estate. The practical question becomes: what type of compute belongs where? A rack optimised for training, a regional inference cluster and an edge deployment may all be part of the same AI system, but they solve different problems.

Data locality becomes a business decision

AI systems depend on data, but moving that data is not free. Latency, bandwidth costs, privacy rules and cross-border restrictions can all influence where a workload should run.

For organisations using AI in manufacturing, logistics, retail or real-time analytics, repeatedly sending large volumes of information to a distant region can undermine the value of the application. In those cases, bringing compute closer to the data can make more sense than bringing the data to the compute.

That makes data locality more than a compliance issue. It becomes an architectural and economic decision. Over time, the location of data may influence infrastructure design almost as much as the demand for raw computing power.

Power is becoming an operating constraint

The most important limitation may be physical rather than digital: electricity supply cannot always expand at the same speed as demand for compute.

The International Energy Agency’s 2026 update projects global data-centre electricity consumption rising from about 485 TWh in 2025 to around 950 TWh in 2030, while electricity use from AI-focused data centres is expected to grow even faster. The issue is not simply the global total. Data centres concentrate very large loads in specific locations, where grid connections, transformers and generation capacity can become bottlenecks.

That changes how operators should think about performance. Not every AI workload needs to run immediately. Training jobs, batch processing and other flexible tasks can sometimes be shifted to periods or locations where capacity is more available. Smarter workload scheduling, therefore, is not just an efficiency feature. It can become part of the energy strategy of the data centre itself.

Model efficiency becomes infrastructure strategy

AI discussion naturally gravitates towards larger and more capable models, but from an infrastructure perspective, bigger is not automatically better.

For many business applications, a smaller specialised model can produce the required outcome with far less compute. Techniques such as quantisation, distillation and model compression can reduce the hardware and energy required to serve a workload at scale.

That means model choice and infrastructure choice can no longer be treated as separate decisions. If two models deliver an acceptable business outcome but one requires materially less compute, the impact reaches beyond software cost: it affects server capacity, cooling, power demand and potentially where the application can be deployed.

Efficiency will become part of performance

The AI boom is often described as a race to build larger data centres. In reality, it may force the industry to adopt a broader definition of performance.

A system with enormous theoretical compute capacity can still be a poor design if it is too expensive to operate, too far from the data it needs or constrained by local power availability. The stronger architecture may be the one that delivers the required result with the right mix of specialised hardware, sensible data placement and efficient models.

The organisations that manage this transition well will look beyond raw computing capacity. They will measure performance not only in tokens, throughput or accelerator counts, but also in cost, latency, energy and the amount of data that has to move to make the system work.

AI will certainly require more compute. The more important transformation is that it will force us to use that compute far more intelligently.

Share this Post: