[Dong-hyuk Kim (김동혁), SA team leader at HS Hyosung Information Systems] More than half of corporate AI investment now goes not into building models but into running them. Problems long handled by data centres, such as storage failures, capacity forecasting mistakes and delayed refresh cycles, are already known, and response manuals are in place, making them familiar to solve. But the burden of investment work in 2026 is different. Corporate AI investment has focused on model training, but the centre of gravity in AI infrastructure has fully shifted from training to inference. It is a workload that runs at large scale all day without stopping.
This shift is no longer a forecast but a reality already unfolding. The real issue is not that inference is technically more difficult than training. It is that not many infrastructure organisations have actually experienced it in live operating environments. By the time they move to respond, the baseline will already be far ahead.
Temporary training and permanent inference
According to Deloitte, inference accounted for about half of total AI computing in 2025. That is a rapid increase given it was only one-third in 2023. Multiple analyses also commonly confirm that 80 to 90 percent of total operating costs for production AI systems occur at the inference stage. This structure arises not because each individual request is costly, but because the system keeps running without stopping.
Once deployed, a model operates continuously. If training imposes a momentary load on a data centre, inference changes the data centre structure itself. As services with new features or workflow automation that embed AI increase one by one, the baseline the infrastructure must handle rises as well, and this baseline does not come back down. It is a cost structure entirely different from the past, and it requires an infrastructure strategy with a different approach.
Why inference is pulling workloads back on premises
Early enterprise AI spread on a cloud basis. In the experimental stage, it was a reasonable choice. But as inference moves into real operations, the calculation has become more complex. To reduce latency, computing must be done close to the data, and as scale grows, cost burdens accumulate quickly. In addition, the so-called Data Gravity effect is regaining strength, where the larger the data or the more it is controlled or regulated, the better it is to compute in place rather than move it.
Cloud remains suitable for training or handling momentary traffic in experiments. But for production-stage inference, hybrid structures are becoming an intentional design from the start, not a transitional alternative. Leading organisations choose not between cloud and on premises, but to place workloads by weighing data location, latency requirements, governance conditions and cost structures. The conclusion is increasingly pointing toward on premises.
A computing problem and a data problem
When AI enters operations, bottlenecks often arise from data accessibility rather than computing capacity. Inference workloads have a high share of reads, are sensitive to latency and, above all, require up-to-date managed data. Models run not on past snapshots but on real-time, live enterprise data. Most of this data exists in on-premises systems and operational databases that were not originally designed with AI pipeline integration in mind. The practical problem arises from the gap between where the data is and where the model needs that data.
Copying data through a separate pipeline increases both latency and governance risk. In regulated industries, each time data moves it becomes a compliance issue in itself. Ultimately what is needed is a structure that accesses data directly near the computing that needs it, from the place where the data originally resides. As inference becomes always-on and evolves toward agent-centred systems, what models need also shifts from fixed datasets to live operational data that can be repeatedly accessed under policies.
Hitachi Vantara addresses this issue from the design stage. VSP (Virtual Storage Platform) 360 is an AI-based control plane for enterprise data infrastructure. It coordinates provisioning and overall operations of storage and data protection platform services to support the consistency, performance and availability that production inference demands every day. On top of that, Hitachi IQ Studio provides AI-based data management and orchestration services used by inference workloads. It supports direct access to enterprise data through policy-based access without needing to copy or separately stage data, and it applies governance, data lineage, quality management and observability to the data provided to models.
The cost structure also changes. Once established, inference is close to a utility embedded deep in operations and difficult to reduce. This is why infrastructure operations organisations need to move away from fixed, project-based thinking to a consumption model that continuously optimises recurring costs.
Data centres moving to the next stage
A common pattern appears among organisations leading this shift. The centre of gravity is moving from planning around peak-time demand to baseline management, from asset ownership to service consumption, and from one-off projects to an always-on operating platform that must function stably every day.
The competition to train models quickly is already past. From now on, competitiveness is determined by how steadily inference can be operated with data that has governance applied, is accessible and is close to computing. Only prepared infrastructure survives into the next phase.