Thought Leadership
Sep 16, 2026

AI Factories Are Forcing HPC to Rethink the Entire Stack

AI Factories Are Forcing HPC to Rethink the Entire Stack

Authored by

Nicole Hemsoth Prickett, Head of Industry Relations

If you’ve followed supercomputing over the last couple of decades in particular the trend was predictable that the largest centers acquired the largest (and fastest) systems and would continue vying for dominance every few years. That model is starting to turn driven by both the emerging requirements for AI factories, not the mention the challenging economics behind building those.

We unpack this and more on this week’s Shared Everything podcast where I sat down with VAST Data’s SLED Director for the Americas (and long-time participant in the exciting 2000s for HC), Chris Ginder.

Chris says that the shift underway now looks less like another supercomputing refresh cycle and more like the consolidation that began across the Department of Energy laboratory system decades ago. The difference this time is that the pressure is coming from several directions at once.

GPUs are expensive and constrained, power is increasingly difficult to secure, new systems require dense infrastructure and liquid cooling and even datacenter construction itself has become more complicated and costly. At the same time, AI is expanding the number and types of researchers who expect access to high-end computing.

Those forces are pushing universities, research institutions and regional organizations toward infrastructure that can be shared across a much broader population he says. This is where the AI factory starts to become more useful as an architectural concept than as an industry buzzword.

The AI Factory Is More Than A GPU Cluster

Traditional HPC environments were often built around highly specialized users, applications and filesystems. A relatively small group of researchers and administrators understood how to extract performance from environments based around technologies such as Lustre or GPFS.AI changes both the workloads and the population using the infrastructure.

An AI factory has to ingest large quantities of unstructured data, move that data through automated pipelines and workflows, train or run models against it, and support inference. Compute remains critical, but it now sits inside a larger operational stack that includes high-performance networking and storage, orchestration through environments such as Slurm or Kubernetes, accelerator software and libraries, and increasingly sophisticated data management.

That makes starting the architecture discussion with GPU count potentially dangerous.

Ginder says many organizations understandably begin there. GPUs are generally the most expensive and visible component of the system, so planning starts with accelerator performance and how much compute can fit within a given budget. But a GPU-heavy design can create another problem: a very expensive pool of compute waiting for the rest of the system to deliver data.

The relevant performance question is therefore not simply how fast the GPUs are. It is where saturation occurs across the complete system.

The bottleneck could be storage. It could be northbound or southbound networking. It could be the processing pipeline. And as models and datasets grow, an architecture that looks balanced for today’s workload can become badly unbalanced later.

For research organizations building infrastructure expected to remain useful for years, feeding the GPUs is as important as buying them.

Shared Infrastructure Changes The Design Problem

The move toward shared AI factories adds another layer of complexity.

Initiatives such as Empire AI and the Massachusetts AI Hub model infrastructure that can serve researchers beyond a single department or even a single institution. That creates economies of scale around equipment that is becoming difficult for individual universities to fund, power and operate independently but it also changes what the underlying infrastructure has to provide.

A shared AI factory cannot assume every user has deep HPC expertise. Researchers, principal investigators and students may come from different universities and disciplines, with radically different technical backgrounds. The system has to provide a more common operating model while maintaining the performance expected from HPC infrastructure.

More importantly, sharing compute does not mean sharing everything else, which is an important distinction Ginder makes. Research groups need their data segregated. Institutions may have sovereignty requirements. Access needs to be controlled at multiple levels. Models and datasets may be tied to grants or regulated research, which means changes throughout the research process need to be tracked rather than disappearing when a model or dataset is modified or deleted.

Data governance therefore becomes part of the architecture rather than an administrative layer added after the machine is installed.

This is one of the fundamental differences between building a large accelerator cluster and building an AI factory. The latter has to combine utilization, accessibility, isolation, security, provenance and performance in the same environment.

And because the hardware involved is so expensive, these systems need to operate at extremely high utilization. Idle infrastructure is not merely wasted capacity when thousands of GPUs are involved. It directly undermines the economics that justified sharing the system in the first place.

The Payoff Is In The Research

The architectural argument becomes clearer when looking at what researchers are actually doing with these systems.

Ginder pointed to medical imaging as one example. At the University of Rochester, infrastructure improvements have reduced the time required to process certain cardiology images from roughly a minute to a matter of seconds. That alone changes the productivity of an extremely specialized human resource.

Models can assist in sorting large imaging datasets, identifying images that appear routine, those that warrant additional review, and those that may require immediate attention. The physician remains in the loop, but the infrastructure allows scarce expertise to be directed toward the cases where it matters most.

Agriculture presents a completely different data problem with a similar systems requirement. Universities are combining drone imagery, weather information, climate patterns and other datasets to understand conditions at much finer geographic resolution. Those models can help determine which crops are best suited to particular areas and potentially improve agricultural yields.

Other research projects are analyzing animal migration and coastal conditions in near real time. Still others are using computer vision to analyze crops in the field, including determining when individual produce is ready for harvesting.

These workloads have little in common scientifically but architecturally, however, they share something important: large and changing datasets need to move through pipelines, reach expensive compute quickly, remain governed and accessible to multiple users, and produce results on useful timescales.

HPC Economics Are Forcing Collaboration

There is another reason this model is likely to persist even if accelerator supply improves, which is that the economics of building research infrastructure have changed.

Power, cooling, construction, accelerators and specialized personnel all make it increasingly difficult to replicate top-tier infrastructure across every institution that could benefit from it. Ginder expects this to drive considerably more collaboration among universities, including institutions that possess strong research talent but lack direct access to large-scale computing.

Instead of asking how large a supercomputer they can build, they may increasingly ask what research communities they need to support, which resources should be shared, how those resources can be funded across institutions, and how the resulting infrastructure can remain useful as AI workloads evolve.

There’s an interesting parallel here with the earlier consolidation of national laboratory supercomputing. Scarcity and scale eventually made isolated infrastructure less practical. The same process is now reaching universities at precisely the moment AI is dramatically broadening demand for advanced computing.

The result could be something larger than a new class of supercomputer.

If AI factories work as intended, they become shared research infrastructure: GPUs, networks, data systems and software operated as one environment, with enough performance to satisfy HPC workloads but enough accessibility to reach researchers who would never have been traditional supercomputing users.

And that may ultimately be the most significant change AI brings to HPC. The defining metric for the next generation of research infrastructure may not be how much compute one institution owns, but how effectively a much larger research community can use it.

More from this topic

Learn what VAST can do for you

Sign up for our newsletter and learn more about VAST or request a demo and see for yourself.

By proceeding you agree to the VAST Data Privacy Policy.

* Required field.