Perspectives
Jul 28, 2026

NOAA's AI Forecasting Shift is Changing How Scientific Infrastructure Gets Built

NOAA's AI Forecasting

Authored by

Jim Crook, Director, Corporate Communications

When GDIT helped NOAA deploy a new GPU cluster, nobody knew whether it would be heavily used.

The agency had spent years evaluating AI forecasting technologies, but there was little production evidence to justify large-scale infrastructure investments. Then a new project appeared, utilization surged, and the once-idle system quickly became NOAA's most contested resource.

The episode illustrates a growing challenge for scientific computing organizations: AI is advancing faster than the procurement processes used to support it.

Leaders from GDIT - the systems integrator responsible for much of NOAA s high-performance computing infrastructure - recently described how the rise of AI weather forecasting forced them to make infrastructure decisions before the workloads themselves had fully arrived.

The result was a departure from a procurement model that has governed HPC environments for years.

We kind of went away from our evidence-based model," said Shawn Needham, HPC Architect at GDIT. "We went to an anticipatory model."

That shift may ultimately prove more significant than any individual hardware deployment.

Building for Workloads That Didn't Exist Yet

Historically, NOAA's forecasting systems have been dominated by physics-based numerical weather prediction. Massive CPU clusters ingest observational data from satellites, aircraft, ocean sensors, and weather stations before running complex simulations multiple times per day.

Those workloads are well understood.

Storage teams can analyze years of I/O patterns, file sizes, capacity growth, and metadata behavior to make highly informed procurement decisions.

The challenge NOAA faced in 2024 was different.

AI forecasting models such as GraphCast, NVIDIA's medium-range forecasting models, and other machine learning approaches were beginning to demonstrate promise, but production demand remained uncertain.

GDIT had recently deployed a GPU cluster containing NVIDIA H100 accelerators. Initially, usage was minimal. "We saw the GPUs idle really until August," Needham said.

Then the demand appeared almost overnight. A new forecasting project emerged, utilization surged, and the AI cluster quickly became NOAA's most heavily contested resource.

"Since then, it's been about 10% of our resources. It's now our most contentious resource," Needham said.

The challenge was timing. Waiting for definitive evidence would have delayed infrastructure deployment by years due to government procurement cycles.

"If we were waiting for evidence-based," Needham explained, "you wait and you have to wait an additional two years after that to propose and actually deliver the system."

Instead, NOAA and GDIT made a calculated bet that AI forecasting workloads would become strategically important before they could be directly measured.

In hindsight, the gamble paid off.

The Emerging Collision Between HPC and AI

The infrastructure requirements driving this decision reveal a broader shift occurring across scientific computing.

Traditional weather forecasting workloads are dominated by large-scale simulations running against highly optimized parallel file systems. Through years of operational experience, organizations have learned how to tune storage architectures around those access patterns.

AI introduces different behavior.

NOAA's environment contains enormous numbers of small and medium-sized files, alongside the large datasets traditionally associated with scientific computing. Metadata performance becomes increasingly important. Training pipelines and machine learning workflows introduce access patterns that don't always align with architectures optimized exclusively for large sequential I/O.

Needham described NOAA's search for a platform capable of handling both worlds simultaneously.

"We were basically looking for something that would really do good at small file performance," he said. "We were hoping this alternate system would match the need for the small file that we knew about, match the need for the AI/ML workload that might not be there yet, and also provide the parallel file system needs as well."

That combination is becoming increasingly common across research organizations.

Scientific institutions are not replacing simulation with AI. They are layering AI onto existing simulation-driven workflows. The result is an environment where storage systems must support both traditional HPC requirements and emerging AI pipelines simultaneously.

The distinction matters because organizations can no longer optimize around a single dominant workload.

Instead, they are increasingly being asked to support workloads whose future characteristics remain partially unknown.

Infrastructure as Optionality

Perhaps the most revealing aspect of NOAA's approach was not performance, capacity, or benchmark results: it was optionality.

Needham repeatedly returned to the idea that infrastructure decisions today must preserve flexibility for tomorrow's workloads.

The objective was not merely to support known applications. It was to create room for experimentation as forecasting techniques evolve.

  • Could the platform support AI training?

  • Could it function as a parallel file system?

  • Could it eventually replace separate home-directory infrastructure?

  • Could multiple storage environments converge into a single architecture?

Those questions remain open.

What's notable is that NOAA now has the ability to test them using production workloads rather than theoretical models.

"It's a really exciting time for us," Needham said. "We have the ability to explore this new use case."

That mindset represents a meaningful departure from traditional HPC planning, where infrastructure is typically designed around clearly defined requirements and expected utilization patterns.

In the AI era, organizations increasingly find themselves investing in flexibility before certainty exists.

A Preview of Scientific Computing's Next Phase

The most important lesson from NOAA's experience may be that AI forecasting is not simply a new application running on existing infrastructure.

It is forcing a change in how scientific organizations think about infrastructure planning itself.

For years, procurement teams could rely on historical data to justify future investments. Emerging AI workloads are reducing the usefulness of that approach.

When workload adoption can shift dramatically in a matter of months, waiting for perfect evidence may become more expensive than making an informed prediction.

NOAA's decision to build for anticipated demand rather than observed demand reflects a broader reality facing research institutions, government agencies, and enterprise AI teams alike.

The challenge is no longer just scaling infrastructure; it's building systems that can accommodate the workloads that haven't arrived yet.

More from this topic

Learn what VAST can do for you

Sign up for our newsletter and learn more about VAST or request a demo and see for yourself.

* Required field.