Thought Leadership
Sep 1, 2026

Failure is a Finding: The Hard Lessons Behind Computing’s Biggest Bets

Failure is a Finding: The Hard Lessons Behind Computing’s Biggest Bets

Authored by

Nicole Hemsoth Prickett, Head of Industry Relations; Glenn Lockwood (Principal Technical Strategist, VAST Data)

Our industry is is very good at documenting success but as you might imagine, failure is harder to find, even when it produces the more useful lesson.

In this episode of Shared Everything, AI and HPC veteran (and current Principal Technical Strategist at VAST) Glenn Lockwood looks back at three technologies he encountered across his career in supercomputing, national labs and hyperscale AI. These are not stories about obviously bad ideas or failed systems, that would be too simple: these are examples of ambitious, technically sound bets inside highly successful systems that did not play out quite as expected.

Together, they produced three findings about building infrastructure at scale. And this is worth a close, thoughtful listen.

Among Glenn’s findings? Three important ideas stand out based on experience with what were then among the world’s largest systems…

First, There Is No Magic

The first comes from Gordon, the supercomputer deployed at the San Diego Supercomputer Center in 2012. Gordon was an important system, including one of the earliest large-scale uses of enterprise flash in supercomputing.

It also included ScaleMP’s VSMP, an attempt to make distributed computing easier. VSMP could combine many physical compute nodes and present them to a programmer as one large shared-memory system. The idea was to let researchers scale applications without having to master MPI and distributed programming.

The problem was adoption. Gordon was initially expected to have 32 VSMP supernodes. Within two years, that had fallen to two.

VSMP hid the complexity of distributed computing, but the complexity was still there. Programmers looking for good performance still had to understand memory locality, NUMA domains and communication between physical nodes. And the people capable of doing that were often already capable of writing distributed applications.

The lesson goes well beyond VSMP. Compilers, storage tiers and other abstractions can make complex systems easier to use, but they cannot repeal the architecture underneath them. There is no magic. You can hide complexity, but you cannot make it disappear.

Everyone’s Problem Gets Everyone’s Fix

At NERSC, Lockwood worked on Cori and its DataWarp burst buffer. At the time, flash was too expensive to use everywhere, so DataWarp put a relatively small flash tier in front of a much larger disk-based parallel file system.

It worked. Users could get roughly twice the I/O performance with relatively minor changes to their jobs.Yet utilization of the burst-buffer nodes averaged only around 15 percent.

The reason was surprisingly simple. The regular parallel file system was already fast enough. For many users, 2X better performance did not justify changing workflows, even slightly. They preferred something familiar that could move easily between systems and computing centers.

Then the larger market changed the equation entirely. Everyone wanted flash, not just HPC. Flash prices fell, and subsequent NERSC systems moved to all-flash storage rather than maintaining a specialized burst-buffer architecture. That led to the second finding.

Everyone’s problem gets everyone’s fix.

A specialized technology may solve one community’s problem exceptionally well, but broadly useful technologies have economics, ecosystems and skills working in their favor. Lockwood sees the same dynamic in the adoption of technologies such as Kubernetes in HPC. It does not have to be the perfect HPC solution if it solves the same problem across a much larger world and is good enough for HPC too.

And a Biggie: Big Bets Turn All at Once

The third finding comes from a radically different scale.

At Microsoft, Lockwood worked on Fairwater, part of a generation of AI supercomputers connecting enormous numbers of GPUs into tightly coupled systems. The rationale for building them was strong. Scaling laws suggested that larger models, more data and more compute could produce predictably better AI.

GPT-4 appeared to validate the idea. The industry responded by building much larger training infrastructure. But infrastructure takes years to build and AI didn’t wait.

Reasoning models emerged while these systems were being developed, shifting more computation toward inference and changing assumptions about what infrastructure the next generation of models would require.

Lockwood is careful not to call Fairwater a failure. It is a productive and extraordinary system. The finding is about the risk inherent in making enormous infrastructure decisions around workloads changing far faster than the infrastructure itself.

A tightly coupled machine designed around one assumption is difficult to redesign halfway through construction when that assumption changes.

That may be the most consequential lesson as AI infrastructure moves into increasingly expensive territory. Sometimes there is no way to know whether the bet is right until someone builds the system.

More from this topic

Learn what VAST can do for you

Sign up for our newsletter and learn more about VAST or request a demo and see for yourself.

* Required field.