Thought Leadership
Jul 29, 2026

The Data Mess the AI Era Exposes

The Data Mess the AI Era Exposes

Authored by

Nicole Hemsoth Prickett, Head of Industry Relations

When people talk about AI infrastructure, the conversation almost always begins with GPUs, power, cooling, or the latest large language model. And while yes, those are certainly important stories, after talking with Howard Marks I found myself thinking about something much less glamorous. AI isn’t just changing what enterprises compute, it’s changing how they think about the information they’ve been collecting for decades.

The organizations that will benefit most from AI may not be the ones with the biggest clusters or the fastest networks. They may simply be the ones that finally understand what data they have, where it lives, and how to make it useful, at least such is the line of thinking the veteran of IT thinks, as unpacked in the episode below.

That sounds like an odd conclusion from someone who has spent more than forty years immersed in computing infrastructure (now at VAST), but perhaps that’s exactly why it’s so compelling. Howard has watched nearly every major architectural shift, beginning long before enterprise storage became its own discipline, and his perspective is refreshingly free of hype.

At one point he describes how AI has quietly broken one of the industry’s longest-held assumptions. For decades, hardware became smaller, faster, and cheaper with remarkable consistency. Organizations could afford to be inefficient because next year’s systems would almost always solve yesterday’s mistakes. As he explains, “the demand from AI has gotten so big that the cheaper part just went away completely.”

That observation reaches much further than the price of flash or memory. If infrastructure no longer gets dramatically cheaper every year, then every unnecessary copy of a dataset, every forgotten archive, and every poorly understood repository becomes more expensive than it used to be. AI doesn’t create those problems. Instead, it shines an uncomfortably bright light on them because AI depends on finding, organizing, and reasoning over information that many organizations barely knew they possessed. Howard’s conclusion is surprisingly practical.

Rather than talking first about bigger systems, he argues that “it’s time to invest more in data management. It’s time to do more understanding of what your data is.”

What struck me most is that this feels like a return to fundamentals rather than another technology cycle. We have spent years talking about collecting more data because storage was inexpensive enough that there was little reason not to. Now enterprises are asking entirely different questions. Which data is duplicated? Which information actually has value? What has been sitting untouched for years that suddenly becomes useful because AI can understand documents, conversations, images, or recordings that traditional analytics never could? The challenge is no longer preserving information. It’s creating enough context around that information that it can participate in an AI workflow instead of remaining another forgotten collection of files.

That may also explain why so many organizations suddenly find themselves rethinking infrastructure. Howard makes the point that people rarely replace systems simply because something newer exists. They replace them when the old approach no longer supports what they’re trying to accomplish. AI has become exactly that catalyst. It isn’t that existing storage suddenly stopped working. It’s that existing architectures were never designed to answer the kinds of questions AI now makes possible.

The conversation eventually turns to governance, and here again Howard avoids dramatic predictions in favor of something much more practical.

If AI agents are going to retrieve information and act on behalf of users, then the policies that determine who can access what information cannot remain an afterthought. They belong in the infrastructure itself because that’s where identity, permissions, and the relationships between data already exist. It’s an argument that feels remarkably sensible. AI may be introducing entirely new capabilities, but the responsibility to protect information hasn’t changed. The only difference is that the infrastructure now has to make those decisions for both people and the software acting on their behalf.

By the end of the conversation, I was left thinking much less about AI models than about the quiet work that has to happen before those models become useful. We have spent the last several years celebrating what AI can do once it has access to information.

Howard’s point is that enterprises are only beginning to appreciate what it takes to make that information understandable, trustworthy, and secure in the first place.

And indeed, that may not be the headline grabbing story of the AI era, but it may prove to be the one that matters longest.

More from this topic

Learn what VAST can do for you

Sign up for our newsletter and learn more about VAST or request a demo and see for yourself.

* Required field.