Ai

VAST Data and AMD Expand AI Infrastructure Alliance to Tackle the Industry's Inference Bottleneck

Kasun Illankoon

By: Kasun Illankoon

5 min read

For most of the past three years, the artificial intelligence industry has measured progress in a single unit: the number of GPUs a company can buy. More chips meant more training capacity, and more training capacity meant a faster path to bigger models. That era is ending. As enterprises shift from building models to running them, inside agents, chatbots, and reasoning systems that operate continuously, the bottleneck has moved.

[For more news, click here]

It is no longer simply how many processors a company owns. It is how fast data can move in and out of those processors, and how efficiently the memory that keeps an AI system reasoning can be managed.

That is the wager behind an expanded collaboration between VAST Data, the AI infrastructure company, and AMD, announced this week. The two companies are combining the VAST AI Operating System with AMD's sixth generation EPYC processors and Instinct GPUs, aiming squarely at what has quietly become the costliest and least understood problem in enterprise AI: inference, the ongoing work of putting a trained model to use.

The Shift From Training to Thinking

Training a large language model is, in infrastructure terms, a contained problem. It has a start date, an end date, and a predictable shape. Inference does not. Once a model is deployed inside a customer service agent, a coding assistant, or a research tool, it runs continuously, often holding onto long threads of conversation and context that can stretch for hours. Engineers call the memory structure that stores this context the key value cache, or KV cache, and keeping it close to the GPU without overwhelming its limited onboard memory has become one of the defining engineering challenges of 2026.

John Mao, VAST Data's vice president of global technology alliances, said inference is fundamentally a data problem. “The industry is discovering that inference is fundamentally a data problem,” Mao said, adding that success depends on bringing data, compute, memory and intelligence together as a single system.

Storage Earns a Seat at the AI Table

The technical core of the expanded partnership is VAST's Disaggregated Shared Everything architecture, paired with AMD's sixth generation EPYC CPUs, code-named Venice, and the newer Instinct MI355X GPU. In early testing, VAST said offloading KV cache data to its storage cluster over a high-speed network connection produced a ninefold speedup in the time it takes a model to generate its first response token, and nearly a tenfold increase in overall token throughput under heavy concurrent use.

Those numbers matter because token generation is now a direct cost line for AI companies, not an abstract performance metric. Slower time to first token means idle GPUs, and idle GPUs are among the most expensive assets in modern computing.

Derek Dicker, AMD's corporate vice president of its Enterprise Business Group, framed the collaboration in terms of choice. “The future of AI will be built on an open ecosystem,” Dicker said, one that lets organizations pick the technologies that fit their performance and business needs.

A Blueprint for the AI Factory

Alongside the hardware pairing, VAST, AMD and networking firm DriveNets published a joint reference architecture that combines AMD's rack-scale Helios infrastructure with DriveNets' AI Fabric networking and the VAST AI Operating System. The document gives AI cloud providers and large enterprises a validated starting point for building what the industry has taken to calling AI factories, data centers purpose-built for training, inference, and reinforcement learning at scale, rather than a checklist they have to assemble themselves.

Yossi Kikozashvili, DriveNets' vice president and head of product for AI infrastructure, said customers increasingly want the flexibility to build AI infrastructure using the technologies that best meet their requirements.

Compliance Gets Built In

One detail in the announcement speaks to how enterprise AI has matured past the experimentation phase. Because KV cache data can contain fragments of sensitive customer conversations, VAST is applying its existing data lifecycle policies to automatically expire and delete cached information, a feature designed to help regulated industries meet privacy and compliance requirements without manual cleanup. It is a small addition, but it reflects a broader pattern. As AI systems handle more real customer data in production, the infrastructure supporting them is being asked to behave less like a research tool and more like a bank's core system.

The Ecosystem Widens

The expanded partnership also folds in two younger companies working specifically on inference optimization. TensorMesh and EmbeddedLLM are both building software layers meant to make KV cache reuse and GPU utilization more efficient on top of AMD's ROCm software stack.

Pin Siang Tan, Embedded LLM's co-founder and chief technology officer, said agentic AI moves the inference bottleneck beyond raw compute to context.

That framing, that AI's hardest problems are increasingly about context and memory rather than raw processing power, was echoed across the roster of cloud providers VAST and AMD named as customers.

Raghu Chakravarthi of Core42 said customers need infrastructure that can securely and reliably support increasingly complex workloads without sacrificing performance. Erwan Menard of Crusoe said the collaboration gives customers a validated foundation purpose-built for AI. Kevin Cochrane of Vultr said the partnership gives customers more flexibility in how and where they deploy AI.

David Bitton of 5C said validated architectures help reduce complexity and accelerate deployment. Piotr Tomasik of TensorWave said the market is looking for high-performance alternatives that provide both scalability and flexibility. Dhanaseker Kandhasamy of Phanos.AI said managing large and concurrent context windows is one of the biggest challenges in large-scale AI.

A Bet on What Comes Next

Taken together, the announcement is less about a single product than about where the AI infrastructure market is placing its bets for the next phase of the buildout. Nvidia's GPUs and CUDA software still dominate AI training. But as inference, not training, becomes the larger and more persistent cost center for enterprises, AMD, VAST and the cloud providers betting on them are wagering that an open, storage-centric approach can compete on the metric that will matter most going forward: how cheaply and reliably an AI system can keep thinking.

Related Articles:

IBM and Red Hat Launch Lightwell to Automate Open Source Vulnerability Fixes for Enterprises

How Outpost24's New CyberFlex Program Helps Security Teams Keep Pace With AI-Era Risk

France Moves to Remove Microsoft From 2.5 Million Government Computers in Push for Digital Sovereignty

Share this article

Related Articles