Dell added data context, preparation, and storage features to its AI data platform.
Earlier this year, Dell surveyed 3,800 enterprise IT decision-makers and AI experts worldwide about their experiences adopting and scaling AI technology. The vendor found that data—its quality, availability, management, and security—emerged as the top challenge enterprises face.
This finding validated what Dell executives have been emphasizing since launching their AI Data Platform two years ago: organizations scaling AI are constrained not by GPUs or hardware, but by data readiness.
"This is the gap that we hear from customers every day," Varun Chhabra, senior vice president of Dell's Infrastructure Solutions Group, told journalists at a recent media briefing. "The infrastructure is ready, but the data isn't, and without data that's ready for AI, investments turn from pilots into isolated pilots that never reach their full potential at scale across the organization. The real bottleneck is not access to compute. It's not often access to models. It's actually access to data. Enterprise data is simply not in a place where it's ready to support scaling AI workloads."
Enterprise data resides across fragmented locations—cloud, datacenters, file systems, databases, applications, and edge infrastructure—Chhabra noted. "Much of it is trapped in silos based on workloads, and there is often no unified way to access it. Much of it is unstructured. It's often dark, which means it's effectively invisible to AI. It is often ungoverned, leaving teams caught between AI's need for access and policies requiring control."
Dell's AI Data Platform, a core component of its AI Factory, comprises three layers. The Data Orchestration Engine handles data ingestion, preparation, labeling, and enrichment through unified pipelines, distributed control that decouples compute from storage, and native access to Nvidia NIM microservices, AI Blueprints, and templates.
The Data Engines for analytics, processing, and search enable organizations to locate the right data in hours instead of weeks. Storage Engines include PowerScale for network attached storage of unstructured data and high-throughput AI workloads, ObjectScale supporting S3-over-RDMA and Nvidia CUDA libraries for large object repositories and AI model snapshots, and the Lightning File System—a fast, software-defined parallel system announced at Nvidia's GTC 2026 in March and released the following month—designed for high-scale training and inference.
Nvidia technology underlies the platform. Beyond CUDA libraries and NIM microservices, it includes Nvidia's Nemotron Retriever models for document parsing, embedding, and reranking, and cuVS, an open-source library of GPU-accelerated algorithms for vector indexing and search.
The unified storage layers provide exabyte-scale capabilities to prevent GPUs from waiting for data and address what Chhabra calls the "pilot reproduction gap": use cases that work in demos or small pilots often fail when scaled, hampered by data governance challenges, missing or dirty data, and resulting poor AI outcomes. "That's really where enterprise scale gets held back," Chhabra said.
This week, Dell is expanding the platform with new capabilities for agentic AI. These features give agents common context and consistent meanings across scattered structured and unstructured data, simpler pathways to finding relevant context, and tools to apply it.
Agents accessing the same data and documents must currently reconstruct their understanding and regenerate tokens with each query, Chhabra explained. "This costs extra tokens and often creates churn from inevitable back-and-forth with humans in the loop, which slows things down. Often the answers aren't fully trustworthy. Context is what matters. All of these features drive fewer steps and lower costs because agents don't have to keep rebuilding context in every request. They have data flowing into the agent that reduces compute costs and token generation."
The Unified Semantic Layer provides rules and definitions ensuring words mean the same thing wherever they appear. In manufacturing, for example, the term "defect" may have different definitions across plants. The Unified Semantic Layer, which includes a searchable glossary, ensures agents understand consistent meanings. Dell integrates Nvidia's Auto-Ontology open-source library to create knowledge graphs from enterprise data.
The Enterprise Knowledge Graph shows how structured and unstructured data relate through metadata, lineage, and query history, continuously refining itself. While the semantic layer defines what words mean, the knowledge graph provides connections between data across systems. "If you're correlating defect frequency with particular suppliers or distributors in the supply chain," Chhabra said, "the Enterprise Knowledge Graph can make those connections for the model, making it much more token-efficient—rather than the model having to rerun and create tokens every time."
Knowledge Agents, built atop these layers, function as trusted agents specialized in particular topics. Enterprises can set rules governing their behavior, token budgets, and data access. These agents also leverage Nvidia's Nemotron Retriever models for reasoning and visual data understanding. Knowledge Agents give enterprises greater flexibility—choosing model sizes, deploying general-purpose models in the cloud, or running open-source models on-premises.
Dell is bringing Nvidia's cuDF GPU-accelerated library to the Data Processing Engine for faster data processing and analytics, and Apache Arrow to efficiently move data between storage and processing, allowing queries to run where data resides. GPU acceleration can deliver 20 times faster batch processing than CPUs and four times faster data processing speed overall. This complements the cuVS-enhanced search and cuDF analytics already in the Processing Engine.
A new open-source Dell Storage Performance Tool for ObjectScale and PowerScale lets organizations accurately size their AI infrastructure and maintain data flow to GPUs. Users can define their own workloads to benchmark—sequential writes, high-concurrency reads, mixed read-write environments, and complex queries—using the same tools Dell Engineering employs internally.