NVIDIA and AWS jointly announced EC2 G7 instances powered by Blackwell GPUs and GPU-accelerated vector search in OpenSea
Building AI systems at scale demands low-latency inference, fast vector search, strong GPU price-performance and infrastructure that grows without multiplying operational complexity. NVIDIA's latest work with Amazon Web Services addresses each of those constraints through advancements across Amazon OpenSearch and Amazon EC2.
Amazon EC2 G7 instances powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs expand the compute layer for AI, graphics, video and data analytics workloads. Compared with G6 instances, G7 delivers up to 4.6 times AI inference performance, up to 2.1 times graphics performance and significantly faster GPU-accelerated data analytics on Amazon EMR using the NVIDIA cuDF library for Apache Spark workloads. With support for up to eight GPUs, 256 gigabytes of total GPU memory, 700 gigabits per second of EFA-enabled networking and up to 7.6 terabytes of local NVMe SSD storage across one, two, four and eight GPU configurations plus bare metal coming soon, G7 instances let customers right-size infrastructure for their workloads instead of over-provisioning. The instances are accessible through AWS Deep Learning Amazon Machine Images, Amazon Deep Learning Containers, Amazon EMR, Amazon EKS, Amazon ECS and graphics AMIs, with availability coming soon to Amazon SageMaker AI.
Amazon OpenSearch Serverless powers agentic AI and dynamic workloads with no infrastructure management required by using GPU-accelerated vector indexing powered by NVIDIA cuVS as the default compute choice for all vector collections. For teams building retrieval-augmented generation, semantic search, recommendation systems and agentic AI applications, this shift turns GPU-powered vector search from a specialized optimization project into a standard AWS capability. Vector indexing runs up to ten times faster at a quarter of the cost compared with CPU-only builds, making billion-scale vector databases practical to build in under an hour.
AWS has achieved NVIDIA Exemplar Cloud status on NVIDIA GB300 for training workloads, meeting the rigorous performance thresholds NVIDIA uses to benchmark AI workloads against its reference architecture. This achievement results from deep co-engineering efforts between AWS and NVIDIA teams, allowing developers and AI leaders to use consistent, high-performance cloud infrastructure for large-scale training with greater confidence in total cost of ownership and faster time to production.