Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

OpenAI has officially unveiled its first proprietary AI chip designed for its Stargate infrastructure.

Reduces hyperscaler reliance on third-party accelerators and signals vertical integration trends that will reshape GPU procurement volumes.
Trade pressSlicast · August 28, 2026 · US · Source: Google News
importance 80

This week in AI, attention centers on custom inference hardware, highly efficient multimodal models, and on-device intelligence. From OpenAI’s development of its own inference chip to Z.ai’s launch of a cost-optimized model, and Liquid AI’s new framework for benchmarking models directly on consumer devices, artificial intelligence is becoming faster, more efficient, and increasingly capable of operating closer to end users.

OpenAI has published initial performance data for Jalapeño, a custom AI inference chip developed in partnership with Broadcom. Engineered specifically for large language model serving workloads, the chip features joint optimizations across compute, memory movement, networking, and inference serving. Early benchmarks indicate that Jalapeño delivers 1.5–1.9× greater compute per watt and 1.7–3.6× lower latency than comparable NVIDIA Blackwell-based systems. Unlike many architectures that force a trade-off between throughput and latency, Jalapeño is designed to enhance both simultaneously. According to OpenAI, the chip has already accelerated ChatGPT response times and improved the responsiveness of Codex sessions and agents, underscoring the advantages of co-designed hardware and software. Jalapeño also marks the first iteration of a multi-generation AI compute platform, with broader deployment scheduled for later in 2026.

Z.ai has introduced GLM-5.3-Flash, the first natively multimodal model within the GLM-5 family, engineered to deliver frontier-level capabilities while minimizing inference costs. The architecture comprises 320 billion total parameters but activates only 18 billion, leveraging a hybrid design that merges sparse and linear attention mechanisms. This approach reduces attention computation by approximately 3× and cuts KV-cache requirements by roughly 4.4× compared to GLM-5.3, significantly improving efficiency for long-context and agentic workloads. Supporting a 1-million-token context window, the model adds native visual processing capabilities for text, images, screenshots, and code. It is particularly optimized for coding, visual understanding, and long-horizon agentic workflows, with competitive API pricing positioning it as a practical solution for high-volume, cost-sensitive applications. The model is accessible through Z.ai’s coding ecosystem and is being marketed for frequent, resource-efficient AI deployments.

Liquid AI has launched Pipette, an open-source benchmarking platform designed to answer a critical practical question: which AI model performs best on your specific device? Rather than evaluating models in isolation, Pipette assesses the entire deployment stack, including the model architecture, quantization method, runtime, hardware, context length, speed, latency, memory consumption, and output quality. Its inaugural public dataset encompasses over 1,000 configurations spanning 30+ models, quantization techniques, runtimes, devices, and context lengths, with benchmark clients available for macOS, Windows, iOS, and Android. The platform features an interactive leaderboard and dashboard, enabling developers to compare models based on real-world on-device performance rather than relying exclusively on cloud benchmarks or parameter counts. As smaller models and edge AI grow more capable, such granular evaluation is essential; a model that excels in server environments may perform markedly differently on a smartphone or laptop. Pipette aims to make these performance variations measurable, reproducible, and actionable for developers building local AI applications.

Claude AI Design enables users to generate polished product demo videos using straightforward prompts, screenshots, and website metadata, functioning as a streamlined design AI for presentations and visual storytelling without requiring advanced editing expertise. X1 serves as an AI app builder that guides developers step-by-step from concept to a publishable iPhone application, avoiding the pitfalls of single-prompt generation. Expertise AI provides a platform for creating, deploying, sharing, and monetizing AI skills. Experts can publish protected playbooks as standalone skills, allowing businesses to install them instantly or customize their own, thereby scaling institutional knowledge across organizations. ChatCut Desktop introduces an AI-powered video editor designed for collaborative human-agent workflows. Users can direct edits via natural language or integrate ChatGPT, Codex, or Claude Code, watching changes render on a fully editable timeline. MCP-Builder.ai simplifies the creation of hosted, secure MCP servers. Functioning similarly to Lovable for MCP connectors, it automatically builds, hosts, and secures connectors based on user descriptions, eliminating manual infrastructure and security configuration. Tellie Prompter 1.5 reimagines teleprompter functionality by adapting to the speaker rather than forcing rigid adherence to text. It listens to live speech, allowing pauses, improvisation, or script deviations, while version 1.5 anticipates upcoming content to maintain seamless delivery. ify integrates directly atop existing helpdesk platforms—including Freshdesk, Zendesk, Salesforce, and HubSpot—rather than requiring migration. It operates across email, chat, WhatsApp, Slack, and other channels to automate customer support workflows. JD.com has open-sourced JoyAI-Echo, a framework for generating long-form, multi-shot audio-video content up to five minutes in duration. The system maintains character and voice consistency across shots and reports a 7.5× inference speedup. Apple has expanded its silicon lineup with the M6 and M5 Ultra chips, delivering substantial gains in performance and AI compute. The M5 Ultra supports up to 512GB of unified memory, significantly enhancing the Mac Studio’s capacity to run demanding AI workloads locally. Google Cloud has launched Gemini Enterprise for Legal, integrating AI agents with enterprise data and legal-specific tools to assist with contract review, document analysis, regulatory tracking, research, and related workflows. ApodexAI’s open-source FrontierAgent empowers developers to build autonomous AI agents for complex, multi-step tasks. It supports ReAct and Agent Team modes alongside a command-line interface to streamline agent development. IBM has expanded its Granite AI model family with Granite 4.2, introducing enhanced coding and reasoning capabilities. The release underscores how synthetic training data can elevate performance on challenging software engineering tasks. Meta has recruited AI researcher Luke Metz for its Superintelligence Labs, marking another high-profile talent acquisition from OpenAI and Thinking Machines amid intensifying competition for top AI researchers.

Finally, a new study reveals that multi-agent AI systems can independently derive the correct answer yet still select a widely circulated incorrect option. Testing across 15,336 questions and 81,390 candidate pools, researchers found that combining answer frequency metrics with LLM-based judgment improved accuracy from 63.82% to 70.95%. The findings suggest that refining answer selection mechanisms may yield greater gains than simply increasing agent count or extending deliberation periods.

Read the original
OpenAI has officially unveiled its first… · Slicast