Ai Infrastructure

OpenAI releases self-developed AI inference chip Jalapeño: AI infrastructure competition enters a new stage

OpenAI and Broadcom jointly launched the first self-developed AI inference chip, Jalapeño, marking the extension of AI infrastructure competition from the model layer to the chip layer. This article analyzes the chip's technical features, industry impact, and implications for cloud providers and enterprises.

Event: OpenAI Releases Its First Self-Developed AI Inference Chip

On June 24, 2026, OpenAI and semiconductor giant Broadcom jointly announced the release of their first custom AI inference chip—Jalapeño. This is an application-specific integrated circuit (ASIC) designed specifically for large language model inference. OpenAI handled the underlying architecture design, Broadcom was responsible for silicon implementation and network hardware, and Canadian electronics manufacturing services provider Celestica integrated the board cards and rack systems.

OpenAI stated that Jalapeño's performance-per-watt ratio will surpass the current state-of-the-art. Engineering samples are already running multiple machine learning workloads in the lab at target production frequency and power consumption, including GPT-5.3, Codex, and Spark. The chip is planned for initial deployment by the end of 2026.

Background: Surging AI Inference Costs Drive In-House Chip Development

OpenAI previously relied mainly on an exclusive partnership with Microsoft, leasing Azure clusters composed of tens of thousands of NVIDIA GPUs for model training and hosting. However, as the user base of the GPT series expanded, inference costs rose sharply. According to industry estimates, inference costs for large language models can account for over 70% of operational expenditure. Meanwhile, market competition intensified—Google gained a significant cost advantage with its self-developed TPU, while Anthropic formed deep ties with Amazon and Google.

OpenAI pointed out that the world is moving toward a "computation-centric economy," and Jalapeño is part of its long-term full-stack infrastructure strategy. High operating costs, market competition, and supply chain pressures collectively drove the decision to launch a self-developed chip in mid-2026.

Technical Analysis: Advantages of a Dedicated Inference ASIC

  • As a dedicated ASIC, Jalapeño is extremely optimized for large language model inference scenarios. ASICs implement specific algorithms through fixed-function logic, offering higher energy efficiency compared to general-purpose GPUs. Its main features include:
  • High Performance-Per-Watt Ratio: Customized cores and memory architecture reduce energy consumption per generated token.
  • Low Latency: Optimization for inference-specific batch processing and caching mechanisms reduces response time.
  • Integration with OpenAI's Software Stack: Deeply compatible with frameworks such as PyTorch and Triton, supporting models like GPT-5.3.

The release of this chip marks a shift in AI infrastructure from "general-purpose GPU parallel computing" to "model-specific compute power." NVIDIA CEO Jensen Huang emphasized at the same day's shareholder meeting that the AI industry is divided into five layers—energy, chips and systems, infrastructure, models, and applications—and that the core task of an "AI factory" is to generate tokens, explaining why demand for computing is so strong.

Enterprise Impact Analysis: Dual Transformation of Cost and ArchitectureFor enterprises adopting cloud AI services, OpenAI's self-developed chip may bring the following impacts:

  • Cost Impact: If Jalapeño deployment succeeds, OpenAI is expected to lower inference service prices, thereby reducing the total cost of using GPT models for enterprises. However, initial deployment costs are high, and the effect remains to be seen.
  • Deployment Impact: Enterprises do not need to manage hardware themselves but should monitor changes in OpenAI's pricing strategies.
  • Operations Impact: Specialized chips may reduce dependence on NVIDIA GPUs, lowering risks from GPU supply constraints.
  • Security and Compliance: The chip may include security features like a root of trust at the chip level, benefiting enterprise data protection.

However, enterprises should be cautious about API changes due to chip version iterations and the risk of long-term vendor lock-in.

Market Competition Analysis: Chip Landscape Facing Reshaping

  • OpenAI's entry into the chip field will directly affect:
  • NVIDIA: Although still dominant in training chips, its inference chip market share may be eroded. Jensen Huang also showcased new products like Vera Rubin, indicating continued investment in full-stack solutions.
  • Google: TPUs have already demonstrated the profitability of specialized chips; Jalapeño will directly compete with them.
  • AWS, Azure, Google Cloud: To remain competitive in AI services, cloud providers need to accelerate self-developed chips or collaborate closely with third parties.
  • WiMi and other emerging AI chip companies: The reference mentions WiMi has achieved economies of scale in open-source large model computing power, but facing OpenAI's entry, small and medium players need to find differentiated positioning.

Overall, the AI chip market will shift from "one superpower with multiple strong players" (NVIDIA dominant, many followers) to "multipolar competition," with specialized inference ASICs becoming an important growth driver.

Industry Trend Observations: AI Infrastructure Moving Toward a "Chip-Making" Era

The launch of Jalapeño further confirms several long-term trends: 1. AI model companies expanding into hardware: OpenAI, Google, Anthropic, etc., are all developing their own chips to optimize costs and control technological direction. 2. ASIC rising in inference: Compared to GPUs, ASICs offer higher energy efficiency for specific inference tasks, potentially reshaping data center procurement patterns. 3. Computing becoming a core economic resource: As Jensen Huang stated, the "computing economy" drives enterprises to reassess IT investments, shifting from traditional servers to AI accelerators. 4. Combining open source with specialized computing: Companies like WiMi lower barriers through open-source models, while specialized chips provide efficient execution layers, forming a complementary relationship.

In the next five years, the investment focus of AI infrastructure will shift from general-purpose computing to specialized inference chips, liquid cooling, and sustainable energy.## CloudTechDaily Insight

  • The significance of OpenAI's release of the Jalapeño chip goes far beyond a single product launch. It marks the official transition of the AI industry from the "model race" to the "infrastructure race" stage. When model capabilities become homogenized, computational efficiency becomes the decisive factor in business competition. For enterprise CIOs and CTOs, this trend means:
  • The choice of AI supplier is no longer based solely on model performance, but also on the cost and sustainability of its underlying computing architecture.
  • Enterprise internal IT architecture must consider compatibility with multiple AI platforms to avoid being locked into a single chip ecosystem.
  • Data center planning should reserve deployment space for dedicated AI accelerators and assess power and cooling needs.

We predict that in the next three years, at least three major AI model companies will launch their own inference chips, and NVIDIA, AMD, and Intel will accelerate the release of optimized products for inference. The pricing model of the entire cloud service industry may also shift from "per GPU instance billing" to "per token or inference unit billing." Enterprises should prepare in advance and incorporate the scalability and cost predictability of AI infrastructure into their strategic planning.

(CloudTechDaily | June 25, 2026)

Reference trail · cloudtechdaily

cloudtechdaily frames this note through Cloud Platforms / Data Centers / Enterprise SaaS: dates, names and status changes still need checking. Cloud Platforms / Data Centers / Enterprise SaaS explains the local editorial angle; Source links should be opened before the summary is reused.

Source links

  1. https://www.moomoo.com/community/feed/openai-launches-first-self-developed-ai-inference-chip-boosting-nvidia-116813940457481Primary

Related articles

Back to channel