New IBM Z Generation: Telum II and Spyre for Advanced AI | IBM | IBM AI | AI | Turtles AI

New IBM Z Generation: Telum II and Spyre for Advanced AI
IBM unveils Telum II and Spyre, two innovative chips designed to enhance AI performance on IBM Z mainframes.

IBM is gearing up to launch a new generation of AI mainframe systems with the Telum II processor and Spyre accelerator. These innovative chips are designed to handle complex and intensive AI workloads, with significant improvements in performance and data management capabilities.

Highlights:

  • Telum II: 8 cores at 5.5 GHz, 360 MB of on-chip cache, and 768 TOPS per system.
  • Spyre Accelerator: Over 300 TOPS, 128 GB of LPDDR5 memory, and support for complex AI models.
  • Energy efficiency: 15% reduction in power consumption of the Telum II processor.
  • Availability: IBM Z AI systems available by 2025, with Spyre in technical preview.

 

IBM has recently unveiled the architectural details of its new Telum II processors and Spyre accelerators, designed for next-generation IBM Z mainframe systems specifically optimized for AI workloads. These advanced systems represent a major evolution in enterprise infrastructure, aimed at supporting complex AI models and generative applications with unprecedented scalability and efficiency.

The Telum II processor, at the core of this innovation, integrates eight high-performance cores operating at 5.5 GHz, a significant upgrade from the previous generation. Each core features 36 MB of L2 cache, with a total on-chip cache capacity reaching 360 MB, a 40% increase over the previous version. This increase is further supported by a virtual L4 cache of 2.88 GB per processor drawer, ensuring minimal latency and more efficient data management.

The true novelty lies in the integration of the AI acceleration unit directly into the Telum II chip, enabling low-latency, high-speed in-transaction AI inference, crucial for applications like fraud detection in financial transactions. This unit provides four times the computing capacity of the previous generation, with 24 TOPS per chip, 192 TOPS per drawer, and 768 TOPS per system. The new I/O Acceleration Unit DPU, integrated into the Telum II chip, improves data handling with a 50% increase in I/O density, contributing to the system’s overall scalability and energy efficiency by reducing power consumption by 70% for I/O management.

Alongside Telum II, IBM is introducing the Spyre AI accelerator, a unit designed for complex AI models and generative applications. Spyre offers over 300 TOPS of performance and features 128 GB of LPDDR5 memory, expandable up to 1 TB through eight cards working in synergy within the IBM Z mainframe. Each Spyre accelerator integrates 32 computing cores that support various data types, including INT4, INT8, FP8, and FP16, with a 75W TDP per card. These features make Spyre an ideal solution for low-latency, high-speed AI applications.

The Telum II is built using Samsung’s 5nm technology, with a die area of 600 mm² housing 43 billion transistors. This advanced design allows for significant improvements, such as a 20% increase in socket performance and a 15% reduction in overall power consumption. Additionally, the processor features new branch prediction capabilities and an increase in register sizes to 160, making it a crucial component for intensive AI workloads.

IBM expects its AI mainframe systems with Telum II processors to be available to clients by 2025, while the Spyre accelerator is currently in technical preview, with availability also expected by 2025. These developments mark a significant step in the evolution of mainframe systems, highlighting the growing role of AI in large-scale data processing and transforming business operations.

As these systems advance, it opens up a broader reflection on the role of AI in enterprise infrastructure. The ability to process massive amounts of data in real-time with reduced latency not only improves operational efficiency but can also transform how businesses address complex challenges, from security to service personalization.