← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

Intel Details Crescent Island AI Accelerator with Larger Caches and Deeper XMX Engines

🔄 Updated 2h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Crescent Island uses Xe3P architecture, targeting inference workloads.
  • It features expanded L1 cache and general register file space per Xe Core.
  • XMX engines have a 16-deep systolic design, improving matrix processing.
  • The accelerator is a 350W air-cooled PCIe card with LPDDR5X memory.

Crescent Island's Niche in AI Acceleration

Intel revealed further architectural specifics of its Crescent Island AI accelerator, powered by the Xe3P architecture, at the Hot Chips symposium. Unlike high-power, liquid-cooled GPUs such as Nvidia's Rubin and AMD's MI455X, which target both AI training and inference, Crescent Island is positioned for lower-power, inference-first applications.

Crescent Island is an air-cooled 350W PCIe card that supports up to 480 GB of LPDDR5X memory. This design allows for deployment in standard servers without requiring specialized power and cooling infrastructure.

Architectural Enhancements

The accelerator is composed of four Xe3P slices, each containing eight Xe Cores, totaling 32 Xe Cores. Each Xe Core integrates eight Xe Vector Engines and eight XMX matrix accelerators, resulting in 256 of each resource across the chip.

The Xe3P architecture refines the cache hierarchy seen in the Xe3 graphics architecture. Each Xe Core now features 1MB of general-purpose register file space, double the 512KB found in Battlemage. Additionally, each Xe Core includes 512KB of L1 cache or shared local memory, an increase from 256KB in Battlemage. The chip also incorporates 32MB of shared L2 cache, with these expanded caches supporting the larger matrix accelerators.

Deeper XMX Engines for Inference

Xe3P's XMX engines feature a 16-deep systolic design, which allows for processing matrices in larger chunks compared to the four-deep systolic design of previous Xe2 and Xe3 architectures. This deeper systolic design enhances the capability for general matrix-multiply operations, which is a critical aspect for AI inference performance.

Intel is also focusing on broad data type support, including FP4 formats, to further optimize the chip for its inference objectives.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~12 min · 12 stories · Aug 25

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Intel provided new architectural details for its Crescent Island AI accelerator at the Hot Chips symposium, highlighting increased cache sizes and deeper XMX engines. This accelerator is designed for lower-power, inference-first AI workloads, differentiating it from high-power training-focused GPUs.