IT Trends & Tips for Users

ARM Neoverse CSS-N4: Up to 128 Cores per Chiplet for Server CPUs

Sep 8, 2026 3 min read
All articles

At its Tech Days event, ARM unveiled new CPU and GPU cores for smartphones alongside a new core generation for servers. For server CPU cores, ARM is taking an unusual approach this time: the fourth generation initially only covers the N-cores, designed for integration density and cost efficiency, while the performance-optimized V-cores go without a successor for now.

Smaller process node, more cores per die

With Neoverse CSS N4, ARM is moving to a 3nm process instead of 5nm. Combined with changes to the interconnect network, that enables 8 to 128 cores per die. Thanks to coherent Arm Chi C2C (chip-to-chip) interconnect, at least two chiplets with CPU cores are possible per socket. As a result, ARM promises double the performance per socket compared to its predecessor, alongside 25 percent higher efficiency. Additional customer-specific chips can be connected via UCIe (Universal Chiplet Interconnect Express).

With Neoverse N4, ARM promises the most extensively configurable compute subsystem for servers yet, while also letting customers get to their own chip faster than ever. The extensive configuration options show up in memory support, for instance: ARM highlights LPDDR6, but DDR5 with support for multi-rank DIMMs at up to 12,000 MT/s is also possible. Memory bandwidth is meant to increase by 75 percent as a result.

Higher clock speeds, more PCIe lanes

ARM is still holding back details on changes to the cores themselves. They're based on version 9.3 of the ARM architecture, and the support page for Neoverse CSS N4 mentions clock speeds of up to 3.8 GHz. For L1 cache, ARM now appears to fix 64 KB each for instructions and data, while L2 cache size remains configurable, capped at 2 MB with error protection, as with Neoverse N3. On PCIe, ARM upgrades to Gen7, though Gen6 remains an option. Up to 128 lanes can be integrated, and the new core also supports Compute Express Link version 4.0.

ARM has also upgraded the Coherent Mesh Network, which connects the CPU cluster to memory controllers, other devices, or chiplets. A 16-by-16 mesh is now the maximum, and ARM says it has doubled bisection bandwidth.

Conflicting positioning for agentic AI

Overall, though, ARM's new Neoverse generation also raises some questions. ARM states that Neoverse N4 was developed for agentic AI, yet its own presentation slides assign such applications to the performance-optimized V-cores, which have no successor for now. How that contradiction resolves will likely only become clear once concrete server products based on Neoverse N4 reach the market.