Arm’s C2-Ultra, G2-Ultra NX, and CSS N4 IP
Chips and Cheese (George Cozma)Hello you fine Internet folks,
While Arm’s last announcement was about their new datacenter focused CPU, the Arm AGI CPU. Their latest set of announcements bring a much wider focus. The incoming C2-Ultra CPU and G2-Ultra NX GPU IP will see broad use across flagship phones, with latter G2 Ultra already shipping in Xiaomi’s XRING O3 chip that launched on August 24th. While the new Neoverse CSS N4 IP will see usage in the datacenter space.
Hope y’all enjoy!
C2-Ultra CPU Core IP
Starting off with a comparison to the Cortex X925, the last ARM core we have good data on, with the limited information provided to us by Arm on C2 Ultra. We see limited changes when comparing to the two generation old Cortex X925.
Across both cores you have the same overall layout featuring 10 wide decode, 8 simple ALUs, 6 lanes of FP, 3 branch ports, and same 4 load/2 store config. So for C2 Ultra to get it’s performance boost, we’re looking at more iterative changes targeting a few critical structures like the branch predictor and OoO buffers.
Arm was unfortunately very vague with how it accomplished these branch predictor improvements. We don’t know if they modified BTB sizes, return stacks, or the branch algorithm as a whole.
Likewise they mention C2-Ultra has a larger execution window compared to C1-Ultra along with improved speculation, but go into no explicit details here.
All this means that C2-Ultra spends less time waiting for data compared to C1-Ultra.
Now, looking at the claimed uplift from C1-Ultra to C2-Ultra we see that Arm is claiming a 15% peak performance uplift with an average 12% uplift in traditional benchmarks and workloads.
However, the endnotes provide some important context for these claims.
Firstly, these numbers are from FPGA simulations so actual hardware may see different uplifts. Secondly, the C2-Ultra platform’s results are estimated with an 8.5% higher clock compared to the C1-Ultra along with a larger L2 cache and a memory system providing nearly twice the bandwidth for the CPU benchmarks.
This means that C2-Ultra likely hasn’t improved the per-clock performance compared to C1-Ultra according to the numbers Arm provided.
Arm says that C2-Ultra uses 38% less power compared to C1-Ultra. However, this number does factor in node and implementation improvements so how much of this 38% decrease comes from the microarchitectural improvements is up in the air.
Arm also announced C2-Nano and C2-Pro as well, however these use the same underlying microarchitecture as the C1-Nano and C1-Pro cores.
G2-Ultra NX GPU Core IP
Moving to the GPU IP, Arm says that Mali G2-Ultra NX is “The largest re-architecting of the GPU IP in 7 generations."
The amount of math that a G2-Ultra NX shader core can do hasn’t changed. It still has 128 FMA units which equates to 256 FP32 FLOPs per clock or 512 FP16 FLOPs per clock. The maximum number of shader cores allowed in a G2-Ultra NX GPU, 24 G2-Ultra NX shader cores, also hasn’t changed.
G2-Ultra NX has increased the number of registers that a warp can access from 64 to 128 along with improving the granularity of register access so that G2-Ultra can now allocate registers at a granularity of 16 registers. And along with these improvements, the register file has also increased by 25%.
Arm has also improved the RT units in G2-Ultra NX by making the triangle structure that the RT unit works on more compact which according to Arm reduces the DRAM traffic by 13% due to an elimination in redundant data which allows more data to fit into cache.
Arm has also added Opacity Micromaps to G2-Ultra NX which brings a desktop GPU feature down into the mobile sector.
Another desktop GPU feature that Arm is integrating into G2-Ultra NX is a matrix accelerator into the Shader Core. This matrix accelerator can do up to 1,024 INT8 MACs per clock or 512 INT16 MACs per clock and it can clock up to twice the clock that the Execution Engines can clock to. However, a glaring omission is that the matrix accelerator only supports INT8 and INT16 but not formats such as lower precision like FP8 or the larger yet very popular format BF16 which also isn’t supported on the standard ALUs either.
Arm also doesn’t require every Shader Core in a G2-Ultra NX GPU to have a matrix accelerator but does require a minimum of 6 “NX cores” for the Ultra branding.
Xiaomi decided for their G2-Ultra NX implementation on their brand new XRING O3 that only half of the Shader Cores would implement the matrix accelerator.
Comparing the structure sizes of the cores with and without the matrix accelerator, a G2-Ultra NX shader core with a NX unit ends up being about 1.88 mm^2 and without the matrix unit a G2-Ultra shader core is about 1.55 mm^2 which means that the matrix unit adds about 21% to the area of a shader core.
The added matrix units have allowed Arm to introduce their own version of ML-powered upscaling called Neural Super Sampling (NSS).
Arm is also introducing their Frame Rate upscaling technology with the G2 generation.
With all of changes to the GPU IP, Arm is claiming up to 14% improvement in games and up to 24% in Ray Tracing benchmarks compared to last generation.
However, the G2-Ultra NX GPU is clocking about 11% higher compared to the G1-Ultra GPU in these comparisons which imply that these changes may not improve non-RT games much.
Neoverse CSS N4
On the server side, Arm has also announced that Neoverse N4 CSS will become available for their partners.
Sadly, this is the one slide that Arm had about Neoverse CSS N4. When asked, Arm did give out more information:
- CPU: 8–128 Neoverse N4 cores per die, the widest core-count range offered in a Neoverse CSS, with frequencies up to 3.8 GHz.
- Scaling: Supports multi-chiplet and multi-socket designs for scaling beyond a single die.
- L1 cache: 64KB instruction and 64KB data cache per core.
- L2 cache: Up to 2MB private L2 cache per core.
- System-level cache: Up to 256MB shared cache per die.
- Memory: Supports DDR5 or LPDDR6, providing flexibility across capacity, bandwidth, power and system design.
- I/O: Up to 128 lanes of PCIe Gen 6/7 and CXL 4.0.
- Chiplet connectivity: Arm chip-to-chip interconnect with support for UCIe or partner-specific PHYs.
There are still a number of questions about what core is N4 using, what is the width of the memory bus on a single die, what is the number of PCIe lanes on a single die, among others.
Conclusion
Something that I should mention is that Arm’s technical disclosure this time around is disappointing. To Arm’s credit, the company did answer questions when asked and that responsiveness is appreciated, however it is frustrating that those questions were necessary to fill gaps that should have been addressed in the presentation.
C2-Ultra is a good example of this problem, with claims about improved branch prediction and speculation giving little information of what actually changed, while the performance figures combine those core changes with higher clocks, a larger L2 cache, and more memory bandwidth. Moving to the power figures which include process and implementation improvements, we are left with the question about how much is the new microarchitecture contributing to the power decrease. These numbers are particularly frustrating for an announcement about CPU IP where comparisons at equivalent clocks and memory configurations would have been useful.
G2-Ultra NX gives more details, particularly around register allocation, the addition of matrix accelerators, and the improvements to the ray tracing units, however describing it as “The largest GPU rearchitecting in seven generations” creates an expectation of technical depth that wasn’t delivered. The matrix accelerator’s limited precision support also raises questions about its usefulness beyond the workloads Arm has chosen to target and I would have liked Arm to spend time explaining in the presentation why. And CSS N4 takes the lack of detail to an extreme with one slide with basic information about the product having to be asked.
Arm has IP and products worth talking about, but the presentation does a poor job of communicating them. The technical substance should be in the presentation from the beginning rather than something we have to assemble through follow-up questions and emails.

Generated by RSStT. The copyright belongs to the original author.