Fujitsu's 144-Core MONAKA Arrives in November, Built for AI Inference Without a GPU

Fujitsu's 144-Core MONAKA Arrives in November, Built for AI Inference Without a GPU

Fujitsu has formally announced MONAKA, the Armv9 server CPU that has been visible on slides for two years. Availability begins in November, with broader European availability in 2027.

The specifications

  • 144 Armv9 cores
  • 2nm process for the cores, 5nm for cache and I/O, combined in a 3D-stacked package
  • 3.8GHz maximum operating frequency
  • 8800MT/s memory transfer speed
  • 12 RDIMMs per socket
  • PCIe Gen6

Server configurations announced:

Form factorCoolingCore clock
1U rackmountAir2.1GHz
1U rackmountLiquid2.9GHz
2U rackmountAir2.1GHz

The gap between the 3.8GHz peak and the 2.1GHz sustained figure in an air-cooled 1U is the honest number. 144 cores in one rack unit is a thermal problem before it is a compute one, and the liquid-cooled option buying 800MHz is a direct measure of how much performance is sitting behind cooling capacity.

Mixing process nodes within one package is the interesting engineering choice. Cores get the expensive 2nm process where density and power efficiency pay off most; cache and I/O get 5nm, where the newer node would add cost without proportionate benefit. It is a sensible answer to leading-edge wafers being extremely expensive.

The AI pitch, and how to read it

Fujitsu’s framing:

“The CPU provides hardware acceleration for matrix operations, which are central to AI inference, using dedicated instructions. Combined with SVE2 vector operations and software optimization, it achieves twice the throughput for AI inference compared to other CPUs. This enables practical AI inference performance from the CPU alone, offering the flexibility to deploy at the necessary scale even in locations with power and cooling constraints.”

Treat “twice the throughput compared to other CPUs” as a vendor claim until somebody independent measures it. The comparison base is unspecified.

The strategic argument underneath it is more solid than the number. Inference on CPUs is a real deployment pattern, not a consolation prize. Not every workload needs a GPU, GPUs are supply-constrained and power-hungry, and plenty of sites cannot add the power and cooling that a GPU rack demands. A CPU that does inference acceptably well is deployable where a GPU is not.

The “sovereign AI infrastructure” positioning is aimed at governments and institutions that want compute inside their own jurisdiction on hardware they can account for. That is a growing market driven by policy rather than by benchmarks.

The detail that says this is real

Compiler support went upstream in GCC 15.

That matters more than a press release. Getting a target upstreamed into GCC means engineers did the work well in advance, it survived review, and any distribution building with a current toolchain can produce optimised binaries on day one. Hardware that shows up without upstream toolchain support tends to stay in evaluation labs.

# once you have access to one
gcc -mcpu=native -Q --help=target | head
lscpu | grep -E 'Model name|Core|MHz'

Why 144 cores is topical

A machine with this many cores lands in the middle of an active kernel discussion. Scheduler patches proposed this month address the fact that Linux’s idle-CPU selection has a scan-order bias that wastes capacity precisely on high-core-count systems, and were tested on a 160-core Ampere Altra.

Arm in the datacentre is no longer novel. Ampere, AWS Graviton and now Fujitsu are all shipping high-core-count Arm server parts, and the software ecosystem has largely caught up, which our Linux on Arm servers guide covers for anyone evaluating the move.

For hosting specifically, Arm instances are already cheaper per core at most providers, and our VPS comparisons note where they are available. Whether MONAKA reaches that market or stays in sovereign and HPC deployments is the open question, and November will start to answer it.

Background reading

Explainers for the concepts behind this story.