2048 NPU cores. 32 TB/s HBM4 bandwidth. 8 PFLOPS INT4. PCIe 6.0. 400W TDP. Surpasses NVIDIA Rubin GB200 and AMD MI400 across every metric at half the power.
Leads in bandwidth, compute density, power efficiency, and memory capacity.
| Specification | APEX X1 | NVIDIA Rubin GB200 | AMD MI400 |
|---|---|---|---|
| Process Node | 3nm TSMC N3E | 3nm TSMC N3 | 3nm TSMC N3 |
| AI Cores | 2048 APEX NPU | ~2000 CUDA | ~1800 CU |
| HBM Bandwidth | 32 TB/s HBM4 | 8 TB/s HBM3e | 9.8 TB/s HBM3 |
| FP8 Performance | 4 PFLOPS | 3.5 PFLOPS | 3.2 PFLOPS |
| INT4 Performance | 8 PFLOPS | 7 PFLOPS | 6 PFLOPS |
| TDP | 400W | 700W | 500W |
| Memory | 192GB HBM4 | 96GB HBM3e | 128GB HBM3 |
| Die Size | 600mm² | 814mm² | 750mm² |
| PCIe Interface | PCIe 6.0 x16 | PCIe 5.0 x16 | PCIe 5.0 x16 |
Four key subsystems working in concert for industry-leading AI performance.
2048 APEX NPUs in 64x32 grid. 16x16 systolic array per core, local SRAM, variable-sparsity engine. FP8/FP16/BF16/INT8/INT4.
8-channel HBM4, 32 TB/s aggregate. 16 Gbps per channel, 1024-bit wide. ECC, in-memory atomics, peer-to-peer.
8x8 2D mesh. 512 GB/s per router, wormhole routing, adaptive congestion avoidance.
PCIe 6.0 x16, 128 GB/s bidirectional. PAM-4 at 64 GT/s. SR-IOV, ATS, PASID, DOE.
Three breakthrough technologies powering the APEX X1 advantage.
Dynamic runtime sparsity: up to 2x throughput on attention layers without accuracy loss. Supports 1:1 to 8:1 ratios vs NVIDIA's fixed 2:4.
Hardware root-of-trust with CRYSTALS-Kyber/Dilithium. NIST FIPS 205/206 compliant. Sub-microsecond key exchange.
Integrated silicon photonics. 8 WDM channels at 100 Gbps each = 800 Gbps per fiber pair. 40x lower latency, 1/10th power.
From architecture freeze to production deployment.
Enterprise-grade AI silicon. Full sovereign deployment. Start designing today.