What is the NIMO AI Mini PC?
The NIMO AI Mini PC is a 2-liter desktop built around AMD’s Ryzen AI Max+ 395 “Strix Halo” APU — 16 Zen 5 cores, 32 threads, and the Radeon 8060S, currently the fastest integrated GPU you can buy — fed by 128 GB of LPDDR5X-8000 unified memory on a 256-bit bus. Launched at CES 2026, it is one of the few Strix Halo boxes you can actually order on Amazon in the US today at $2,999.99, rather than gambling on a China-direct storefront.
That framing matters. This is not a general-purpose mini PC that happens to have an NPU sticker on it. It’s a local-AI workstation in the same conversation as the NVIDIA DGX Spark — except it runs plain Windows 11, plays games, and costs less than most GB10 machines.
What it’s good for: local LLMs in 128GB of unified memory
The whole point of AMD’s Ryzen AI Halo platform is that the GPU shares one big pool of memory with the CPU. On the NIMO you can allocate up to 96 GB of the 128 GB directly to the Radeon 8060S as dedicated graphics memory (and more via shared allocation on Linux). In practice:
- 70B-class dense models (Llama 3.3 70B, Qwen 2.5 72B at Q4) fit entirely in memory — impossible on any consumer discrete GPU short of stacked RTX 5090s.
- MoE models like Llama 4 Scout and the gpt-oss-120B class are the sweet spot: huge total parameter counts, but few active parameters per token, so they run at genuinely interactive speeds.
- 7B–32B models fly, with room left over to keep several loaded at once for agent pipelines, RAG embedding, and a Stable Diffusion instance on the side.
The XDNA 2 NPU contributes ~50 TOPS (about 126 TOPS platform-wide) for Windows Studio Effects and Copilot+ features, but be clear-eyed: LLM inference on this machine runs on the GPU via ROCm, Vulkan, or LM Studio — the NPU is a bonus, not the engine.
For creator work, 40 RDNA 3.5 CUs put the 8060S in mobile RTX 4060–4070 territory: DaVinci Resolve and Premiere 4K timelines are comfortable, and 1440p gaming is a legitimate side benefit. As an office machine it’s frankly overqualified, but the 120W ceiling and small footprint make it an easy fit on any desk.
Memory bandwidth — the real ceiling for token generation
Here’s the honest part every Strix Halo buyer needs to hear: token generation speed is limited by memory bandwidth, not compute. LPDDR5X-8000 on a 256-bit bus works out to roughly 256 GB/s theoretical — enormous for an iGPU, but a fraction of the ~1.8 TB/s an RTX 5090 enjoys.
Every generated token requires reading the model’s active weights from memory, so throughput scales with bandwidth. Expect ballpark figures like single-digit tokens per second on a dense 70B at Q4, and much faster — often 30–70 tok/s — on smaller or MoE models. A discrete GPU with the same model fully in VRAM generates several times faster; the NIMO’s advantage is that 70B+ models fit at all. Notably, the $3,000–$4,000 NVIDIA DGX Spark lives with nearly the same constraint (~273 GB/s) — this is simply what unified-memory AI boxes are in 2026.
Build and connectivity
The chassis is a clean ~2-liter design (about 19 × 20.5 × 7 cm, 1.4 kg) with three performance modes scaling the APU up to 120W sustained, which matters for holding GPU clocks during long inference runs. Reviewers report the cooling handles the full 120W mode with fan noise that’s noticeable but not intrusive.
Thirteen ports cover essentially everything:
- USB4 (40 Gbps) — eGPU-capable, dock-capable
- HDMI 2.1 + DisplayPort 1.4 — up to 8K@120Hz output
- 2.5 GbE Ethernet, plus Wi-Fi 7 and Bluetooth 5.2
- A healthy spread of USB-A and USB-C for peripherals
No 10 GbE and no dual-LAN — a mild miss on a machine this network-relevant for home-lab AI serving, though USB4 adapters close the gap.
Memory, storage, and upgrades
The 128 GB of LPDDR5X is soldered — non-upgradable, non-negotiable. That’s physics, not stinginess: the 256-bit bus at 8000 MT/s requires LPDDR5X mounted millimeters from the package. Buy the memory you’ll need on day one — which is why we’d skip any smaller-RAM Strix Halo config for AI work.
Storage is properly serviceable: dual M.2 PCIe 4.0 slots, configurable from 1 TB to 8 TB. Model libraries are huge (a single Q4 70B is ~40 GB), so the second slot is genuinely useful rather than a spec-sheet flourish.
NIMO AI Mini PC vs GMKtec EVO-X2: which Strix Halo box?
Same silicon, same 128 GB ceiling. The GMKtec EVO-X2 pushes a slightly higher power limit and undercuts on street price during sales, while the NIMO counters with clean US Amazon availability, Wi-Fi 7, and a tidier port stack. The newer GMKtec EVO-X3 raises the platform’s power ceiling further — worth a look if you’re not in a hurry.
Pricing and where to buy
The 128 GB / 1 TB configuration lists at $2,999.99, with Amazon-refurbished units appearing around $2,299.99. It’s sold direct at nimopc.com (stock rotates; it was listed out of stock there at review time) and on Amazon, with larger-SSD variants also circulating. Against $2,700+ import-only rivals, the pricing is fair for a US-warrantied, in-stock unit — TechRadar’s running roundup of all 37 Ryzen AI Max+ 395 machines is a useful market sanity check.
What we’d flag
- Bandwidth is the ceiling. ~256 GB/s means dense-70B token generation is usable, not fast. If your workload is one mid-size model at maximum speed, a discrete-GPU desktop wins.
- RAM is soldered. 128 GB forever. Fine — but know it going in.
- ROCm software friction. AMD’s AI stack has improved dramatically, but expect occasional flag-hunting in llama.cpp/LM Studio versus the smoother CUDA path.
- NIMO is a young brand. The hardware is standard Strix Halo and the Amazon channel gives you return protection, but there’s no long support track record yet.
Verdict
The NIMO AI Mini PC delivers exactly what the Strix Halo platform promised: a quiet 2-liter Windows box that runs 70B-class and large MoE models locally, handles serious creator work, and doesn’t require importing hardware on faith. Its limits — soldered RAM, bandwidth-bound token rates — are platform limits shared by every unified-memory AI box up to and including the DGX Spark, not NIMO-specific flaws. If you want the most Amazon-buyable path into 128 GB local-LLM territory in mid-2026, this is an easy machine to recommend.