Desktop workstation with a compact computer, AI brain visualization, keyboard, and futuristic office decor.
Honor has put Nvidia’s RTX Spark platform into a roughly 2-liter Windows 11 workstation, the Tiangong AXB35 Ultra, with 128GB of unified memory and a 2TB NVMe SSD. The real attraction is not the chassis size or the “Mac Studio rival” framing: it is the prospect of running memory-hungry local AI workloads on a desk without building a conventional multi-GPU tower.

Notebookcheck first detailed the system’s appearance, while Honor’s Tiangong product site now lists the AXB35 Ultra alongside Intel and AMD alternatives. Chinese outlet IT Home and VideoCardz independently confirm the family, shared 70 × 192 × 203 mm enclosure, 1.5 kg weight, and the Ultra model’s RTX Spark configuration. Honor has not published a retail price, shipment date, or a country-by-country availability list; its own site asks prospective buyers to contact the company for configuration advice.

That makes the AXB35 a product announcement rather than a purchasing option today. But it is also one of the first signs that Nvidia’s Windows-on-Arm RTX Spark push is moving beyond laptop announcements and into compact systems intended for on-premises enterprise AI.

The 128GB figure is the meaningful specification​

RTX Spark combines an Arm-based Nvidia Grace CPU with a Blackwell RTX GPU on a shared-memory design. Nvidia specifies the highest desktop-class configuration as a 20-core Arm processor, up to 6,144 CUDA cores, and up to 128GB of LPDDR5X unified memory. Its headline performance figure is up to 1 PFLOP at sparse FP4 precision.

For Windows and IT buyers, the 128GB pool deserves more attention than the petaflop claim. A conventional desktop with a discrete GPU divides available memory between system RAM and VRAM. The AXB35 Ultra’s unified-memory design lets the CPU and GPU work from one coherent pool, which is useful when a local model does not fit into the 16GB, 24GB, or even 32GB of VRAM common in enthusiast graphics cards.

There is an important limitation: 128GB of unified memory is not 128GB reserved solely for a model. Windows, background services, applications, runtime overhead, model metadata, prompt context, and the operating system’s own allocations all consume part of it. Memory bandwidth also remains central to how quickly a model generates output. A model that technically loads is not automatically practical for interactive use.

Nvidia’s own DGX Spark documentation gives a more restrained reference point than Honor’s marketing. Nvidia rates the DGX Spark, which uses the related Grace Blackwell GB10 platform, for models up to 200 billion parameters. Nvidia’s May RTX Spark announcement for Windows PCs used a different practical example: a 120-billion-parameter LLM with up to one million tokens of context. Both figures depend heavily on quantization and workload conditions, but they show why Honor’s claim of support for deployment of models “up to 300B” should be read as a configuration target rather than a promise of comfortable local performance.

Honor does acknowledge that qualification on its product page, saying usable model scale changes with quantization, context length, concurrency, and software. That caveat is doing considerable work. A 300B model reduced to low-bit precision may be possible to place in memory under constrained conditions; serving it to multiple users, retaining a large context window, connecting retrieval data, and maintaining responsive generation are separate capacity questions.

Windows 11 support comes with an Arm transition​

Honor lists Windows 11 and Ubuntu support across the AXB35 range. On the Ultra, that is significant because RTX Spark is an Arm-based platform rather than another x86 PC with an Nvidia add-in card. Nvidia’s Windows-on-Arm porting documentation describes RTX Spark as an Arm SoC, and Nvidia has explicitly positioned it with Microsoft as a foundation for Windows-native local agents.

For organizations that have standardized on Windows endpoints, that can lower the barrier to experimenting with local inference, retrieval-augmented generation, and agent workflows. The AXB35 Ultra offers native CUDA and RTX acceleration rather than asking developers to move their project to an Apple Silicon or Linux-only machine. Nvidia’s current software documentation also lists Windows on Arm support for RTX Spark in its PyNvVideoCodec package, a small but concrete indication that the software stack is being adapted rather than merely emulated.

Still, buyers should treat the machine as an Arm development system, not assume it will behave exactly like an x64 workstation. Native Arm64 applications and tools should be preferred. Windows emulation can help with ordinary desktop software, but it is a poor planning assumption for a production AI toolchain, drivers, plug-ins, custom Python environments, low-level dependencies, or legacy management agents.

The Ubuntu option could be the more straightforward route for teams already using containers, CUDA-focused frameworks, and Linux-based inference stacks. Honor has not said whether the system will ship with Windows 11 Pro, whether Ubuntu is factory-installed or self-installed, which Windows edition and Arm driver image it will validate, or how firmware and driver updates will be delivered. Those omissions matter more to an enterprise buyer than the enclosure volume.

Honor is selling a family, not one fixed hardware platform​

The Tiangong AXB35 name covers three materially different computers in the same case. The Standard model uses Intel’s Core Ultra X7 358H with 64GB of memory and a 512GB SSD. The Pro uses AMD’s Ryzen AI Max+ 395 with 128GB of memory and a 1TB SSD. The Ultra uses RTX Spark, 128GB of memory, and a 2TB SSD.

Honor rates the Intel system at up to 180 TOPS and the AMD system at up to 126 TOPS, while the Nvidia system carries the 1 PFLOP FP4 rating. Those numbers are not a clean cross-platform comparison. TOPS and FP4 petaflops describe different precision formats and acceleration paths, while local-AI performance depends on the model, runtime, memory bandwidth, CPU contribution, prompt processing, and whether the software is optimized for the relevant NPU or GPU.

The configurations are better understood by workload class. The Intel Standard model is positioned for knowledge-base work, office agents, and quantized models around 35B parameters. The AMD Pro and Nvidia Ultra are aimed at larger models and multi-agent use. The GPU-enabled CUDA stack gives the Ultra a potentially broader path for developers using existing Nvidia-oriented AI tooling, while the Ryzen AI Max+ platform may be attractive where x86 compatibility is non-negotiable.

Honor’s official page also labels the 128GB memory and 2TB SSD in the Ultra as project configurations. That wording suggests these are not necessarily shelf-ready, globally identical SKUs. It leaves open questions about component substitutions, regional specifications, storage options, memory tiers, warranties, and support contracts.

The network and storage layout fits a small on-premises node​

The shared chassis specification is unusually enterprise-minded for a mini PC: two 10GbE RJ-45 ports, Wi-Fi 7, two USB4 ports, four USB-A ports, HDMI 2.1, DisplayPort 1.4, and three M.2 expansion slots. It also uses an external 240W power supply and offers 55W, 85W, and 132W operating modes.

Dual 10GbE is more consequential than the display outputs. A compact inference appliance can sit near a file server, a workstation cluster, a lab network, or a segregated internal service without immediately becoming network-bound when several users pull documents or send requests. Three M.2 slots also offer room to separate boot media, local model weights, and a retrieval corpus, although Honor has not published lane allocation, supported drive capacities, RAID behavior, or whether all three slots maintain full performance simultaneously.

The 132W performance mode also undercuts the casual assumption that this is simply a low-power mini PC. Nvidia documents a 140W thermal design power for the comparable GB10 system-on-chip in DGX Spark hardware, before the rest of the platform is considered. Sustained performance, fan behavior, and thermals will need independent testing, especially in the AXB35’s narrow vertical enclosure. No outlet has yet published benchmarks, acoustic measurements, or teardown evidence for Honor’s implementation.

Enterprise positioning leaves the consumer case unresolved​

Honor’s presentation places Tiangong under a broader industry-solutions brand that includes model deployment, inference acceleration, knowledge-base, agent, and operations-management software. That context explains the company’s emphasis on local enterprise models and multi-agent workflows, as well as its contact-sales approach.

It also means Windows enthusiasts should resist treating the AXB35 Ultra as a confirmed Mac Studio replacement for ordinary desktop work. There is no announced price, no availability date, no confirmed retail channel, and no published Windows image or support matrix. A system with two 10GbE ports and 128GB of soldered unified memory could be compelling if priced against high-end workstations, but it could also land closer to specialist AI appliance pricing.

Honor has shown the outline of a useful machine: a compact node with enough shared memory to make local AI experiments that are impractical on mainstream PCs more feasible. The next meaningful facts are less glamorous than 1 PFLOP claims—price, regions, Windows-on-Arm validation, sustained benchmarks, and the exact portion of that 128GB memory that real local workloads can use.