Apple’s new Mac Studio is now available to pre-order with M5 Max and M5 Ultra options, but the headline number needs a correction: the company’s claimed 4.3x AI-performance increase applies to the M5 Ultra against the M3 Ultra, not to every new Mac Studio configuration. The M5 Max model has its own, lower vendor benchmark claim—up to 3.9x faster LLM prompt processing than the prior M4 Max Mac Studio in Apple’s LM Studio test.

Apple announced the refresh on August 25, with U.S. availability scheduled for September 22. The entry M5 Max configuration starts at $2,499, while the M5 Ultra model starts at $5,499, according to 9to5Mac and Macworld. That means this is a workstation-class update rather than a broadly accessible desktop upgrade—and the most consequential changes are memory capacity, storage I/O, and multi-machine AI inference rather than the CPU core count alone.

Geeky Gadgets’ report correctly identifies the M5 Max, M5 Ultra, Thunderbolt 5, Wi-Fi 7, Bluetooth 6, and a 512GB unified-memory ceiling as the major changes. It also leaves out an important availability limitation: Apple says the 512GB M5 Ultra Mac Studio will not arrive with the September 22 launch wave. That top-memory configuration is due in late October.

For developers and workstation buyers weighing a Mac against a Windows desktop, the new Studio is less a generic “fast computer” than Apple’s most aggressive effort yet to make high-memory local AI work practical on a compact system.

A futuristic AI workstation displays data visualizations beside stacked computers and glowing neural networks.M5 Ultra makes memory capacity the real differentiator​

The M5 Max model has an 18-core CPU, up to a 40-core GPU, up to 128GB of unified memory, and up to 614GB/s of memory bandwidth. The M5 Ultra raises those ceilings to a 36-core CPU, up to an 80-core GPU, 512GB of unified memory, and 1.2TB/s of bandwidth.

On paper, those figures resemble a straightforward doubling. Apple’s design is more specific: the M5 Ultra combines two dual-die M5 Max packages through its next-generation UltraFusion interconnect, creating a four-die processor that macOS can address as one system. Apple says the inter-die connection exceeds 4.4TB/s.

The important practical distinction is unified memory. A Mac Studio configured with 512GB is not offering 512GB of conventional system RAM plus a separate GPU with its own VRAM. CPU and GPU workloads draw from the same large memory pool. For local inference, that can reduce one of the persistent constraints on conventional GPU workstations: fitting a model and its working data entirely into available accelerator memory.

That does not make the Mac Studio an automatic replacement for a multi-GPU Windows or Linux server. CUDA-first tooling, NVIDIA-specific acceleration, expandable PCIe GPU configurations, and established data-center frameworks remain major reasons to use an x86 workstation or server. But for developers who specifically need a quiet desktop able to load very large models locally, Apple’s 512GB configuration is aimed at a real gap between a single professional GPU workstation and a rack-mounted accelerator server.

Apple says its top M5 Ultra can run large language models with hundreds of billions of parameters locally. That is a capacity claim, not a universal performance guarantee. Whether a particular model is usable depends on quantization, context length, framework support, token speed, and how much memory is unavailable to the workload. Buyers should treat Apple’s headline benchmark multipliers as vendor-supplied results tied to selected applications and test conditions, not as independently reproduced results across every AI stack.

Apple’s AI claims are model-specific, not a blanket speed promise​

The new chips integrate Neural Accelerators into each GPU core, a change Apple is emphasizing for matrix-heavy AI work. For the M5 Max, Apple claims up to 3.9x faster LLM prompt processing in LM Studio compared with an M4 Max Mac Studio, and up to 3.5x faster text-to-image performance.

For M5 Ultra, Apple claims up to 4.3x higher peak GPU compute for AI than M3 Ultra. In individual workload examples, Apple cites up to 4x faster LM Studio prompt processing and up to 4.3x faster text-to-image generation against the M3 Ultra. MacRumors and 9to5Mac both reported the same broad performance positioning from Apple’s announcement.

The distinction between prompt processing and token generation matters. Prompt ingestion evaluates the supplied context; token generation is the serial work of producing a response. A machine can post an impressive prompt-processing result yet deliver a less dramatic improvement in the interactive generation rate users notice in a chatbot window. Apple’s public material emphasizes prompt processing, image generation, and selected creative applications; it does not establish a universal win over NVIDIA-based Windows systems in all inference or training workloads.

Apple is also pitching the Studio for training and fine-tuning through MLX, its Apple-silicon-optimized open-source machine-learning framework, and through the newer Core AI framework. That gives macOS developers a more coherent first-party path for local model development than previous Mac Studio generations had. It also narrows the audience: teams standardized on PyTorch extensions, CUDA kernels, TensorRT, or Windows-centric desktop management tools should verify their exact pipeline before treating the Mac Studio as a drop-in workstation replacement.

Thunderbolt 5 clustering has a meaningful limitation​

The most unusual new capability is Apple’s support for clustering multiple Mac Studios through Thunderbolt 5 and RDMA, or remote direct memory access. Apple says this can create a shared memory pool spanning systems, and that a cluster of four Mac Studios can provide up to three times the AI-inference performance of one system.

This could appeal to small research groups and creative teams that want more local capacity without immediately moving a workload to a cloud GPU provider. A four-node cluster using 512GB machines would present a potentially enormous aggregate memory resource for compatible software.

But Apple’s own “up to 3x” figure is also the caution. Four systems do not turn into a four-times-faster inference node. Distributed AI workloads pay for coordination and data movement across the interconnect, particularly where model layers or requests cross machine boundaries. The shared-memory approach may simplify certain deployments, but it does not erase scaling overhead.

For enterprise IT, the setup also raises ordinary operational questions Apple’s announcement does not answer: how the cluster is monitored, whether node replacement affects model deployment, what administrative tooling is supported at scale, and which frameworks can use the capability without custom engineering. The hardware arrives before those operational details are fully clear.

PCIe Gen 6 storage and pro-video features broaden the workstation pitch​

The new Mac Studio also adopts a PCIe Gen 6 SSD architecture that Apple says can deliver up to twice the storage performance of the previous generation. It adds Thunderbolt 5 ports with bandwidth up to 120Gb/s, Wi-Fi 7, Bluetooth 6, support for as many as eight displays, and genlock over USB-C.

The storage change may be more useful to daily workstation users than the flashier AI claims. Large local model files, high-resolution source footage, cache-heavy video workflows, and datasets can all become I/O-bound. Faster internal storage will help only where the workload is reading and writing locally; it does nothing to improve a bottlenecked NAS, external drive, or network share.

Genlock is similarly specialized but material. It allows compatible cameras and displays to synchronize timing, a requirement in certain multicamera virtual-production, broadcast, and capture workflows. This is one of several signals that Apple is pursuing professional production environments rather than merely selling a compact desktop with a bigger benchmark number.

The M5 Ultra also doubles the Media Engine’s encode and decode blocks compared with M5 Max. Apple says it can support up to 33 concurrent 8K ProRes 422 streams. That is a credible reason for high-end post-production shops to consider the Ultra, but it should not be mistaken for a general-purpose graphics comparison. ProRes acceleration benefits workflows built around Apple’s preferred media formats; Windows shops centered on other codecs, GPU renderers, or NVIDIA-accelerated applications need application-specific testing.

The launch is a bigger upgrade for M3 Ultra owners than M4 Max buyers​

The product line Apple is replacing was uneven: the previous Mac Studio paired an M4 Max option with an older M3 Ultra option. The M5 Ultra therefore supersedes a chip that was already a generation behind the M4 Max in some architectural respects. That makes the M5 Ultra’s gains substantial, but it also makes comparisons important.

M5 Max buyers are moving from M4 Max, and Apple’s own M5 Max results are generally framed against that immediate predecessor. M5 Ultra buyers are moving from M3 Ultra, which explains why the Ultra’s headline comparisons can look more dramatic. The right comparison is between the model a buyer owns and the exact new configuration they are considering, not between the lowest-end prior Studio and the highest-end M5 Ultra.

The launch pricing reinforces that point. The $2,499 M5 Max entry price is unchanged from the M4 Max Mac Studio’s starting price, according to 9to5Mac. The $5,499 M5 Ultra entry price is $200 higher than the M3 Ultra model it replaces. Memory, storage, and higher-tier configurations will matter far more than that base-price difference, especially for AI buyers who need the Ultra’s larger unified-memory options.

Apple’s sustainability figures—35% recycled material overall, 100% recycled aluminum enclosure, 100% recycled rare earth elements in magnets, and 40% renewable electricity in manufacturing—are company disclosures, not independent lifecycle comparisons with competing workstations. They are worth noting, but they do not change the purchasing calculus nearly as much as software compatibility, memory capacity, service arrangements, and workload performance.

The immediate consequence is straightforward: the M5 Max Mac Studio is a faster compact professional Mac, while the M5 Ultra version is Apple’s bid for local AI and media work that previously pushed buyers toward larger, louder, more expandable systems. Buyers needing 512GB of unified memory, however, should not plan around September 22; Apple has placed that configuration on a separate late-October timetable.