Marvell’s newly announced Bravera SC6 SSD controller is aimed at one of AI inference’s most expensive choke points: keeping long-context model sessions supplied with key-value cache data after it no longer fits in GPU high-bandwidth memory. SDxCentral reports that the hyperscale-focused controller combines a PCIe 6.0 host interface with 16 NAND channels running at up to 3,600 MT/s, positioning flash as a much larger and cheaper warm storage tier behind accelerator memory. The significant detail is not that Marvell has introduced another NVMe controller. The company is trying to make SSD latency and bandwidth predictable enough that inference platforms can put a meaningful portion of an active session’s KV cache on flash without leaving costly GPUs idle while data is fetched. That is a tougher task than selling high sequential-throughput drives for conventional storage arrays, and it moves the controller into a workload-sensitive part of the AI stack.
SDxCentral’s report places the Bravera SC6 alongside Marvell’s Structera CXL memory products and its Photonic Fabric work. Those are related attempts to address the same memory-capacity problem, but they solve it at three sharply different latency and cost tiers. Treating them as one product story risks overstating what the SC6 alone can deliver.

Infographic of an AI inference GPU cluster with HBM, CXL DRAM, and NVMe SSD memory tiers.Bravera SC6 targets the cache that grows with every token​

A generative AI model’s KV cache stores the intermediate attention data created as it processes a prompt and generates a response. Keeping that history available avoids recalculating it, which can improve responsiveness for long conversations, retrieval-heavy queries, and multi-turn agent workloads. The catch is that the cache expands with model size, context length, batch size, and concurrent users, competing directly with model weights and active compute for scarce GPU memory.
Offloading is therefore already a common systems strategy: retain the data needed immediately in GPU memory, use host DRAM for a larger near tier, and place colder or less frequently reused cache segments elsewhere. Flash is attractive because it offers capacity at a fraction of the cost and power of HBM or server DRAM. It is also persistent, although persistence is generally a secondary benefit for ephemeral inference cache data.
But SSD offload does not make the memory problem disappear. Recent academic work on KV-cache offloading has found that naïvely moving cache pages to SSDs becomes bandwidth-bound, with PCIe transfer time and device-level access patterns capable of dominating inference latency. Other work has shown that the gains come from software that coordinates placement, prefetching, parallel I/O, GPU DMA, and computation rather than merely attaching a faster drive.
That distinction is central to the Bravera SC6 announcement. A controller can raise the ceiling for an SSD’s bandwidth, random IOPS, NAND concurrency, and latency consistency, but it cannot decide which KV blocks the inference engine should preserve, fetch, or recompute. The storage firmware, host driver path, model-serving stack, and scheduler have to work together. Without that stack-level coordination, a PCIe 6.0 SSD can still turn GPU memory pressure into I/O stalls.
Marvell says the controller can support multigigabyte-per-second throughput and millions of random IOPS. Those are plausible categories of performance for an advanced hyperscale flash controller, but the announcement as described by SDxCentral does not provide the numbers administrators and SSD makers need for a serious comparison: sequential read and write rates, random read and write IOPS at defined queue depths, tail latency, power draw, capacity targets, endurance ratings, or the NAND type used in a finished drive.
Those omissions are material. KV-cache offload is not a read-only workload. Session churn, cache eviction, prompt reuse, and checkpointing can create a demanding mixed-I/O pattern. A controller’s peak random-read figure says little about whether an SSD will retain latency consistency under sustained writes, garbage collection, error correction, and multi-tenant traffic.

PCIe 6.0 raises the transport ceiling, not the delivered result​

The SC6’s PCIe 6.0 interface matters because PCIe 6.0 doubles PCIe 5.0’s signaling rate to 64 GT/s. PCI-SIG specifies up to 256 GB/s of bidirectional bandwidth for a full x16 PCIe 6.0 link, using PAM4 signaling, fixed-size flow-control units, and forward error correction.
An SSD controller does not receive all of that x16 bandwidth. Enterprise NVMe drives normally use much narrower links, and the practical limit is shaped by the drive’s lane count, platform topology, CPU root complex, switch configuration, firmware, NAND parallelism, and thermal envelope. The 3,600 MT/s figure cited for the SC6 is a per-channel NAND-interface speed, not a guarantee of host-visible throughput.
That is why the 16 NAND channels are consequential. A controller can only turn a fast host interface into useful throughput if it has enough NAND dies, channels, and internal parallelism behind it. For inference-cache SSDs, the objective is less about one benchmark number than sustained access to many small, scattered objects without unpredictable tail-latency spikes.
Marvell has previous form here. Its 2021 Bravera SC5 line was an early PCIe 5.0 hyperscale SSD-controller family with either eight or 16 NAND channels, and Marvell said then that its top configuration could enable up to 14 GB/s and 2 million random-read IOPS. The SC6 is the generational successor in branding and interface speed, but Marvell has not publicly supplied a comparable top-line performance figure for the new part in the material cited by SDxCentral.
That makes it premature to call the SC6 a finished performance leader. It is a controller, not a retail drive or a complete server-storage appliance. SSD manufacturers and cloud operators will determine the NAND population, overprovisioning, firmware behavior, cooling, form factor, endurance profile, and ultimately the performance a deployed product exposes.

The controller, CXL expansion, and Photonic Fabric are three different tiers​

Marvell is pitching three ways to push memory capacity outward from the GPU:
  • Bravera SC6 places data on local NVMe flash, offering the greatest capacity per dollar but also the highest access penalty of the three approaches.
  • Structera X memory-expansion controllers attach DRAM over CXL, while Structera A products place compute near expanded memory. Marvell’s current Structera material describes CXL 2.0 memory expansion with up to 200 GB/s of memory bandwidth and multi-terabyte memory capacity in certain configurations.
  • Photonic Fabric is Marvell’s more ambitious disaggregated-memory concept: optical connectivity and CXL-based sharing intended to make memory available across multiple hosts and racks rather than tied to one server.
These tiers can work together. A large inference cluster might keep the hottest KV pages in HBM, place a wider active working set in CXL-attached DRAM, use a pooled optical-memory tier where the architecture supports it, and retain colder data on NVMe SSDs. The point is to avoid buying enough HBM and local DRAM to cover every worst-case context window on every accelerator.
They are not interchangeable, however. CXL memory expansion is still DRAM-backed and is designed for far lower-latency access than SSD flash. A photonic shared-memory fabric adds a networking and orchestration challenge. Bravera SC6 is a storage controller, and its success depends on whether inference software can use flash selectively enough to keep its latency from leaking into token generation.
For Windows Server administrators, there is no announced Windows-specific deployment path, driver package, supported model-serving framework, or validated server list tied to the SC6. Standard NVMe support alone would not create a KV-cache tier. The relevant work would sit above the OS in the AI serving platform and beneath it in SSD firmware and fleet management. This is a hyperscaler component announcement, not a storage upgrade that enterprises can apply to existing GPU servers.

Marvell’s published Photonic Fabric capacity claim needs clarification​

One figure in the SDxCentral report deserves particular scrutiny. The report says Marvell’s multi-rack Photonic Fabric architecture could provide “up to 32 terabits” of warm KV-cache offload across XPUs and racks up to 50 meters apart. Thirty-two terabits equals 4 terabytes.
Marvell’s own April Photonic Fabric material, meanwhile, describes a pod-scale appliance that can dynamically share up to 32 TB across as many as 16 servers. A July research paper by Marvell engineers describes a Photonic-CXL memory appliance with 32 TB of shared memory across 16 hosts. Those are eight times the capacity implied by 32 terabits.
The figures may describe different configurations or different measures: 32 terabits of a designated warm-cache allocation versus 32 TB of total addressable shared memory. But neither the SDxCentral report nor the readily available Marvell material establishes that distinction. Until Marvell publishes a product brief defining the topology, usable capacity, bandwidth per host, and the relationship between the 4 TB and 32 TB figures, customers should not convert the higher figure into a deployment assumption.
The same caution applies to the claimed potential for three-times higher token throughput within an existing data-center footprint. That is an architecture-level projection tied to a proposed disaggregated-memory design, not a measured result attributable to the Bravera SC6 controller. It should not be read as a promise that replacing SSD controllers will triple inference throughput.
Bravera SC6’s near-term importance will be clearer when Marvell or an SSD partner identifies sampling status, shipping drives, supported form factors, endurance specifications, and benchmark methodology for real KV-cache workloads. Until then, the announcement establishes where Marvell wants flash to sit in AI infrastructure: below CXL memory in latency, far above bulk storage in responsiveness, and under software control as a capacity valve for GPU-bound inference clusters.

References​

  1. Primary source: SDxCentral
    Published: 2026-08-05T00:00:00+00:00
  2. Related coverage: marvell.com
  3. Related coverage: marvell.com
  4. Related coverage: cn.marvell.com
  5. Related coverage: hyperframeresearch.com
  6. Related coverage: chinaainews.org
  7. Related coverage: linkedin.com
  8. Related coverage: nand-research.com
  9. Related coverage: topcpu.net
  10. Related coverage: d1io3yog0oux5.cloudfront.net
  11. Related coverage: d1io3yog0oux5.cloudfront.net
  12. Related coverage: techradar.com
  13. Related coverage: tomshardware.com
  14. Related coverage: investor.marvell.com
  15. Related coverage: hpcwire.com
  16. Related coverage: aithority.com
  17. Related coverage: nextmsc.com
  18. Related coverage: investor.marvell.com