Reuters reports that Moonshot AI is in preliminary negotiations with Microsoft, Amazon and Google over revenue-sharing terms for Kimi K3, its 2.8-trillion-parameter open-weight model. But for Azure customers, the headline needs an important correction: Kimi K3 is already deployable in Microsoft Foundry through Fireworks AI, a partner-hosted offering Microsoft announced on July 28.

The proposed deal Reuters describes would therefore not be K3’s first appearance in Microsoft’s AI platform. It would be a potentially more consequential direct commercial relationship between Moonshot and Azure, AWS, and Google Cloud—one that could determine who receives inference revenue, what usage data each party sees, and whether the model receives broader first-party cloud distribution.

Reuters, citing three people familiar with the discussions, says Moonshot wants up to 30% of revenue from K3-related services on Azure, AWS, and Google Cloud. Microsoft, AWS, Google, and Moonshot either declined comment or did not respond, and Reuters says no agreements have been signed. No other outlet has independently reported the negotiating terms or timing.

Futuristic control room managing a glowing AI network linking servers and global cloud systems.Azure access exists, but it is not the deal Reuters describes​

Microsoft’s Foundry announcement is unusually clear about the present arrangement. Customers can deploy “FW Kimi K3” through Fireworks AI, while Fireworks supplies optimized inference infrastructure and Microsoft Foundry supplies the deployment experience, governance, and billing layer.

That means an enterprise already using Foundry can provision K3 without downloading the weights, assembling a GPU cluster, or operating an inference stack. Microsoft lists Data Zone pricing at $3.30 per million input tokens, $0.33 per million cached-input tokens, and $16.50 per million output tokens.

The distinction is more than branding. The current Azure option is a three-party route: Moonshot owns the model, Fireworks operates the model-serving infrastructure, and Microsoft provides the enterprise platform surface. Reuters’ reported talks concern an agreement under which Microsoft itself, as well as AWS and Google, could host or commercially distribute K3 under direct revenue-sharing terms with Moonshot.

For IT buyers, that leaves several unanswered questions. A direct Azure arrangement could potentially alter regions, support terms, model catalog placement, procurement channels, data-processing commitments, pricing, or capacity. None of those changes has been announced, and customers should not assume that a Reuters-reported negotiation changes any existing Foundry deployment.

There is also a practical billing lesson in Microsoft’s own support channels. A Foundry customer reported seeing apparently flat token charges that did not match the published K3 prices. A Microsoft moderator said the public K3 rates are differentiated by input, cached-input, and output tokens, but that the billing records required meter-level reconciliation. That does not establish a broad billing problem, but it does mean administrators testing K3 should validate actual Cost Management exports against token categories rather than relying only on a blended credit balance.


Kimi K3 is open-weight, not an easy self-hosting project​

Moonshot released K3’s weights under its own Kimi K3 License and describes the model as an open-weight, native multimodal agentic model. The company’s documentation lists 2.8 trillion total parameters, with 104 billion activated parameters per token across a mixture-of-experts architecture. It also supports text and image inputs and a context window of 1,048,576 tokens.

Those numbers explain why managed access matters. A model can be downloadable while remaining impractical for most organizations to serve at production throughput. K3 activates only 16 of 896 experts per token, reducing inference work compared with running every parameter at once, but its model footprint, expert routing, memory requirements, and interconnect needs are still far beyond what a normal Windows workstation—or even a typical single-server AI deployment—can reasonably absorb.

Moonshot’s public materials recommend vLLM, SGLang, and TokenSpeed as inference engines and provide an API compatible with OpenAI- and Anthropic-style client patterns. Its technical report describes deployment work involving expert parallelism, memory management, and specialized attention mechanisms. In other words, K3’s open weights give organizations flexibility, but the usable product for most enterprises will still be an API or managed endpoint.

Microsoft’s Foundry listing makes that reality concrete. The customer is not simply “running an open model on Azure”; it is consuming a partner-served model through Microsoft’s governance and billing framework. A direct agreement with Moonshot could change the commercial plumbing, but it would not remove the fundamental need for large-scale inference infrastructure.

AWS has published a deployment path; Google has not announced one​

Reuters framed the talks as a possible route for K3 to run on the three major U.S. clouds. The public record shows that the situation is already uneven.

AWS published a July 30 technical guide for deploying K3 on Amazon SageMaker HyperPod and Amazon EKS. That is meaningful support for organizations willing to operate K3 themselves on AWS infrastructure, but it is not the same thing as a managed K3 API in Amazon Bedrock or a revenue-sharing distribution deal with Moonshot. AWS’s guide is a deployment procedure, not a marketplace announcement.

Microsoft’s Foundry option is further along for customers seeking a managed endpoint, though Fireworks AI sits between Microsoft and Moonshot. Google Cloud has not announced a comparable K3 availability arrangement. Reuters’ report, if negotiations bear fruit, could close those product-distribution gaps—but that is prospective, not current capability.

The reporting also identifies token-use auditing and data access as unresolved issues. Those are not routine legal details. With usage-priced AI services, the token meter determines the revenue base that both Moonshot and a cloud provider are trying to divide. Data-access terms determine what telemetry a provider may collect, retain, or use to operate the service—questions enterprise security teams will care about at least as much as model benchmark scores.

The policy risk sits outside the model catalog​

The commercial discussions come after U.S. officials publicly accused Moonshot of using outputs from Anthropic’s Fable model to develop K3 and of obtaining restricted Nvidia AI chips. Reuters reported on July 22 that White House technology official Michael Kratsios made those allegations, while Treasury Secretary Scott Bessent said he was considering a trade blacklist designation and sanctions.

Moonshot has rejected the claim that K3’s performance came from distillation, telling China’s National Business Daily that the performance gains came from original changes to its underlying architecture. The Chinese embassy also called the U.S. accusations unfounded in Reuters’ July reporting.

Those are allegations and denials, not a published enforcement action. Still, they create the central business risk for Microsoft, Amazon, and Google: a cloud provider could negotiate, integrate, validate, and market a model only to face an Entity List or sanctions decision that changes whether the relationship can continue.

That uncertainty also limits what customers should infer from a cloud listing. K3 being available through Microsoft Foundry via Fireworks does not settle the legal or policy questions around Moonshot. Conversely, Treasury scrutiny alone does not mean an existing deployment is automatically prohibited. A formal government action, if one occurs, would matter far more than the current political rhetoric—and neither Microsoft nor the other cloud companies has announced a policy change affecting K3 access.


K3’s appeal is capability per dollar, not a universal benchmark win​

Moonshot says K3’s Kimi Delta Attention and Attention Residuals architecture improves scaling efficiency relative to Kimi K2. Its published technical report says the model trails the strongest proprietary systems overall, including Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol, while performing competitively across coding, agentic, reasoning, and vision evaluations.

Reuters cited Arena.ai’s web-interface-building ranking and Artificial Analysis comparisons on complex, multi-step tasks as evidence that K3 has commercial appeal. Those results are useful indicators, but they should not be turned into a blanket claim that K3 is the best model for every enterprise workload. Benchmark harnesses, tool access, reasoning settings, model versions, latency targets, and safety policies all materially affect outcomes.

For Windows administrators and developers, the sensible evaluation is narrower. Test K3 on a representative internal codebase, document corpus, or agent workflow; log token use and latency; assess the model’s output against organization-specific security and accuracy requirements; and verify where the selected partner-hosted endpoint processes and retains data. A one-million-token context window is valuable only when the model can reliably use the information and when sending that much material to an external endpoint is permitted.

The immediate consequence is straightforward: Azure customers who want Kimi K3 do not need to wait for Reuters’ reported negotiations. They can already deploy the Fireworks-hosted model through Microsoft Foundry, subject to its availability and commercial terms. What remains unresolved is whether Microsoft will turn that partner route into a direct Moonshot relationship—and whether U.S. policy will allow that business arrangement to survive long enough to matter.