A man consults a friendly robot assistant displaying code and analytics in a modern office.
Microsoft has published a primer on Agent Experience, or AX. It is a discipline for testing and improving how AI coding agents find and use your SDK, API or CLI. The post comes from Waldek Mastykarz, a Principal Developer Advocate at Microsoft who works on AI coding agents. It matters beyond Redmond. If you ship anything developers touch, an AI agent now sits between you and them. Whatever it produces becomes the verdict on your product.

What AX actually is​

AX is the experience an AI agent has when it discovers, chooses and uses your technology. Practicing it means systematically improving that experience and then measuring whether the changes helped.

The term comes from Netlify. Netlify's own page says AX doesn't replace DX, it extends it. That quote is attributed to Zeno Rocha of Resend. Several secondary sources credit Netlify CEO Mathias Biilmann with coining the phrase. One says he did so in a January 2025 post, "Introducing AX," defining AX as "the holistic experience AI agents will have as the user of a product or platform." Netlify's page confirms the 2025 date but not the month, so treat "January" as secondary-sourced.

Biilmann has since described four areas of AX. They are Access, Context, Tools and Orchestration.

The Microsoft post makes the practical case. Developers ask agents to build with your product and then judge it by what comes back. If the agent picks a competitor, calls a deprecated API or reports success on a broken build, your product looks broken, however good your docs are for humans.

The two questions: propensity and efficacy​

Microsoft boils AX measurement down to two questions.

  • Propensity: does the agent find and choose your technology? Test it with a task that doesn't name your product, such as "add email sending to this app."
  • Efficacy: once the agent uses your technology, does it use it correctly? Name the product in the task and check whether it calls the current API or a deprecated one.

Some evaluation tools use other labels, such as discovery and selection for propensity, or quality for efficacy. Both questions are scored on what each run produced and on what it cost.

Cost is part of the score​

Microsoft's example is a warning about token pricing. The post says Claude Sonnet 5 launched with 33% lower per-token pricing than Sonnet 4.6. Yet on SharePoint Framework (SPFx) upgrades in GitHub Copilot Chat, it reportedly cost 3.7x more per run, across 3 scenarios and 15 runs per model.

I couldn't independently verify those specific figures. They are Microsoft's reported results for its tested setup, not a general statement about model pricing. The principle holds regardless: a cheaper token doesn't guarantee a cheaper finished task. Task cost should always be reported next to quality.

Where you can influence an agent, and where you can't​

Microsoft splits the stack into three layers:

  • The model. You can't change it.
  • The harness. This is the agent app that runs the model, such as GitHub Copilot Chat. You can only influence it through its vendor.
  • Everything the agent reads and calls. This part is yours.

The part that is yours includes:

  • Docs. A model that learned your docs in training won't know about updates until the next model ships. An agent that reads them live sees changes on its next run.
  • Agent extensions. These are skills, MCP servers, instruction files and custom agents. They only help developers who install them.
  • Interfaces. These are APIs, SDKs, CLIs, error messages and scaffolders. Agents trust what these return. Microsoft cites a case where an unpinned scaffolder generated a project on a July 2020 version and the agent counted it as a success. The post doesn't name the scaffolder, so treat this as an attributed anecdote.

The SPFx case study​

The overview's worked example is upgrading SPFx projects. Microsoft's separate Microsoft 365 Developer Blog write-up gives more detail. It describes asking GitHub Copilot Chat in Visual Studio Code on Windows, using Claude Sonnet 4.6, to upgrade a project from SPFx 1.21.1 to 1.22.2. The prompt was deliberately plain: "Upgrade the project to 1.22.2".

Every run produced a project that passed the execution gates, which looks like success. A closer inspection showed the agent had bumped the headline version but missed dependencies and migration details.

Two sets of figures are in circulation, and they shouldn't be mixed:

  • The overview post reports that the baseline passed 30 of 80 configuration checks. Telling the agent to use CLI for Microsoft 365 raised that to 75 of 80.
  • The detailed write-up reports 30/80 on configuration correctness and 34/50 on dependency currency for the baseline. With the CLI, it reports 75/80 and 49/50.

The detailed write-up explains why the agent went wrong:

  • It fetched the 1.22 release notes but not every intermediate release page. That is risky because SPFx upgrades are incremental.
  • It formed its plan before reading the docs, then used the docs to confirm that plan.
  • A tip suggesting the CLI didn't change its approach, because the same page offered detailed manual steps that looked actionable.

Fixing the docs instead of shipping a skill​

The authors tested their changes with Dev Proxy, so they didn't have to publish untested ideas to Microsoft Learn. Moving the tip made no difference. Removing the migration guide hurt manual runs without reliably increasing CLI adoption.

What worked was a warning placed directly before the tip, one that challenged the agent's manual approach. In the write-up's simulations, CLI adoption went from none in the baseline to agents finding and using it on their own. Restoring the step-by-step guide sent agents back to the manual route. The team then rewrote the guide to explain what changes in the gulp-to-Heft migration while pointing to the CLI to apply them. Agents chose the CLI again.

The team says the fixes landed in SPFx documentation as PRs #10855 and #10921. The result is that every developer and agent benefits, including those without any skill installed.

The write-up also says the investigation produced a platform fix. Microsoft Learn can return pages as Markdown. The team shared that with the GitHub Copilot teams, who added Accept: text/markdown support to GitHub Copilot Chat and GitHub Copilot CLI.

"Sounds agent-friendly" isn't the same as "is"​

The post's most useful section tests ideas that sound sensible. Each result applies only to the scenarios and agent profiles measured.

IdeaWhat Microsoft measured
Add a JSON input mode to your CLIWith regular arguments, every agent profile got all 5 runs right. In JSON mode, Claude Haiku 4.5 got 2 of 5 deployments right, and every model cost 4x to 11x more per task.
Add a docs tip pointing to the right toolA more direct tip led to tool use in 1 of 5 runs. A warning naming the failing approach got 5 of 5.
Give the agent another documentation sourceThe context7 MCP server added no meaningful lift on SPFx upgrades.
Switch to the model with cheaper tokensSonnet 5 cost 3.7x more per run than Sonnet 4.6, despite 33% lower per-token pricing.

On the context7 result, the two Microsoft accounts differ slightly. The overview says its tools didn't load in 3 of 5 runs and weren't called in the other 2. The detailed write-up says they didn't load in 3 of 5 runs and were available but never invoked in the other 2. Either way, availability didn't make the tools relevant to the agent's plan.

The write-up also tested a separate anti-hallucination skill. It improved results and cut average token use by roughly 9%. It still left the project only partially upgraded, with configuration correctness at 37/80.

The vocabulary​

  • Agent extensions: skills, MCP servers, instruction files and custom agents.
  • Bare baseline: the same tasks with no extensions, using only the harness and model. Every change is compared against it.
  • Lift and drag: a change either improves outcomes (lift) or leaves them the same or worse (drag).
  • Task cost: the full cost of one task, with every token counted at its price.
  • Agent profile: the operating system, harness, model and settings, and loaded extensions. A result belongs to the profile it was measured on. It isn't the same as the agent profile files some tools use.
  • Propensity and efficacy: whether the agent chooses your technology, and whether it uses it correctly.

How AX relates to DX, GEO and llms.txt​

  • DX: AX builds on it. Agents are users who are easier to observe, because you can run the same task five times with and without a change.
  • GEO (generative engine optimization): this is about getting AI search to mention your product. It overlaps with propensity. AX also asks whether the agent uses the product correctly once it has picked it.
  • llms.txt and MCP: these are surfaces you can offer agents, so they count as part of AX. Like any best practice, they are hypotheses until you measure them.

How to start: a practical procedure​

  1. Pick your technology. Write a few tasks your developers actually give agents, in their own words.
  2. Include some tasks that don't name your product (propensity) and some that do (efficacy).
  3. Run each task several times on a bare baseline with a fixed agent profile.
  4. Read the agent's trace, meaning what it fetched and where its plan diverged.
  5. Fix the cause on surfaces developers already use, usually docs or interfaces. Build a new extension last.
  6. Rerun the same tasks. Compare quality and task cost, and label the change lift, drag, or lift that costs too much to keep.
  7. Scope your conclusion to that agent profile, and rerun when the profile changes.

If you've already shipped a skill or MCP server, start at step 2 and compare runs with and without it.

Analysis: what to take from it​

The strongest point is the method. Five runs per scenario is small, and Microsoft says so. The post stresses that results hold only for the measured profile. The value is in the loop of baseline, trace, fix, remeasure.

Caveat on perspective. This is a Microsoft developer-advocacy post about Microsoft's own product, SPFx. The case studies are favorable to Microsoft's tooling and docs. Microsoft Learn and the CLI for Microsoft 365 are the fixes that worked. The principle is vendor-neutral, but the numbers are specific to one scenario, one harness (GitHub Copilot Chat on Windows) and one model.

For Windows and Microsoft 365 admins and developers, two takeaways stand out:

  • Agents reach for plausible, actionable content over tips. A warning that challenges the agent's plan beat a gentle suggestion.
  • "Passes the build" isn't the same as "upgraded correctly." Review agent output before it ships, a point Microsoft itself makes about its SPFx Dev Skills preview.

AX is an evolving field. Agent behavior will shift with each new model and harness, which is exactly why the post argues for measuring rather than trusting a demo that looked fine.

 

References

  1. What is Agent Experience (AX)? Microsoft Developer Blogs 2026-10-06T09:18:05+00:00
  2. Agent Experience netlify.com
  3. Behind SPFx Dev Skills: testing what agents know and fixing what they miss - Microsoft 365 Developer Blog devblogs.microsoft.com