A study of 137 college students using Microsoft Copilot found that users rated its answers as less credible and less relevant when the assistant’s conversational style did not fit the role they expected it to play. The result is useful for Copilot users and administrators because it separates a familiar complaint — “the answer feels wrong for this job” — from the narrower question of whether the answer is factually wrong.

The research, by Elon University’s Qian Xu and Oregon State University’s Cheng Chen, was published online in Behaviour & Information Technology on March 3, 2026. Elon publicized the paper on August 24. The university’s announcement and the journal record agree on the core design: participants worked with Copilot in Creative, Balanced, or Precise styles while producing a public-health strategy for their home county across at least three rounds of interaction.

The important finding is not that users prefer one canned tone over another. Xu and Chen found that the perceived usefulness of Copilot’s output depended on the combination of style and expected role: whether a user approached the tool as a co-author, creator, curator, or conversational partner.

A woman studies a laptop displaying AI collaboration dashboards, analytics, and warning indicators.The study tested communication fit, not Copilot accuracy​

The paper measured readability and language features in both user prompts and Copilot responses, along with participants’ perceptions of credibility and relevance. It did not establish that a Creative, Balanced, or Precise response was more accurate on the underlying public-health task. Nor did it test whether participants made better decisions after following Copilot’s suggestions.

That boundary is central for organizations deploying Copilot. A response can be well written, aligned with a user’s preferred tone, and rated highly credible while still needing factual review. Conversely, an answer that is accurate but terse, overly cautious, or stylistically mismatched can be dismissed by a user before its substance is checked.

The study therefore describes a trust and usability issue rather than an evaluation of model truthfulness. For IT leaders, that points to a practical risk: people may treat communication quality as evidence quality, especially when a tool is positioned as a collaborator rather than a search interface.

Microsoft has long presented Creative, Balanced, and Precise as choices intended to tailor Copilot conversations. Xu and Chen’s work adds evidence that such labels affect more than presentation. They can influence the language users themselves produce and the way they judge what comes back.

Copilot changed the users’ prompts, too​

The researchers found that Copilot’s selected conversational style affected the readability of participant prompts. In other words, users did not simply receive differently phrased outputs; they adapted their own writing as the interaction progressed.

That is a consequential finding for anyone using prompt logs to assess adoption, training needs, or employee behavior. A prompt is often treated as a direct record of what a worker wanted from an AI tool. This experiment suggests it can also be a record of the interaction the product has already shaped.

Xu and Chen also identified linguistic alignment between users and Copilot in emotional tone and clout, a language-analysis measure associated with confidence and social status. The alignment appeared across all three styles. But the users and Copilot diverged on analytical-thinking language.

The mismatch on analytical language should not be overread as proof that Copilot is incapable of reasoning or that users became less analytical. The accessible journal abstract does not make either claim. It does show that a system can converge with a user’s emotional and assertive language without converging on the language cues associated with analytic expression.

For workplace use, that is a warning against assuming that a smooth exchange reflects shared understanding. Copilot may sound aligned with the employee who is prompting it while still framing the work differently.


“Co-author” and “curator” are different jobs​

The most actionable part of the paper is its treatment of expected AI roles. A user who sees Copilot as a co-author may want drafts, alternatives, argument structure, and revisions. A user treating it as a curator is more likely to want relevant material organized and synthesized. Someone expecting a creator may want breadth and originality; someone treating the tool as a conversational partner may prize responsiveness and dialogue.

Those expectations shape what “good” looks like before Copilot writes a word.

The study’s result suggests that product defaults are blunt instruments when the task is unclear. A concise, highly structured reply may satisfy a user who wants a research aide but disappoint one who expected collaborative drafting. A more expansive, idea-generating answer can have the opposite effect. The cost is not only annoyance: participants reported lower credibility and relevance when style and expected role diverged.

Neither the Elon release nor the publicly accessible journal abstract identifies which specific style-role pairings produced the strongest or weakest ratings. That missing detail limits any claim that administrators should assign Creative to one job title and Precise to another. The data supports a narrower conclusion: a one-size-fits-all conversational setting will create avoidable friction where employee tasks differ.

It also matters that “role” is not necessarily a setting the Copilot interface exposes or enforces. In the study, anticipated role is an experimental concept describing what the participant expected from the system. In a live deployment, those expectations are created by training, sample prompts, interface wording, leadership messaging, and the task employees bring to the tool.

Training should define the deliverable before the tone​

Many Copilot deployment guides focus on prompt specificity: provide context, name the desired output, set constraints, and verify the response. The Xu-Chen study gives organizations a reason to add one more step: state what job the assistant should perform in the interaction.

For example, a request for a “co-author” should specify the draft’s intended audience, source material, argument, and revision goal. A request for a “curator” should specify inclusion criteria, exclusions, and whether the output should summarize, rank, or compare information. A request for a “creator” should ask for alternatives and explain the level of novelty or risk that is acceptable.

A workable internal prompt pattern would be:

  • “Act as a co-author and produce a first draft that I will revise, using these source notes and this audience.”
  • “Act as a curator: extract the decisions, owners, deadlines, and unresolved issues from these meeting notes.”
  • “Act as a creator: generate five distinct campaign concepts, then explain the assumption behind each one.”
  • “Act as a conversational partner: help me explore options one at a time, and challenge unsupported assumptions.”

These are not magic phrases, and the study does not validate them as a universal formula. Their value is operational: they make the user’s role expectation visible. That gives Copilot more usable context and gives the employee a clearer basis for judging whether the answer meets the task.


The experiment is informative, but narrow​

The study used 137 college students completing a specific strategic-planning exercise. That makes it a controlled look at interaction patterns, not a census of how engineers, finance teams, lawyers, support agents, or Windows administrators use Copilot in production.

The task also involved at least three turns, which is enough for observable adaptation but much shorter than the extended, document-grounded work that defines many Microsoft 365 Copilot deployments. The results should therefore be treated as evidence about a mechanism — expectations affect interaction and perception — rather than a forecast of adoption, productivity, or error rates in an enterprise.

There is another timing issue. The paper’s online publication predates Elon’s August announcement by nearly six months, and the study necessarily examined a prior Copilot experience. Microsoft’s Copilot interfaces, models, grounding options, and product boundaries change frequently. No public evidence attached to this announcement shows that the exact effect sizes or style behavior would replicate unchanged in the August 2026 Copilot product.

Still, the underlying lesson is durable: users do not approach an AI assistant as a neutral text generator. They bring a model of what the system is supposed to be doing. When the product’s behavior conflicts with that model, confidence can fall even before anyone checks whether the answer is correct.

For Copilot administrators, the immediate consequence is straightforward. Configure governance and data access carefully, but do not treat enablement as a technical switch-flip. Training and prompt templates should tell employees whether they are asking Copilot to draft, retrieve, organize, brainstorm, or debate — and they should require a human check when credibility is being inferred from style rather than evidence.