IT Brief reported on August 28 that IBRS adviser Joe Sweeney sees this pattern particularly among mid-tier organisations: AI arrives through an existing licensing relationship, individual departments begin experimenting, and the organisation has not decided which work should change, who owns the risk, or how employees should judge an answer before acting on it. That diagnosis is supported by a much larger, public Australian Government evaluation of Microsoft 365 Copilot: deployment was comparatively easy, but adoption and reliable value were much harder.
The distinction deserves attention from Microsoft 365 administrators. Copilot can be provisioned at scale within a familiar stack of Teams, Word, Outlook and SharePoint. The difficult work begins after assignment: configuring permissions and information architecture, selecting high-value tasks, training people in those tasks, and making verification an operational habit rather than a warning in an acceptable-use policy.
Australia’s Copilot trial showed where the adoption work begins
The Australian Digital Transformation Agency’s whole-of-government trial put more than 5,765 Microsoft 365 Copilot licences into agencies between January and June 2024. It offers a useful counterweight to vendor adoption figures because it measured use in workplaces that already had a substantial Microsoft 365 footprint and a defined reason to experiment.
The results were positive but narrower than a simple “AI saves time” story. The evaluation found that 77% of participants were optimistic about Copilot by the end of the trial and 86% wanted to continue using it. Yet only one in three used it daily. Use centred on summarising material, drafting and rewriting content, with Microsoft Teams and Word accounting for the bulk of activity.
That is important when boards calculate return on investment from seat counts. A company can buy or inherit hundreds of Copilot licences and truthfully say it has rolled out AI, while most employees only use a small set of familiar functions—or do not use the product often enough to alter a business process. The licence is an input cost; it is not evidence of changed work.
The government evaluation also documented a constraint Sweeney highlighted to IT Brief: confidence rises with training that applies to the employee’s actual job. Seventy-five percent of participants who received three or more forms of training were confident using Copilot, a 28-percentage-point advantage over those who received only one form. The evaluators said generic prompt training was insufficient on its own; participants needed agency-specific examples, clear guidance about accountabilities and prompt information, and internal champions who could demonstrate practical benefits.
Microsoft Australia has its own reason to emphasise capability. In July, the company said 18 of the ASX Top 20 and 66 government agencies use Microsoft 365 Copilot, while 16 Australian organisations operate deployments above 10,000 seats. Those figures describe significant market penetration. They do not establish how deeply Copilot has been incorporated into decisions, workflows or measurable customer outcomes—a gap Microsoft itself acknowledges when it urges customers to focus on impact rather than the number of pilots and licences.
Training must include the decision to reject an answer
The most useful part of Sweeney’s warning is not the familiar instruction to keep sensitive information out of a public AI prompt. That remains necessary, but it does not address the operational risk of an apparently plausible answer generated inside an approved environment.
According to IT Brief, Sweeney described an AI-generated recommendation that, if followed, could have led to the deletion of an entire CRM database. The article does not identify the system, model, prompt or controls involved, so it cannot establish whether the failure was hallucination, ambiguous instruction, excessive access or operator error. But it illustrates the correct lesson: a tool can comply with a data-handling policy and still produce advice that is unsafe to execute.
The government’s Copilot evaluation reached the same practical conclusion from a broader trial. It found that output unpredictability and limited contextual knowledge created verification and editing work that reduced some efficiency gains. Up to 7% of participants said Copilot added time to their activities. Sixty-one percent of managers surveyed could not confidently identify Copilot-generated output.
For IT administrators, verification training should therefore be tied to the consequences of the task. A generated meeting summary may need a spot check against the transcript. A proposed policy or financial analysis requires the owner to validate sources, assumptions and calculations. A suggested PowerShell command, SQL statement, retention change or CRM action needs testing in a non-production environment, peer review where appropriate, and permissions that make a single bad instruction incapable of causing an irreversible result.
This is where a generic “use AI responsibly” course fails. Staff need examples from their own systems showing where Copilot is useful, what data it can reach, which actions it may recommend but cannot approve, and when the human operator must stop. The Australian Government’s AI adoption guidance makes a similar division: responsibility must exist both across the organisation and for each specific use of an AI system, because the same model creates different risks when drafting marketing copy, reviewing job applications or acting on enterprise data.
Department-by-department AI creates a shadow problem
Sweeney told IT Brief that AI adoption is frequently being driven by managers within individual departments instead of a coordinated organisational strategy. This is a predictable result of AI features being embedded in productivity suites: finance, HR, sales and marketing leaders can each find a credible use case quickly, without waiting for an enterprise transformation program.
The danger is not that every team must wait for a central committee to approve every prompt. The danger is that the company ends up with separate rules, duplicate experiments, uneven training and no shared record of which data, connectors, agents and models are in use. Employees who see other teams receiving better AI access or training may turn to personal accounts and public tools to close the gap. That is shadow AI: AI use outside an organisation’s security review, governance and visibility.
Australian governance data suggests the risk is neither theoretical nor limited to the public sector. The Governance Institute of Australia’s 2025 survey reported that 46% of respondents had no AI training, 88% had struggled to integrate generative AI with legacy systems and 93% could not effectively measure AI return on investment. Its sample was 344 respondents, so those figures should not be treated as a census of Australian business. They do, however, corroborate the implementation failures raised in the IBRS interview: skills, system integration and measurement are being left behind by procurement and experimentation.
A central AI program does not need to ban departmental initiative. It should give it a controlled path. Teams should be able to propose a use case, classify the data involved, name a business owner, document the expected outcome, test the process and report what happened. That supplies IT with an inventory of actual AI activity and lets successful practices travel between departments rather than remaining dependent on a single enthusiastic manager.
Measure work quality and rework, not prompts and licences
Sweeney’s argument that productive AI users employ a “scientific mindset” is more than a rhetorical preference for critical thinking. It points to the measurement problem that has distorted many Copilot discussions. Counting prompts, active users or documents summarised may reveal adoption, but those numbers do not show whether the work was improved or merely produced faster.
The Australian Government trial found that 69% of survey respondents believed Copilot improved the speed of tasks and 61% believed it improved work quality. About 40% reported redirecting time to higher-value work such as mentoring, planning, stakeholder engagement and product improvement. Those are promising results, but they were self-reported perceptions in a non-randomised trial whose participants volunteered or were nominated; the government explicitly noted selection bias and lower response rates in post-use surveys.
Organisations should treat that as a reason to measure a smaller number of concrete workflows rather than force a universal productivity target. A service desk may track first-response quality, escalation rates and time spent preparing case notes. A sales team may assess whether account briefs are more complete and whether follow-up work declines. A software team can examine whether AI-assisted code review increases defects found before production without inflating review time. Finance or legal teams should measure correction rates and the time spent validating generated drafts, not simply documents produced.
The result may be that some use cases do not justify paid licences, at least yet. That is not failure. A stopped pilot that exposes bad data, unclear ownership or an untestable workflow is cheaper than a broad rollout that hides those problems behind monthly active-user charts.
The first corrective action is an inventory, not another purchase
For Australian organisations that have already assigned Copilot seats, the short-term response is straightforward: identify every licensed group, its enabled capabilities, its connected data and its actual use case. Then compare active use with a role-based training plan and a defined review process for outputs that affect customers, records, money, access or production systems.
The organisation should name an accountable business owner for each consequential AI use, while IT retains visibility into identity, permissions, connectors, retention and logging. The Australian Government guidance specifically calls for a senior AI governance owner and a responsible person for every AI system, alongside documentation that can be audited and revisited as use changes.
Copilot’s availability in a Microsoft agreement may make deployment feel like a procurement detail. The evidence from Australia says it is a change-management and control problem instead. Companies that cannot identify which work is better, who validated the output and where their data travelled have not yet captured AI value—they have simply made the technology easier to access.