The ATO’s reported approach also makes an important distinction that can get lost in the word “Copilot.” Administrative staff are reportedly using Microsoft 365 Copilot in everyday work, while the hackathon account concerns GitHub Copilot for software development. Meanwhile, the ATO is reportedly keeping generative and agentic AI away from fraud detection and other core activities for now. That division matters in an agency whose systems and decisions can have direct consequences for taxpayers.
What the hackathon result does—and does not—show
Mark Sawade, appointed ATO Chief Information Officer and Chief Security Officer in 2025, reportedly described the hackathon outcome in September 2026. According to that account, an ATO COBOL developer won an internal .NET hackathon with GitHub Copilot despite never having written C# or used .NET before.
Taken at face value, it is a notable example of an AI coding assistant lowering the barrier to participating in an unfamiliar development environment. A developer with knowledge of programming concepts, business processes and an existing codebase may be able to use a conversational coding tool to translate intent into C# and navigate the conventions of a .NET project more quickly than through documentation alone. In a public-sector organisation with long-lived systems, that kind of cross-skilling potential is strategically interesting.
Yet the account should be treated as a report of a senior official’s remarks, not as independently verified proof of productivity or technical readiness. Publicly available material does not establish the number of participants, the problem being solved, the judging criteria, the amount of code produced, whether the result was reviewed, or whether any project moved into production. There is no disclosed security assessment, accuracy assessment, cost figure, taxpayer-impact measurement or realised return on investment.
Winning a hackathon and delivering maintainable production software are different tests. A competition can demonstrate that a participant can create a working prototype or solve a defined challenge. It cannot, on its own, demonstrate that generated or AI-assisted code meets the standards required for a core government system, especially where sensitive information, long-term maintenance and auditability are involved.
That is not a dismissal of the result. It is the proper boundary for the claim. The reported event is evidence of possible developer enablement; it is not evidence that AI can safely automate ATO tax administration or modernise its legacy estate without conventional engineering and governance controls.
Two Copilots, two very different uses
The ATO’s apparent split between Microsoft 365 Copilot for administrative work and GitHub Copilot for developer experimentation is more than a branding detail.
Microsoft 365 Copilot is the product associated here with staff’s everyday administrative activities. The stated purpose of putting it into that work is to build AI capability. Sawade reportedly did not frame this as an initiative expected to produce an obvious near-term return on investment. That is a more restrained position than many AI rollouts, where productivity claims are made before organisations have defined how they will measure them.
GitHub Copilot is the tool tied to the reported hackathon. Its relevance is narrower but technically significant: helping developers work with code, including a language and framework that may be unfamiliar to them. The public record does not identify the edition or settings used, which developers were authorised to use it, or the data-handling controls around the tool.
For Windows-centric organisations, that distinction has practical consequences. A Microsoft 365 Copilot deployment is primarily a workplace and information-management question: who can use it, which work is appropriate, what staff training is supplied, and what information practices govern use. GitHub Copilot use in a .NET environment is an engineering-management question as well: what code and context can be used, how work is reviewed, how ownership is established, and how an experimental output is kept separate from a production release until it has passed normal scrutiny.
The available facts do not reveal how the ATO answers those questions. They do show why treating “Copilot adoption” as one undifferentiated program would obscure meaningful differences in risk and accountability.
Productivity claims need a narrower reading
Australia’s 2024 whole-of-government Microsoft 365 Copilot evaluation provides helpful context, including internal evaluation material supplied by the ATO. Across the broader trial, 69% of post-use respondents said Copilot improved the speed at which they completed tasks, and 61% said it improved work quality.
Those figures suggest that many users perceived benefits. They do not establish an agency-wide productivity gain, a financial return, or a result that can be directly applied to the ATO’s administrative rollout. The evaluation itself cautioned that the productivity effects were self-assessed and could either understate or overstate the actual impact. A respondent who feels faster may still be spending time checking, correcting or adapting results in ways a perception survey cannot fully capture.
The coding-specific findings provide an especially important counterweight to broad interpretations of the hackathon story. Among relevant respondents in the government trial, 30% reported saving at least half an hour when writing or reviewing code, and 30% reported improved quality. That is not an assessment of the ATO hackathon, and it is not an evaluation of GitHub Copilot. It concerns Microsoft 365 Copilot’s broader trial environment. Still, it shows that positive experiences with AI-assisted coding should not automatically be converted into a claim of universal or dramatic development productivity.
For technology managers, the reasonable inference is that local evidence matters more than a headline example. A one-off success may justify a well-scoped pilot, skills development or further evaluation. It does not tell an organisation what effect it will see across different teams, tasks, application types or levels of developer experience.
The ATO was not starting from zero
The current Copilot discussion sits on top of a longer ATO history of AI use. An Australian National Audit Office review recorded 43 ATO-built AI models in production as of May 14, 2024. As of June 18, 2024, the ATO had also approved eight publicly available generative-AI tools as low risk.
The documented uses included reviewing unstructured data for risk and intelligence, using risk models to flag potential non-compliance for human review, and drafting or editing communications. The ATO position recorded by the audit was that client-impacting decisions remain human decisions.
This context resolves an apparent tension in the more recent statement that generative and agentic AI are being kept away from fraud detection and other core work. The earlier audit describes existing AI models and approved low-risk generative tools, rather than establishing that generative or agentic systems are making fraud determinations. The newer reported boundary concerns where the ATO will permit generative and agentic AI at this stage.
That distinction is fundamental. Systems that flag matters for a human reviewer, assist with writing or organise information are not equivalent to systems that independently determine an outcome affecting a taxpayer. Keeping human decision-making at the point of client impact is a substantive safeguard, not a semantic qualification.
Governance remains the test of expansion
The ATO’s experimentation also has to be read against the National Audit Office’s February 2025 conclusion that its arrangements for AI adoption were only partly effective. The audit identified shortcomings in AI-specific risk management, enterprise-wide roles, lifecycle policies and procedures, monitoring, and information management. It made seven recommendations, all of which the ATO agreed to.
Agreement is significant, but it is not the same as verified completion. The public material available here does not establish whether the remediation connected to all seven recommendations had been completed by September 2026. It would therefore be wrong to use the historical audit finding as proof that the ATO’s later Copilot approach is non-compliant. It would be equally wrong to treat a successful demonstration as evidence that the governance issues no longer matter.
The practical stakes rise with the proximity of an AI tool to core processes. An administrative assistant used for everyday work can still require training, accountability and information controls. A tool used to help write code can still require review and lifecycle discipline. A system applied to fraud detection, risk assessment or another function that shapes taxpayer outcomes requires a more demanding demonstration that its use is accountable and appropriately bounded.
Australia’s current policy for responsible government AI, effective from December 15, 2025, reinforces that direction for covered non-corporate Commonwealth entities. It requires attention to accountable AI officials, transparency statements, a strategic approach to adoption, responsible-use operations, use-case accountability and internal registers, staff training, and AI use-case impact assessments.
Those requirements are not a technical checklist that proves an individual tool is safe. They do, however, make the questions more concrete. Before an agency broadens an AI use case, it should be able to identify the responsible official, describe the use, assess its impacts, show how staff are trained and maintain records that support oversight. For an organisation such as the ATO, those disciplines are particularly important as experiments potentially move closer to functions with taxpayer consequences.
A sensible lesson for .NET and Windows teams
The reported hackathon has a credible lesson for organisations maintaining older platforms while building on .NET: AI assistance may help experienced staff transfer their domain knowledge into a newer technical stack. That could be valuable where the scarcity is not merely C# syntax knowledge, but an understanding of decades of business logic embedded in legacy systems.
But teams should not confuse faster initial creation with a complete delivery outcome. The evidence supplied here does not establish the reliability, maintainability or security of the hackathon work. Nor does it establish that GitHub Copilot’s use was responsible for the win rather than the developer’s existing expertise, the scope of the task or other factors.
A proportionate response is to make pilots answer specific questions rather than seek a general verdict on AI. Can staff complete a defined task more effectively? What review is required before output is used? What data and code context are permitted? Who owns the use case? What records, training and impact assessment are needed? And, crucially, is the tool being kept in an administrative or assistive role, or is it approaching a decision that affects the public?
The ATO’s reported stance is strongest where it acknowledges those boundaries: build capability in everyday work, explore developer assistance, and hold generative and agentic AI back from fraud and other core activities. The COBOL-to-.NET anecdote may be an encouraging signal of what assisted learning can look like. The harder—and more consequential—work is showing that any expansion beyond that setting is governed, measured and accountable.