Two officers review an AI-generated aircraft plan on screen, enforcing two-person verification and an audit trail.
The chief constable of West Midlands Police, Craig Guildford, has retired after an independent review and parliamentary scrutiny found that his force’s recommendation to bar away supporters from Maccabi Tel Aviv’s Europa League visit rested in part on a fabricated item generated by Microsoft’s Copilot — an AI hallucination that slipped into operational intelligence and precipitated a political crisis. ice.uk](])

Background​

West Midlands Police advised Birmingham’s multi‑agency Safety Advisory Group (SAG) that the Europa League match between Aston Villa and Maccabi Tel Aviv on 6 November 2025 posed sufficient risk to justify recommending that Maccabi’s travelling supporters should not attend. That recommendation, presented as a public‑safety measure, was later discredited when the force’s intelligence pack was shown to contain multiple inaccuracies — including a reference to a non‑existi Tel Aviv fixture that had been generated by a generative AI assistant.

The Chief Inspector of Constabulary’s preliminary review, led by Sir Andy Cooke, identified a catalogue of shortcomings: overstated claims about numbers and injuries, weak provenance for intelligence items, poor community engagement, and confirmation bias in compiling material that supported a pre‑determined operational option. The Home Secretary publicly declared she had “no confidence” in the chief constable after the report’s release, and the political pressure that followed culminated in Guild16 January 2026.


What exactly happened: a concise timeline​

October–November 2025: Decision and match​

  • West Midlands Police presented intelligence to Birmingham SAG ahead of the Aston Villa v ure.
  • On 6 November 2025, the match proceeded without visiting fans after the SAG recommendation; policing operations reported arrests and heightened security, though no major stadium disorder occurred.

December 2025–January 2026: Scrutiny and thealists and parliamentary committees scrutinised the force’s intelligence pack and identified discrepancies, including the fictitious West Ham fixture that could not be substantiated.​

  • Initially senior officers attributed the error to a routine web search or sociaequent enquiry showed the item had been produced by Microsoft Copilot during open‑source research and incorporated into briefings without adequate verification.
  • Chief Constable Craig Guildford later wrote to the Home Affairs Select Committee to apologise and correct the record, acknowledging Copilot’s involvement.

Mid‑January 2026: Pretirement​

  • The inspectorate’s public summary and parliamentary exchanges produced a strong political reaction, with the Home Secretary saying she had lost confidence in Guildford. Daily pressure, the inspectorate’s findings, and continuing oversight from the Independent Office for Police Conduct (IOPC) led to Guildford’s decision to retire on 16 January 2026. lucination became operational evidence

Generative AI assistants such as Microsoft Copilot are designed to summarise, synthesise and assist with open‑source research. They produce coherent prose by predicting probable next tokens, not by querying an infallible database of facts. Under the right (or wrong) conditions, that statistical prediction produces plausible but false outputs — the phenomenon commonly referred to as a hallucination. When such outputs are treated as primary evidence rather than provisional leads, they can migrate from a researcher’s draft into briefing documents and then into policy. In this case, an officer used Copilot during open‑source research; the assistant produced a reference to a past fixture that never occurred. That item was not caught during normal checks, and it migrated into an intelligence product presented to the SAG. The operational chain therefore contained three failure points:

1. a generative AI output that contained fabricated content;

2. human failure to verify the claim against primary sources; and

3. inadequate documentation and provenance in the intelligence product, which allowed t influence a decision that restricted people’s movement.


Why this matters: rights, trust and high‑stakes decision making​

Policing decisions that limit civil liberties — including recommen defined group from attending an event — require the highest standards of verification and traceable evidence. An unverified AI‑generated claim, even if plausible, dohold. When such claims are used to justify exclusions that disproportionately impact a minority community, the consequences go beyond reputational damage: they erode trust, inflame community tensions, and invite legal and political scrutiny.

The West Midlands episode highlights three connected risks:

  • Evidential risk: AI outputs lack automatic provenance; they must be traced back to original sources before being treated as facts.
  • Institutional risk: processes and leadership must ensure verification and a culture of challenge; otherwise, confirmation bias can legitimise weak evidence.
  • Political risk: once public bodies are seen to use unverified AI outputs in decisions affecting civil liberties, political accountability and calls for reform accelerate. The Home Secretary’s public loss of confidence exemplifies that dynamic.

Technical and governance gaps exposed​

Generative assistants are not oracles​

Copilot and similar assistants perform well at summarisation and drafting but are not replacements for primary‑source verification. Vendors explicitly warn users about hallucinations; nonetheless, organisations are increasingly using these tools in time‑sensitive workflows withouthe vendor disclaimers are insufficient on their own if internal controls are weak or absent.

Auditability and provenance were absent or insufficient​

The inspectorate found poor record‑keeping and weak command strgence product. For any claim used to limit rights, police forces must be able to show the chain of custody: who asked the question, what AI prompt or query produced the output, what primary sources were checked, and who authorised inclusion in an operational briefing. That audit trail was missing here.

Training and procurement shortcomings​

  • Analysts who rely on AI need accredited training that emphasises verification first.
  • Procurement must include contractual requirements for provenance features, archiveable queries, and vendor transparency about retrieval modes (local indexing, web retrieval, or model synthesis).

Wider landscape: this is not an isolated failure​

Other organisations have suffered consequences from unverified generative AI outoitte agreed to partially refund a government contract after a report produced with generative AI contained fabricated references and quotes; that episode led to a corrected report and extra scrutiny of consultancy practices. The parallels are instructive: plausible‑looking fabrications can appear in professional outputs when AI is used without strict human verification. Legal filings and academic citations have also been shown to include AI‑generated errors in multiple jurisdictions — a pattern that underlines the need for systemic controls rather than ad hoc fixes.


What the inspectorate and oversight bodies recommended (and what they did not)​

The inspectorate’s preliminary review did not find evidence that antisemitism or political interference alone explained the force’s recommendation, but it did identify an imbalance in how evidence was gathered and presented. The report highlighted:

  • Eight demonstrable inaccuracies in the intelligence product.
  • Confirmation bias in privileging evidence that supported a pre‑existing decision.
  • Weak community engagement, particularly inadequate outreach to the Jewish community prior to the SAG decision.

Oversight bodies such as the IOPC have stated they will continue examining available evidence and may open independent conduct investigations if appropriate. That oversight can proceed regardless of Guildford’s retirement, and it underscores that organisational accountability can outlast personnel changes.


Practical lessons and concrete safeguards for police forces (and public bodit rely on open‑source research and AI must convert lessons into operational rules. Recommended measures include:​

  • Require an AI usage register: list approved tools, users and business purposes.
  • Mandate prompt and output archiving for any AI query that contributes to official reporting.
  • Implement a two‑person verification rule for any claim used to restrict movement or rights.
  • Enforce a ‘no‑inclusion’ default: AI outputs can be used to identify leads, but not to provide final evidence, unless retraced to primary sources.
  • Contractually demand provenance and logging from vendors and refuse black‑box retrieval modes for sensitive workflows.
  • Provide accredited training for analysts that emphasises provenance, bias testing and adversarial review.
  • Maintain an audit trail for intelligence products: who produced, who validated, and who approved.

These are practical, implementable steps that reduce tcinations becoming operational facts.


Policy and legal implications​

This episode has reignited debate over the legal mechanisms available to hold senior public servants accountable. The Home Secretary publicly declared loss of confidence in the chief constable but lacked the statutory authority to remove him directly; that gap prompted calls for reform of the frameworks that regulate senior police appointments and dismissals. The political element is consequential because accountability structures shanisational change. At a sectoral level, Parliament and inspe push for sector‑wide guidance on generative AI in public services, including mandatory auditability and provenance in intelligence workflows. These changes will require both statutory clarity and operational investment.


Where responsibility sits — and why single‑actor narratives are incomplete​

There is a natural tendency in public controversy to single out an individual as the locus of blame. In this case, Guildford’s retirement answers an immediate political question but does not by itself fix the systemic failures that allowed an AI hallucination to influence a high‑stakes decision.

Accountability in this episode is layered:

  • Individuals who used AI and failed to verify outputs share proximate responsibility.
  • Supervisory and managerial structures that failed to enforce verification and documentation share organisational responsibility.
  • Procurement and IT functions that allowed ungoverned AI usage — or bought tools without demanding provenance features — share responsibility at the institutional level.

Fixing the problem therefore requires a mix of personnel accountability, process re‑engineering and procurement reform.


The vendor question: what can platform providers do?​

Vendors of generative assistants can help by:

  • Building stronger provenance features that identify and link generated claims to indexed sources.
  • Providing enterprise‑grade logging and immutable archives of prompts and outputs.
  • Offering conservative default settings for sensitive workflows (e.g., “no synthesized citations”).
  • Supplying clear documentation and training materials tailored to public‑sector use cases.

However, vendor improvements are necessary but not sufficient: public bodies must demand these features contractually and adapt their internal processes. The vendor ecosystem will improve under procurement pressure; the West Midlands case is a clear signal that procurement terms and vendor accountability matter.


Final analysis: a turning point for AI governance in public services​

The collapse of trust in the West Midlands intelligence product — catalysed by a generative AI hallucination — is a vivid, real‑world example of how weak governance can convert a technology flaw into a political and civilis neither a verdict against AI per se nor an excuse for inattention: the tool produced an error, the organisation failed to catch it, and the combination produced an outcome that affected people’s rights and confidence in policing.

If public bodies act decisively — codifying AI‑use registers, enforcing provenance and archiving, training analysts, and toughening procurement — this episode can lead to durable, worthwhile reform. If they do not, similar incidents will recur as generative assistants are adopted more widely to manage volume and complexity.

The salient lesson is operational and organisational: treat AI outputs as provisional, not primary; require traceable verification for decisions that impact rights; and design systems so that a plausible senodel can never become an operational fact without primary‑source confirmation. The costs of complacency are now visible in the political, legal and reputational fallout that followed a single fabricated match.


Recommendations checklist for immediate imple forces and councils)​

  • Publish an AI usage register and update it quarterly.
  • Mandate archiving of all AI prompts and outputs used in research that feeds operational reporting.
  • Require two‑person sign‑off for any intelligence claim that affects civil liberties.
  • Insert a hard stop in procurement contracts: no black‑box retrieval for sensitive use cases.
  • Roll out practitioner training focused on provenance, verification and adversarial review.
  • Create a public summary of AI governance for event‑risk decisions to rebuild community trust.

These steps are short, implementable and proportionate to the risks exposed by the West Midlands episode.


The West Midlands case is a cautionary tale: generative AI can speed research and streamline analysis, but it amplifies existing organisational weaknesses when used without strict procedural controls. The retirement of a chief constable answers a political moment; the longer test is whether policing bodies translate the episode into enduring operational and procurement changes that safeguard rights and restore public confidence.