A desktop displays an AI voice-cloning process beside a phone call and prompts to call back or use a security key.
Giorgia Meloni's EUIPO filing got the headlines on October 5, but the more useful story for IT readers is what AI voice cloning means for identity checks. Local voice cloning has also reached ordinary Windows PCs. This piece separates what the evidence supports from the more dramatic numbers circulating in the coverage, and it ends with what to do about it.

What actually happened in Rome​

ANSA reported on October 5 that the Italian prime minister filed a four-second audio recording of the words "Io sono Giorgia Meloni" with EUIPO, the EU's intellectual property office. ANSA framed the move as protection against voice cloning by generative AI. It also noted that singer Giusy Ferreri had done something similar in recent months.

Three details matter:

  • It is an application, not a granted right. ANSA says the EUIPO record lists the filing as under examination. Nothing in the reporting shows a registration or enforceable protection yet.
  • It is a sound mark. That is a trademark category, not copyright in a voice.
  • It is tied to specific services. ANSA says the application covers three kinds of services: downloadable multimedia content, cultural activities, and the organization of political events.

How far a sound mark could reach beyond those categories is a legal question I can't settle from the filing coverage. A trademark is, by design, a tool for commercial-context disputes. It is not a technical control. It does nothing to stop a model from imitating a voice, and it is a weak match for a private phone scam.

"Three seconds" and "70%": handle with care​

Much of the coverage repeats two statistics as if they described the same thing. They don't.

  • Security.org's deepfake page reports that, in a McAfee survey, 70 percent of people said they weren't confident they could tell a real voice from a cloned one. That is self-reported confidence, not a measured failure rate. It does not show that 70% of listeners are fooled by a clone.
  • The same page says three seconds of audio is sometimes all that's needed to produce an 85 percent voice match. The word "sometimes" matters. This is a similarity claim, not a guarantee that any three-second clip yields an indistinguishable fake.

The original coverage also cited a 900% rise in deepfakes, a 1,300% jump in voice-cloning crime, a 45–50% detector error rate on compressed audio, and multi-million-euro Italian bank frauds. I could not independently confirm those figures in the material I reviewed. Treat them as unverified until you can see the underlying studies and their methods.

The Windows angle: local models are real​

Cloud cloning services leave accounts, billing records and logs. Open-source models that run on your own hardware change that picture, and two projects show it.

  • OmniVoice (k2-fsa). The project describes itself as a zero-shot text-to-speech model covering over 600 languages, with voice cloning among its features. Its documentation covers NVIDIA GPUs, Apple Silicon and Intel Arc GPUs. It also includes a local web interface.
  • GPT-SoVITS. Its README advertises zero-shot TTS from a roughly five-second sample, plus few-shot fine-tuning with about a minute of data. For Windows, it offers an integrated package for Windows 10 and later, launched through a batch file that opens a web interface. The project's tested-environment table is narrower than some tutorials imply, so setup varies by hardware and software versions.

I'm deliberately not reproducing install commands. A step-by-step guide to cloning someone's voice helps attackers more than defenders. The point for administrators is simpler: the capability is free, documented and runs on a mid-range gaming PC.

"No trace" is overstated​

Some commentary says local cloning is invisible to investigators. That claim goes too far. Running inference offline means no cloud provider sees the request. It does not erase the rest of the chain. Downloading model files, installing software, handling audio files and delivering the result over a call or message can all leave evidence on endpoints, networks and accounts. What local use mainly removes is one convenient source of logs, and that makes provider-side monitoring a weaker defense.

The EU rules: what the dates really are​

The coverage said the Digital Omnibus "stretched out" the AI Act's labeling rules. That is only partly true, and the difference matters for any organization publishing synthetic media.

  • Article 50(4) requires deployers who publish deepfakes to disclose visibly that the content is artificially generated or manipulated, and it has applied since 2 August 2026.
  • Legalithm notes that the Digital Omnibus, Regulation (EU) 2026/1744, did not move that date.
  • The Omnibus transition is narrow. Generative systems already on the market before 2 August 2026 must deliver machine-readable marking under Article 50(2) only from 2 December 2026. New systems are covered from day one.
  • Fines for Article 50 breaches sit in the second tier, up to EUR 15 million or 3% of worldwide annual turnover, according to Legalithm's summary of Article 99(4).

There is also a structural gap. Marking duties fall on providers and deployers. A person running downloaded open-source weights to impersonate someone has no interest in labeling the output, and a labeling rule won't deter fraud. Watermarking helps the well-behaved part of the market and does little about the rest. That is general industry reasoning, not a finding from any one source.

What defenders should do​

Security.org's practical advice is plain: if someone calls, texts or video messages asking for help, call the person directly to verify. Its guidance also warns that detection software often fails. As an example, only one of four free tools flagged the "Biden" robocall audio as AI-generated. That page is a general consumer guide, and its detector example is a single case, not a benchmark.

For IT and security teams, the sturdier approach is to stop relying on a voice at all:

  1. Never let a familiar voice be the only authorization for payments, bank detail changes, password resets or access grants.
  2. Call back on a number you already have on file, not one the caller gives you.
  3. Require a second channel or strong authentication for verbal approvals, such as a ticket in your service desk, a signed request or a hardware token.
  4. Train help desk and finance staff on urgency, secrecy and sudden channel switches as warning signs.
  5. Agree a code word with family and executives' assistants for emergency requests.
  6. Don't treat detectors as a gate. Use them as one input. A clean result is not proof of authenticity.

Public audio is the raw material. Podcasts, conference talks and social video give attackers clean samples, so executives who speak publicly should assume their voices are already available.

Bottom line​

Meloni's filing is a legitimate, if symbolic, attempt to treat a voice as a protectable asset. It is still only an application under examination. The practical risk sits elsewhere: free, local, documented cloning tools mean a familiar voice on the phone is no longer evidence of identity. The fix is process: verified callbacks, second channels and strong authentication. A new detector or a new trademark won't do it.

 

References

  1. Cloning a Voice With AI Takes Just a Few Seconds of the Real One - Pasquale Pillitteri Pasquale Pillitteri 2026-10-05T18:17:00+00:00
  2. Giorgia Meloni registra la sua voce all'ufficio marchi contro i fake con l'IA | ANSA.it ansa.it
  3. AI Act Transparency: Article 50 and Deepfake Rules - Legalithm legalithm.com