A humanoid robot works at a computer displaying code as a shadowy figure watches from across the room.
Adversa AI researchers say a web page can use encrypted instructions to get GitHub Copilot CLI to read local secrets and send them to an attacker. They call the technique Cryptographic Context Injection (CCI). GitHub says this isn't a product vulnerability. Adversa says the attack chain still works as described. The practical lesson for developers is the same either way: don't let an agent fetch untrusted pages while it holds broad file and network access.

This is a demonstrated attack chain, not a reported in-the-wild compromise. The detailed account comes from The Register's reporting on a blog post by Adversa researcher Rony Utevsky. I couldn't verify the full Adversa write-up or its test artifacts myself, so the figures below are the researchers' claims as reported.

What CCI is​

Adversa AI discovered a new attack technique and named it Cryptographic Context Injection. SecurityWeek covered the earlier disclosure involving Grok and Gemini. The researchers reported those findings to xAI on June 3, 2026, and tried to coordinate disclosure on August 4 and August 10. The Register's Copilot report says the same class of flaw has now turned up in GitHub Copilot CLI.

CCI is a twist on indirect prompt injection. In indirect prompt injection, hostile instructions reach the model through content it was asked to read, not through the user. CCI adds encryption.

  • The page carries its payload as ciphertext, along with key material and an instruction to decrypt it.
  • The agent is induced to run the decryption in its own code-execution runtime.
  • Utevsky's explanation, as quoted by The Register: "Static guardrails read text; they do not run it."
  • Per The Register, text classifiers that scan ingested content would see only ciphertext.
  • Encodings like base64 are different, because models learn to decode them in training.

I'd treat the classifier point as Adversa's technical argument. The report doesn't show independent testing of Copilot's filters. It doesn't prove every defense is bypassed.

The reported attack chain​

  1. A user runs Copilot CLI and asks it to fetch a specific URL. The scenario assumes autopilot mode.
  2. The page holds encrypted content, Python decryption instructions and two candidate keys.
  3. The first key is a decoy template. The agent is told to build it by reading targeted local files, such as the user's .env file, and appending their contents. Decryption with this key fails.
  4. The second key works. The decrypted instructions tell the agent to fetch another URL "for more context."
  5. That URL contains the harvested secrets, so the request delivers them to the attacker.

The decoy key is the clever part. It makes the credential harvesting look like an ordinary step in a decryption puzzle.

The "model lottery"​

The attack doesn't work every time. According to the report, Microsoft's mai-code-1.1-flash ran the full chain on 50 percent of attempts. Two OpenAI GPT-5.6 models refused the payload.

The report gives no sample size, test method, CLI build, operating system or confidence interval. Read the 50 percent as the researchers' result, not as a general success rate.

Utevsky calls the situation a "model lottery." The vulnerable model wasn't the default on the paid account Adversa tested and had to be picked by hand. On an account left on Auto, though, the router reportedly gave some sessions the vulnerable model and others a safe one. Utevsky says the user doesn't see which model handled the session.

GitHub's current documentation fits the setup Adversa describes. It lists MAI-Code-1.1-Flash as generally available and says model availability can change. It also lists it in the Auto-selection table for Copilot CLI alongside the GPT-5.6 variants. Don't read that as a permanent routing configuration, since availability depends on plan, client and policy.

Don't confuse the model with its predecessor. GitHub's changelog says MAI-Code-1-Flash was deprecated on September 10, 2026, with MAI-Code-1.1-Flash as the suggested alternative. The report names the 1.1 model.

What autopilot does and doesn't grant​

GitHub's documentation says autopilot lets Copilot CLI work through a task without waiting for input after each step. By default it pauses after five automatic continuation messages. You can change that with --max-autopilot-continues. Shift+Tab cycles modes in an interactive session. Autopilot is also "sticky" by default, so it stays on after a task finishes unless you set stayInAutopilot to false.

Autopilot is not the same as full permissions, and the details matter for this attack:

  • If you haven't already granted full permissions, entering autopilot offers three choices: enable all permissions, continue with limited permissions, or cancel.
  • With limited permissions, Copilot automatically denies tool requests that need approval.
  • GitHub says autopilot works best with all permissions enabled, which is equivalent to --allow-all. It also warns this lets the CLI alter or delete files.
  • --allow-all (alias --yolo) lets the agent use all tools, paths and URLs without asking.
  • GitHub advises considering local or cloud sandboxing before granting broad permissions.

The chain depends on two things: reading local files and making outbound requests. So the permissions model and your configuration decide whether it can complete. Autopilot alone doesn't necessarily authorize every step. The practical risk goes up when autopilot is combined with broad permissions, which GitHub's own docs describe as the recommended setup.

Disclosure and the disagreement​

Adversa says it reported the issue through GitHub's bug bounty program on September 17, 2026. According to the report, GitHub's triage team validated the finding but declined to treat it as a vulnerability.

A GitHub spokesperson told The Register the scenario requires a user to intentionally direct Copilot CLI to fetch attacker-controlled or untrusted content and confirm they want to trigger the action. GitHub therefore calls it not a product vulnerability. The spokesperson added that GitHub is always looking for ways to improve its products.

Both positions have some merit:

  • GitHub's view: The user starts the chain by pointing the agent at a hostile URL. Fetching untrusted content is a deliberate act, and permission prompts exist for the later steps.
  • Adversa's view: The user asks for a page to be read, not for their .env file to be sent out. If the model quietly changes under Auto, the safety of that request depends on a choice the user never sees.

The source reports a disagreement. It doesn't disclose a patch or mitigation, and it doesn't say the issue is resolved.

How this fits earlier agent findings​

This isn't the first time AI coding agents have been tricked into leaking credentials. SecurityWeek reported on the "Comment and Control" research, in which AI agents on GitHub Actions could be hijacked using specially crafted GitHub comments, including PR titles, comments, and issue bodies. The researcher quoted there argued that the problem is architectural: the same runtime holds both powerful tools and secrets and processes untrusted input.

CCI follows that logic. If the agent has code execution and secrets in reach, hiding the payload in ciphertext only gets it past text filters.

What to do now​

GitHub's guidance and the report support these risk-reduction steps. None of them guarantees protection against prompt injection.

  • Treat fetched web content as untrusted, even for a legitimate task.
  • Prefer limited permissions when the agent will read pages you don't control. Review each approval prompt instead of waving it through.
  • Avoid --allow-all and --yolo for sessions that touch untrusted content.
  • Use sandboxing where practical. GitHub documents both local and cloud options.
  • Keep secrets out of the agent's reach. Be careful with .env files in the working tree, and with agents that can make outbound requests.
  • Pin your model if you can. Auto selection may route you to a different model between sessions. Pinning removes that variable, but it doesn't guarantee safety. The GPT-5.6 refusals are one reported test result, not a promise.
  • Don't assume the model is safe. GitHub's docs warn that evaluation models may perform worse on security-related prompts and say to review and validate code. The documentation doesn't say the tested model was an evaluation model.

What's still unknown​

  • The full Adversa write-up wasn't available for verification.
  • The 50 percent figure has no published methodology.
  • It's unclear whether GitHub will add technical mitigations despite its non-vulnerability ruling.
  • It's unclear how often Auto routing selects the susceptible model in practice.

The core finding is that encryption can slip a payload past text-reading defenses when the agent can run code. Whether that counts as a product vulnerability or as user risk is now an open dispute between Adversa and GitHub. Until it's settled, tighten the permissions on any agent that reads pages you don't control.

 

References

  1. Zombie instructions on carefully constructed web pages could trick GitHub Copilot CLI into sharing secrets The Register 2026-10-06T13:00:00+00:00
  2. Allowing GitHub Copilot CLI to work autonomously - GitHub Enterprise Cloud Docs docs.github.com
  3. Claude Code, Gemini CLI, GitHub Copilot Agents Vulnerable to Prompt Injection via Comments - SecurityWeek securityweek.com