A handheld retro game console connects to glowing network diagrams and coin stacks on a warmly lit desk.
An AI model reached the Hall of Fame in Pokémon Red for about the price of a vending-machine coffee. It also had a lot of help, and the help is the more useful story for developers.

Christian Mathiesen, CTO of Frigade, pointed TypeSafe AI's decision model Jev at the 1996 Game Boy game. The run lasted 37 hours and 40 minutes. Jev made 16,150 decisions and used roughly 39.2 million input tokens, for a reported model charge of about $1.65 (around 260 yen, by GIGAZINE's conversion). GIGAZINE's headline also carries a second finding that matters just as much. When Jev was asked to choose raw button presses, it never got out of Pallet Town, the town where the game starts.

Both halves of the story teach the same lesson for anyone building AI into software. The model makes choices. The code around it decides which choices exist and carries them out.

The run by the numbers​

The project's public GitHub repository is the main source for the results:

MetricReported figure
Total time37h 40m
Decisions made16,150
Input tokens~39.2 million
Total Jev cost~$1.65
Typical decision time~0.4 seconds
Team wipes16 (14 of them at the Elite Four)
Elite Four attempts15
Final teamCharizard 83, Graveler 62, Nidoqueen 45, Beedrill 44, Haunter 39, Primeape 29

The repository says the run was livestreamed on YouTube from September 25 to 26, 2026. The stream has ended and highlights are on the project's landing page.

A correction to the early coverage: GIGAZINE listed a level 39 Gastly in the final team. The project README says Haunter, and Archyde's report also lists an LVL 39 Haunter. The README is the primary source, so Haunter it is.

GIGAZINE adds two figures that are not in the README's results table: about 1,660 battles, and "Champion Road" as the hardest area. Champion Road is the Japanese name for Victory Road, the last dungeon before the Elite Four. Those two figures come from GIGAZINE's account.

The cost checks out. TypeSafe lists Jev at $0.042 per million input tokens, with output free. 39.2 million × $0.042 comes to about $1.646, which matches the reported total. Keep in mind this is only the model charge. It does not include engineering time, hardware, the livestream, or the weeks spent building the harness.

Two numbers don't quite line up, and the README doesn't explain why:

  • Tokens per decision. 39.2 million tokens across 16,150 decisions works out to about 2,400 tokens per decision. The README describes a typical call as about 1,200 tokens.
  • Decision rate. 16,150 decisions over roughly 37.7 hours is about 430 per hour. The README estimates 800–1,300 calls an hour at real-time speed.

One explanation is that "calls" and "decisions" aren't counted one-to-one. That's a guess, not something the project documents.

Summary: the headline numbers are internally consistent and mostly come from the primary source. The $1.65 is a real model bill, but it doesn't cover the whole project.

What Jev decided and what the harness did​

The system is a loop. A Game Boy emulator running under Node reads the game's RAM to work out the current state: map, party, battle and on-screen text. It gives Jev a list of legal options with supporting facts. Jev picks one. Then conventional code turns that choice into button presses.

Aggregator AI/TLDR summarizes the setup the same way. A Node.js harness runs a Game Boy emulator, reads game memory, and turns the current situation into a list of legal options with context. Jev is called through Vercel AI Gateway to pick one, and the harness only presses buttons — it never writes to game memory.

Jev decided:

  • Every menu answer, including names, the starter Pokémon, yes/no prompts, shopping, healing, and learning or forgetting moves
  • What to focus on: progress, healing, training, catching, shopping, exploring or team management
  • Which Pokémon to catch, and which to swap in and out of the party at the PC
  • Where to go and who to talk to
  • Every battle action: move, switch, item, Poké Ball or run

The harness handled:

  • Movement on the map, using A* pathfinding. If Jev chose a destination, code planned the route and pressed the buttons.
  • Menu navigation
  • Supporting facts attached to each option, such as type matchups, damage estimates and progress toward the objective
  • A story-milestone list saying what the next goal is and where it happens, checked against the game's real event flags

The project calls itself a run with "no scripts or cheats," and a few design choices back that up. The harness never writes to game memory. Hidden items were not shown to Jev, because a human player wouldn't know where they are.

The house rules were:

  • Text speed set to FAST and battle animations turned off, once, at boot
  • The player named JEV and the rival BLUE
  • Every caught Pokémon given a made-up nickname that Jev spelled one letter at a time (no species names or names already in use)

Summary: Jev chose from menus of structured options. Deterministic code built those menus and carried out the choices.

The Pallet Town problem​

According to GIGAZINE, a Hacker News commenter said they had expected an AI that picks which button to press, not one that chooses abstract goals like "head east to Lavender Town." Mathiesen reportedly replied that this was exactly how the first version worked, and that Jev couldn't leave Pallet Town.

That anecdote comes from GIGAZINE's account of the discussion. The README describes only the final design, not the failed first attempt.

It's still an instructive failure. A model built for fast choices among a few options is badly suited to making hundreds of low-level moves in a row with no memory. Another independent project, milanboers/jev-plays-pokemon, describes the same constraint: Jev has no conversation history. Unlike an LLM in a chatbot, Jev doesn't get a growing transcript of past turns — each turn is a fresh, self-contained request. That project also ended up with Jev picking a high-level goal each turn (talk to Mom, explore, leave the room, advance text); code executes it with A* pathfinding.

If a model can't remember that it just walked into the same wall three times, something else has to remember for it.

Loop protection: the safety net​

Mathiesen's harness has three layers of recovery for when things go wrong:

  • Tagging: options that were already tried without changing the game state get flagged, and Jev is told to try something else.
  • Sampling: if the same failing choice keeps coming back, the harness picks an alternative itself.
  • Rollback: as a last resort, the harness reloads the latest milestone checkpoint.

The rollback matters when reading the results. Some of the 37 hours came from the system rewinding, not from Jev working its way out of trouble. This result belongs to the model and the harness together. It doesn't show that Jev can recover from every failure on its own.

It wasn't the only Jev run​

Jev became a popular test subject within days of its launch. Tom's Hardware reported that LangChain described it as having "had a pretty outsized response" since its launch on Sept. 15, had multiple developers playing the game within 10 days.

A separate run by Andrew Boyd, founder of Standard Agents, reached the Hall of Fame first. Boyd's project page says Jev "beat the Elite Four and the Champion and entered the Hall of Fame on September 23, 2026". That run had its own kind of help: Anthropic's Claude Opus 5 monitored the game log and adjusted options and their wording as Jev played, effectively acting like a coach.

Mathiesen's run, which finished on September 26, used a fixed harness instead of a second model editing the options mid-game. That doesn't make it better or worse. It does mean these projects test different things:

ApproachWho writes the options?Recovery
Mathiesen (Frigade)Fixed harness code with a milestone listTagging, sampling, checkpoint rollback
Boyd (Standard Agents)Harness, adjusted live by Claude Opus 5An LLM coach revising options and wording

Tom's Hardware's longer-running comparison point is Anthropic's own stream: Claude Plays Pokémon stream, running Opus 4.5 at the time, still hadn't finished Red as of January. Treat that comparison with care. General-purpose LLM runs usually receive much less structured help, so this isn't a like-for-like contest.

Reading the vendor claims​

TypeSafe founder Diogo Almeida, who says he helped build the instruction-following research behind ChatGPT, announced Jev on September 15, 2026 as the first "System One" model. Its outputs are typed decisions with probabilities and confidence scores, not free text.

Several of TypeSafe's claims should be read as vendor claims:

  • Pricing: the company admits it can't prove the price isn't subsidized.
  • Benchmarks: its workflow evaluations were written by its own team, and TypeSafe acknowledges that some bias could exist.
  • "Can't hallucinate": this means Jev's output always fits the defined structure. It doesn't mean every choice is correct. Sixteen team wipes show that a perfectly structured decision can still be a bad one.

Can you reproduce it on a Windows PC?​

The code is public under GPL-2.0-or-later, but the README sets clear limits.

Prerequisites:

  • macOS or Linux, Node 20 or newer, and Git. Windows is not listed.
  • Your own legally obtained US/EU Pokémon Red ROM with SHA-1 ea9bcae617fdf159b045185467ae58b2e4a48b9a. Other versions won't work because the harness reads RAM addresses from the pret/pokered disassembly. No ROM is included.
  • A Vercel AI Gateway key if you want to run the real Jev model

Documented steps:

  • Clone the repository and run npm install.
  • Run npm run setup. This installs rgbds via Homebrew, clones and builds pret/pokered, and generates symbol data.
  • Copy your ROM to roms/red.gb.
  • Copy .env.example to .env. Set JEV_MODE=gateway and add your AI_GATEWAY_API_KEY.
  • Start the run with npm start -- --speed 1. A local viewer runs at localhost port 8787. Add --resume to continue from the latest milestone checkpoint.

Common problems:

  • Wrong ROM: a regional or revised version won't match the RAM addresses. Check the SHA-1 first.
  • Mock mode by accident: the default JEV_MODE is mock, which the README calls a free, "dumb stand-in." A run in mock mode is not a Jev run.
  • Windows: WSL is an obvious route to a Linux environment, but the project doesn't document or test it, and setup depends on Homebrew. You'd be on your own.
  • Cost: the default throttle is 300 ms between calls and at most 90 calls a minute. The README says that at real-time speed the bot costs roughly $1–1.70 per 24 hours, with a worst case of about $7 a day.

What success looks like: the local viewer shows the game advancing. logs/jev-calls.jsonl records every call: the full state, the options with their facts, the probabilities and the latency. logs/events.jsonl records map changes, milestones, saves and errors.

The takeaway for people who build software​

Beating Pokémon Red isn't really the point. The run is a clear example of a design pattern TypeSafe is promoting: a cheap, fast model that picks among options your code defines, while ordinary code handles pathfinding, validation and recovery.

The Pallet Town failure is the honest caveat. The model's performance depends on how well the options are designed. Give it too little structure and it stalls in the first town. Give it good structure and it can beat the Elite Four for under two dollars in tokens, with a lot of engineering underneath doing much of the work.

 

References

  1. The AI 'Jev' achieved Hall of Fame status in 'Pokémon Red' in 37 hours and 40 minutes, costing only about 260 yen, but it requires an assistance system and, if fully autonomous, cannot leave Pallet Town. - GIGAZINE GIGAZINE 2026-09-28T05:40:00+00:00
  2. Jev Plays Pokémon Red github.com
  3. GitHub - milanboers/jev-plays-pokemon: Playing Pokemon Red using TypeSafe Jev · GitHub github.com