
12.5M-Parameter World Model Learns Pokémon Red on One RTX 3080 Ti
A 12.5M-parameter JEPA-style world model, trained from scratch on a single RTX 3080 Ti, learned to pick Squirtle in Pokémon Red — succeeding in 52 of 100 planning runs.
- By
- Rebecca Stone
- Filed
- Channel
- AI & Compute
- Read
- 3 min read
A developer training a 12.5-million-parameter world model on a single NVIDIA GeForce RTX 3080 Ti has taught it to pick a starter Pokémon in Pokémon Red — a modest but concrete demonstration that small models can be built and trained entirely on consumer hardware.
The developer, who goes by "stmonty," documented the effort in a Sept. 20 blog post titled "Teaching a World Model to Play Pokémon." On a discussion forum, stmonty described the project as a way "to learn and have some fun" and said the model was chosen precisely because it could be trained locally on a gaming-grade graphics card.
The architecture follows the LeWorldModel research line co-authored by Yann LeCun, the AI scientist who left Meta last year to pursue world models at his AMI Labs. The specific design traces to a March 2026 paper describing a single JEPA model of roughly 15 million parameters — trainable on a single GPU — which stmonty's 12.5-million-parameter implementation closely mirrors. The model was trained from scratch for the project.
Training consumed 42,382 grayscale frames collected across more than 1,000 short runs, all starting from a save inside Professor Oak's lab. The dataset mixed three sources: scripted routes, the same routes with random button presses injected, and pure random wandering. The random component matters. A model trained only on clean runs "might see A pressed whenever a dialogue box appears and never learn what B does there," stmonty explained.
Rather than predicting full future frames, the model works with compressed summaries of just 192 numbers, keeping the compute footprint small enough for a GeForce-class card. Initial training is "reward-free" — the model receives no signal about winning conditions. Goals enter later, at planning time.
Planning works in 14-press chunks. Each round samples 512 candidate plans, and the search prunes them to one in eight, leaving 64 to carry forward. This continues until a final plan is executed in the game emulator — crucially, plans are evaluated inside the model's own predictions, not against the live game.
That design likely caused the first failure, which left the player without a Pokémon. Stmonty hypothesizes that small prediction errors compound when the model plans from its own predicted outcomes rather than real frames. After fine-tuning, the second attempt succeeded: the model selected Squirtle.
The results, while narrow, are quantified. Over 100 runs from the same save, the trained model produced 52 plans that obtained a starter Pokémon. Random button presses achieved zero successes, and an untrained model managed exactly one. Stmonty points out that simply mashing the A button already works from that particular save — the point was that the model discovered the solution on its own, without being told.
The developer is realistic about the ceiling. The next milestone, which stmonty "may still try," would start in the lab, walk to the Professor, sit through the dialogue, and only then pick a Pokémon. "I suspect that difficulty scales exponentially with plan length," the developer said. Scaling the current model larger would probably not beat the game either: "I don't think simply making my current model bigger would get us there," stmonty stated in a forum reply.
The project sits in a busy niche. At least two separate, successful attempts to beat Pokémon Red outright have appeared in recent weeks using the Jev decision model with a harness or custom code — more sophisticated approaches than stmonty's. Small models remain the realistic option for hobbyists without datacenter budgets.
The code is open source on GitHub as lePokeRed, and a CUDA GPU is recommended for anyone aiming to reproduce the work. Whether the approach can stretch beyond short-horizon tasks is the open question: stmonty's own exponential-difficulty hypothesis suggests plan length, not parameter count, will be the binding constraint.
Source: Tom's Hardware
More from Rebecca Stone
Show full bio
Correspondent covering media and advertising at Chip Dispatch.
104 articles
Related articles
openai-turns-to-samsung-for-next-gen-ai-chips-as-partnership-deepens-76564f5f
OpenAI Turns to Samsung for Next-Gen AI Chips as Partnership Deepens
openai-triggered-bidding-war-for-hugging-face-ahead-of-nvidia-s-13b-deal-fe5aa2b7
OpenAI Triggered Bidding War for Hugging Face Ahead of Nvidia's $13B Deal
amd-buys-world-labs-for-8-2b-lands-fei-fei-li-as-chief-scientist-46c62f61
AMD Buys World Labs for $8.2B, Lands Fei-Fei Li as Chief Scientist
amd-buys-fei-fei-li-s-world-labs-for-8-2-billion-in-all-stock-deal-bcb44983
AMD Buys Fei-Fei Li's World Labs for $8.2 Billion in All-Stock Deal



