Back to all blogs
#ai
#python
#neat
#ml
January 2, 2026
5 min read
A blog post byThereallo

Teaching an AI to Play Chrome Dino

Watching a neural network figure out physics from scratch still feels magical.

I know what year it is. It's 2025. We have LLMs that can write code, generate video, and argue about philosophy. Building an AI to play the Chrome Dino game isn't exactly groundbreaking anymore.
But sometimes you just want to see a machine learn something from scratch. No pre-training, no terabytes of internet data. Just a population of random brains, a fitness function, and natural selection.
So I built one using Python, Selenium, and NEAT.
The journey was fun. I watched the AI find the laziest way to cheat, punish it, and force it to actually get good.

The architecture is simple. I didn't want to use computer vision because screenshotting and processing frames adds latency. I wanted raw data.
I used Selenium to inject JavaScript directly into the browser's console. This lets us pull the game state instantly.
python
return {
    crashed: r.crashed,
    playing: r.playing,
    score: actualScore,
    speed: r.currentSpeed,
    dinoY: r.tRex.yPos,
    obs: nextObs ? {
        x: nextObs.xPos,
        y: nextObs.yPos, 
        w: nextObs.width,
        h: nextObs.typeConfig.height
    } : null
};
This feeds into a neural network with:
Inputs: Speed, Distance to Obstacle, Obstacle Y Position, Obstacle Dimensions.
Outputs: Jump, Duck, or Run.
The algorithm uses NEAT (NeuroEvolution of Augmenting Topologies). It starts with random connections. Most dinos run into the first cactus. Some jump randomly. The ones that survive the longest get to breed and mutate.

The first few generations were terrible. But around Generation 5, the AI figured out a strategy.
It realized that being in the air is generally safer than being on the ground. So it just started spamming the jump button. It was bunny hopping through the desert.
text
2025-12-15 08:37:36,680 [SpawnPoolWorker-6] Died. Score: 123
2025-12-15 08:37:38,293 [SpawnPoolWorker-5] Died. Score: 47
2025-12-15 08:37:39,034 [SpawnPoolWorker-10] Died. Score: 149
It worked for a while. But it's not playing the game. It was basically cheesing it.
I tried to fix this by adding a penalty. I deducted points if the AI jumped when the obstacle was far away.
text
2025-12-15 08:55:17,486 [SpawnPoolWorker-8] Died. Score: 45 | Penalty: 0.0
2025-12-15 08:55:17,655 [SpawnPoolWorker-3] Died. Score: 45 | Penalty: 0.0
But the math was broken. If the penalty for a "bad jump" is 50 points, and the reward for surviving a cactus is 30 points, the AI does the math and decides:
"It is more profitable to die immediately than to jump and get fined."
So I might have accidentally taught my AI that death was preferable to taxes.

To fix this, I had to balance the economy of the fitness function.
  1. Survival = Goal: Clearing an obstacle grants +30 points.
  2. Bad Jumps = Tax: Jumping unnecessarily costs about 3 points total.
  3. Result: Bunny hopping is profitable (27 points net), but perfect play is more profitable (30 points net).
The AI will learn to bunny hop to survive, but over time, the "clean" players will outscore the spammers and replace them in the gene pool.
I also had to add a "Training Wheels" logic. At low speeds, we punish spamming. At high speeds (> 9.0), we disable the penalty because prejumping is indeed a valid strategy. This is to prevent the AI from spamming the jump button at lower speeds.
python
if raw_dist > jump_limit:
    if action == 0 or state['dinoState'] == "JUMPING":
        fitness_penalty += 0.1  # tax
    else:
        fitness_penalty -= 0.01 # reward for efficiency

Once the math was fixed, evolution took over.
By Generation 14, we hit a score of 7,000. By Generation 17:
text
2025-12-15 09:34:04,556 Gen Summary: Best=11271.3, Avg=1257.4
2025-12-15 09:43:14,087 Gen Summary: Best=15000.0, Avg=1560.1
The AI hit the fitness cap. Ingame, that's a score of 10,000. The game is moving at maximum velocity. The reaction windows are milliseconds. And the AI doesn't miss.

Generation 17 Replay

0:00 / 0:00

The best part about NEAT is that the resulting networks are often shockingly simple. It strips away anything it doesn't need.
Here is a visualization of the winner's brain using Graphviz:
The graph means that:
  1. Obs Y (Obstacle Height) is the steering wheel.
    • Green line to Jump: If Y is high (Cactus), Jump.
    • Red line to Duck: If Y is high, don't Duck.
  2. Obs Dist (Distance) is like the gas pedal.
    • It powers everything. If there's an obstacle, get ready to act.
  3. Speed is the governor.
    • The red line from Speed to Run suppresses the "Run" action as the game gets faster, making the Dino trigger happy on the jump button.
It ignored Obstacle Height and Width almost entirely. It figured out that position and distance are the only variables that matter.

It doesn't. We have GPT-5. We have agents that can code full apps.
But there is something incredibly satisfying about watching a 1KB pickle file outplay a human at a game purely through trial and error.
It reminded me that intelligence isn't always about massive datasets and billions of parameters.
Sometimes, it's just about optimizing a few weights until you stop running into a cactus.