Skip to content

A 12-unit GRU that runs a game cat: how Jello Shop's cat decides what to do

Project documentation · 2 min read · updated

The cat in Jello Shop is controlled by a gated recurrent unit (GRU) with 12 hidden units and 1,194 weights, which is 4,776 bytes as 32-bit floats. Twice a second it reads 18 numbers about itself and the room, updates its memory, and picks one of four actions: sleep, sit, walk or groom. Two more outputs set the walking direction and pace. The weights were trained offline with evolution strategies and ship as a 6,592-byte JavaScript file.

Network
GRU with 12 hidden units, 18 inputs and 6 outputs
Weights
1,194 float32 values: 4,776 bytes (base64 in a 6,592-byte file)
Decisions
One every 0.5 seconds
Training
Offline, evolution strategies, 400 generations; final reward 0.028 per decision
Runs in
Your browser, plain JavaScript, no library

What are the 18 inputs?

InputsMeaning
0, 1How rested the cat is, and how restless
2Whether it is on its cushion
3, 4Offset from the cat to the cushion (x, y)
5, 6The cat's own position on the floor
7How close Jell-omo is
8, 9Offset from the cat to Jell-omo (x, y)
10Whether Jell-omo is moving
11 to 14Which action the cat is doing now (one-hot)
15How long it has been doing it, capped at 30 seconds
16Whether it was just poked
17A little Gaussian noise

What are the 6 outputs?

Outputs 0 to 3 are scores for sleep, sit, walk and groom. They go through a softmax at temperature 1, and the action is sampled from the result, not simply the highest score. Outputs 4 and 5 pass through tanh to give a heading; the pace is the length of that heading vector, capped at 1.

A new action only takes effect once the current one has been held for a minimum time: 20 seconds for sleep, 4 for sit, 4 for walk and 5 for groom. Because the actions are sampled and the hidden state carries over from one decision to the next, the cat never plays the same day twice.

Where do 1,194 weights come from?

A GRU has three blocks, the update gate, the reset gate and the candidate. Each has 12 × 18 input weights, 12 × 12 recurrent weights and 12 biases, which is 216 + 144 + 12 = 372. Three blocks make 1,116. The read-out adds 6 × 12 weights and 6 biases, which is 78. In total 1,116 + 78 = 1,194.

How was it trained?

Offline, with evolution strategies, in a simulator that runs the same catbrain.js file as the game. Candidates were rewarded for keeping the cat rested and comfortable: sleeping on the cushion, getting up when restless, stretching its legs and coming home when tired. The shipped weights are the result of 400 generations, ending at a reward of 0.028 per decision. Nobody wrote a schedule for the cat; the trainer lives in dev/scripts/train-cat-brain.mjs, and running it with --eval reports how the shipped weights behave.

What does the network remember?

Only its hidden state: 12 numbers that carry from one decision to the next, so a choice can depend on the recent past and not only on the moment. It does not learn while you play; the browser runs the trained network forward and nothing else.

Quick answers

How big is the cat's brain?
A GRU with 12 hidden units and 1,194 weights: 4,776 bytes of 32-bit floats, shipped as base64 in a 6,592-byte JavaScript file.
Does the cat learn while I play?
No. It was trained offline with evolution strategies. In the browser the trained network only runs forward, twice a second.
Does the cat do the same thing every time?
No. Actions are sampled from the network's scores and its hidden state carries over, so behaviour never repeats exactly.

More on Jello Shop