A 12-unit GRU that runs a game cat: how Jello Shop's cat decides what to do
The cat in Jello Shop is controlled by a gated recurrent unit (GRU) with 12 hidden units and 1,194 weights, which is 4,776 bytes as 32-bit floats. Twice a second it reads 18 numbers about itself and the room, updates its memory, and picks one of four actions: sleep, sit, walk or groom. Two more outputs set the walking direction and pace. The weights were trained offline with evolution strategies and ship as a 6,592-byte JavaScript file.
- Network
- GRU with 12 hidden units, 18 inputs and 6 outputs
- Weights
- 1,194 float32 values: 4,776 bytes (base64 in a 6,592-byte file)
- Decisions
- One every 0.5 seconds
- Training
- Offline, evolution strategies, 400 generations; final reward 0.028 per decision
- Runs in
- Your browser, plain JavaScript, no library
What are the 18 inputs?
| Inputs | Meaning |
|---|---|
| 0, 1 | How rested the cat is, and how restless |
| 2 | Whether it is on its cushion |
| 3, 4 | Offset from the cat to the cushion (x, y) |
| 5, 6 | The cat's own position on the floor |
| 7 | How close Jell-omo is |
| 8, 9 | Offset from the cat to Jell-omo (x, y) |
| 10 | Whether Jell-omo is moving |
| 11 to 14 | Which action the cat is doing now (one-hot) |
| 15 | How long it has been doing it, capped at 30 seconds |
| 16 | Whether it was just poked |
| 17 | A little Gaussian noise |
What are the 6 outputs?
Outputs 0 to 3 are scores for sleep, sit, walk and groom. They go through a softmax at temperature 1, and the action is sampled from the result, not simply the highest score. Outputs 4 and 5 pass through tanh to give a heading; the pace is the length of that heading vector, capped at 1.
A new action only takes effect once the current one has been held for a minimum time: 20 seconds for sleep, 4 for sit, 4 for walk and 5 for groom. Because the actions are sampled and the hidden state carries over from one decision to the next, the cat never plays the same day twice.
Where do 1,194 weights come from?
A GRU has three blocks, the update gate, the reset gate and the candidate. Each has 12 × 18 input weights, 12 × 12 recurrent weights and 12 biases, which is 216 + 144 + 12 = 372. Three blocks make 1,116. The read-out adds 6 × 12 weights and 6 biases, which is 78. In total 1,116 + 78 = 1,194.
How was it trained?
Offline, with evolution strategies, in a simulator that runs the same catbrain.js file as the game. Candidates were rewarded for keeping the cat rested and comfortable: sleeping on the cushion, getting up when restless, stretching its legs and coming home when tired. The shipped weights are the result of 400 generations, ending at a reward of 0.028 per decision. Nobody wrote a schedule for the cat; the trainer lives in dev/scripts/train-cat-brain.mjs, and running it with --eval reports how the shipped weights behave.
What does the network remember?
Only its hidden state: 12 numbers that carry from one decision to the next, so a choice can depend on the recent past and not only on the moment. It does not learn while you play; the browser runs the trained network forward and nothing else.
Quick answers
- How big is the cat's brain?
- A GRU with 12 hidden units and 1,194 weights: 4,776 bytes of 32-bit floats, shipped as base64 in a 6,592-byte JavaScript file.
- Does the cat learn while I play?
- No. It was trained offline with evolution strategies. In the browser the trained network only runs forward, twice a second.
- Does the cat do the same thing every time?
- No. Actions are sampled from the network's scores and its hidden state carries over, so behaviour never repeats exactly.