Self Improving Rat: an artificial-life reinforcement-learning agent in C++ with no ML framework
Self Improving Rat is a persistent artificial-life study: a recurrent reinforcement-learning agent with a body (homeostasis), a world model, curiosity and episodic memory, living in a procedurally generated maze. It learns online and saves its whole state between sessions. It is a single-threaded C++20 desktop application with an SDL2 renderer and needs no GPU, ML framework or network. It is not a claim of general intelligence.
- Platform
- Linux (tested on Ubuntu 26.04)
- Language
- C++20
- Licence
- Not stated in the repository
- Requires
- cmake 3.16+, g++ and make or ninja
- libsdl2-dev (the test suite needs no SDL)
- Run
./run.shfrom the repository folder
How do I run it?
./run.sh # configure, build, and run the application
./run.sh --test # build and run the test suite (no SDL needed)
./run.sh --clean # remove the build directory only
./run.sh --sanitize # build and run with AddressSanitizer/UBSan
./run.sh --install-deps # sudo apt-get install libsdl2-dev (opt-in)
It needs 64-bit Ubuntu (tested on 26.04 LTS), cmake 3.16 or later, g++ with C++20, make or ninja, the SDL2 runtime and the SDL2 development files. Optional environment hooks include SIR_SEED=N for a fixed seed, SIR_HEADLESS=1 for no window, SIR_MAX_STEPS=N to exit after N steps and SIR_SCREENSHOT=path.bmp to save a frame. ./run.sh never touches data/checkpoints or data/logs.
Does the rat actually improve?
The README answers this carefully rather than promising it. The rat reliably attempts to improve: its parameters change from experience. Long-term, monotonic growth is not guaranteed, and maze navigation with sparse, moving rewards is hard for the built-in learner. The reported measurements are:
| Check | Result reported in the README |
|---|---|
| Matched online benchmark: 100,000 simulation steps per seed, seeds 1, 2, 3, 42 and 12345, fresh checkpoints, unchanged maze difficulty | 186 cheeses in total (mean 37.2) before the update the README describes; 1,167 (mean 233.4) after it, which is 6.27 times as many |
| What that comparison covers | The whole system, including wall masks, exploration and learning; it does not isolate neural learning |
Frozen evaluation (sir_eval) | Saved weights on fresh seeds with training disabled, compared with a legal random walk and an untrained network with the same memory-assisted exploration |
The README adds that online reward alone is not evidence of reliable generalisation.
What is inside the agent?
The agent combines a body with homeostatic variables, a world model, curiosity, episodic memory, structural plasticity and lifelong development. Its learner saves recent replay, reconstructs recurrent sequences, avoids observed walls and uses episodic action counts to explore. Resource use is bounded: a preallocated replay buffer of 4,096 transitions, 256 episodic memory entries, a 1,024-slot novelty table, a size-rotated log with at most two files, and a default replay snapshot of about 1.31 MiB.
What are the documented limitations?
- Q-learning with a small GRU adapts slowly when the cheese is re-placed after every collection, and the learned greedy policy can still stall. Long-horizon navigation in large mazes such as 47×31 is not reliably learned within practical run times.
- There is no guaranteed intelligence growth, only persistent adaptation attempts with bounded, verified learning dynamics.
- Homeostasis, curiosity and "satisfaction" are deterministic numeric state, not feelings; the project makes no consciousness claims.
- Cross-platform determinism is not guaranteed; it is asserted within one platform and build.
Quick answers
- Does Self Improving Rat need a GPU or an ML framework?
- No. It is a single-threaded C++20 application with an SDL2 renderer, with no GPU, ML framework or network requirement.
- Does the rat remember between sessions?
- Yes. It saves its whole organism state to disk and learns online, so what it has learned carries over between runs.
- Is it a claim of artificial general intelligence?
- No. The README calls it a self-contained artificial-life study and states that it makes no claim of general intelligence or consciousness.