Chessko · Machine learning
Machine learning in a chess engine: logistic regression, a k-NN book and a UCB1 bandit
Chessko's learned parts are small on purpose: a 16-number evaluation, a book of 33 positions, and a three-armed bandit. They are simple enough to inspect, and they come with a result worth being honest about: the learned evaluation is a better predictor than it is a player.
The learned evaluation: 16 numbers
The learned evaluation is a logistic regression. A position is turned into 16 features, all measured from the side to move's point of view (own value minus opponent value), so swapping colours cannot change the result:
bias, material for each piece type, center control, activity, king shield, pawn structure, passed pawns, tempo, game phase, and three phase-interaction terms (pawns, rooks and queens weighted by how much of the game is left), so a linear model can say that a rook matters more in the endgame.
The model multiplies each feature by a learned weight, adds them, and squashes the sum into a probability that the side to move wins. That probability is converted back to centipawns so the search can use it. The whole model is 16 weights, about 2 KB of JSON.
Where the training data comes from
Chessko trains on positions from 200 self-play games (7,805 positions after filtering). The labels are not game results: most self-play games are drawn, and the same opening recurs with different outcomes, which makes results a very noisy signal. Instead the labels are distilled from a deeper search: a stronger “teacher” search scores each position, and the model learns to predict that score from the cheap features.
Training runs offline in Python (standard library only): mini-batch stochastic gradient descent on soft targets, 80 epochs, followed by a temperature calibration and neutral-offset fit on a held-out split of games (not positions, so the validation set has no positions from training games).
What the model learned, and how good it is
On the held-out games, the model's log loss is 0.437 against 0.643 for a constant predictor, and it agrees with the teacher on which side is better in about 93% of positions where the teacher has an opinion.
The learned material weights are the most readable part: pawn 0.83, knight 2.05, bishop 2.06, rook 3.31, queen 8.49. The ordering pawn < knight ≈ bishop < rook < queen is the one every chess book teaches, rediscovered from data. The queen weight, in particular, comes out well above the rook's, as expected.
The honest negative result
A better predictor is not automatically a better player. In a quick round robin at search depth 4 (8 games per pairing, fixed openings, both colours), an evaluation that was 50% learned scored 34%, the pure hand-tuned evaluation 56%, and a 25% blend 59%. The sample is small, so the honest reading is that 50% is clearly worse and that 25% versus none is too close to call. That is why the learned model contributes 25% of the blend on every level, with the classical evaluation supplying the rest. There is also a known weak spot: the training set starts at ply 8, so the model extrapolates in the first moves and reads roughly a pawn off there. The eval bar in the game calls a wide band around zero “balanced” for that reason instead of pretending to half-pawn precision.
That is a real limitation of a 16-feature linear model trained on 7,805 positions, and it is why the top two levels hand over to Stockfish.
The k-NN opening book
To keep weak levels from playing the same game every time, Chessko has a small opening book built from the same 200 games: 33 positions, up to 12 plies deep. Lookup is two-stage. First an exact match on the position's hash. If there is none, a k-nearest-neighbours vote: the position is turned into a short vector (material, centre control, phase), the five nearest book positions inside a distance threshold vote for their moves, and each vote is weighted by 1 / (0.15 + distance). Only legal moves are ever returned. All five jelly levels use the book for the first four plies.
The UCB1 difficulty bandit
The third component learns about you, in your browser. After each move you play, Chessko compares the position's score with what it expected and turns the difference into a move-quality number between 0 and 1. A UCB1 bandit with three arms (−1, 0, +1: make it easier, keep, make it harder) treats each arm as a difficulty offset. It uses the upper-confidence-bound rule, average reward plus an exploration bonus that shrinks the more an arm has been tried, to decide which offset fits you best. It suggests a change only when another arm clearly beats the current one, and it can move the level for you if you turn on “auto-goo”. Its counts live in localStorage, so it never leaves your device.
See how the difficulty levels work for what a level actually changes.
Quick answers
- Does Chessko use a neural network?
- The learned evaluation is a logistic regression with 16 weights, not a neural network. The strongest levels use Stockfish 19, which does use an NNUE neural-network evaluation, compiled to WebAssembly.
- What is a UCB1 bandit?
- UCB1 is a multi-armed-bandit algorithm that picks the option with the highest average reward plus an exploration bonus. Chessko uses it to decide whether the difficulty should go down, stay or go up, based on how well you are playing.