A framework based on AlphaZero capable of learning to play single- and two-player games (tic-tac-toe, Connect 4, and CartPole) using a unified model architecture, with self-play training and Monte-Carlo tree search. Both ResNet and MLP network backbones were explored and compared. Built as a team project for the course Deep Learning: Architectures & Methods at TU Darmstadt.

The cover shows a self-play game between two copies of the trained agent (600 MCTS simulations per move, ending in a hard-fought draw) and the search policy at the game’s most contested decision.

Technologies: PyTorch, AlphaZero, self-play RL