A completed AlphaZero self-play Connect Four game with move numbers, next to the MCTS policy distribution of a mid-game decision

Reinforcement Learning Agent in Games (AlphaZero)

AlphaZero-based framework that learns tic-tac-toe, Connect 4, and CartPole with a unified model architecture.