Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

AlphaGo did not win by searching every possible Go game, nor by letting a neural network pick the move with the highest probability. The 2016 system combined Monte Carlo Tree Search (MCTS) with learned policy and value networks, fast rollout evaluation, reinforcement learning, and large-scale parallel computation. The network supplied intuition about promising moves; search tested those moves against plausible futures.

That division of labor made a strategically difficult game tractable. MCTS built only a selective partial tree, repeatedly concentrating computation where the current evidence suggested it could matter most.

What Monte Carlo Tree Search does

MCTS incrementally grows a tree of possible actions instead of constructing an entire game tree in advance. Each simulation starts at the current position, follows a selected path, evaluates a new position, and uses the result to improve future selections.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Traditional MCTS is often explained with four stages:

#1 Best Overall
YMI Magnetic Travel Go Game Set, 11-Inch Compact 19x19 Board, 361 Stones
  • Magnetic Stones Stay Put: 181 black and 180 white magnetic single convex plastic stones (361 total, each 5 x 12.5 millimeters) cling to the board through bumps, tilts, and travel. Packaged in two plastic bowls that tuck inside the folded case.
  • Sized for Carrying Around: Open, the board measures 11 x 11 x 0.6 inch (28.5 x 28.5 x 1.6 centimeters). Folded, it's a compact 11.2 x 5.7 x 1.2 inches (28.5 x 14.5 x 3 centimeters), great for beginners or games on the go. If you want a larger board for regular home play, check our full size Go sets instead.
  • Grab and Go Design: Quality plastic construction with a folding hinge for quick setup on a table, floor, or countertop in seconds. No assembly, no loose parts to track down.
  • Lightweight and Portable: The complete set weighs just 1.72 pounds (0.78 kilograms), light enough for a bag, backpack, or car.
  • A Game Worth Learning: Go is one of the world's oldest strategy games, easy to pick up in an afternoon but deep enough to for a lifetime of rewarding play.
  1. Selection: begin at the root and choose child nodes using an exploration–exploitation rule such as UCT. High-value moves are exploited, while under-visited moves are explored.
  2. Expansion: add one or more previously unexplored legal actions.
  3. Simulation or evaluation: estimate the new position. A classic implementation plays a fast rollout, using random or heuristic moves until the game ends. AlphaGo added learned evaluators and policy guidance to make this estimate substantially more informative.
  4. Backup: propagate the result back along the visited path, updating visit counts and value estimates.

After many simulations, the root’s search statistics indicate which action has the strongest support. The exact selection score differs among UCT, the original AlphaGo, AlphaGo Zero, and AlphaZero; they should not be treated as one interchangeable algorithm.

MCTS is not minimax

Minimax with alpha–beta pruning systematically searches alternating best responses and normally depends on a position-evaluation function. MCTS samples a subset of possibilities and improves statistical estimates as its simulation budget grows. This makes it useful when exhaustive search is infeasible and a simulator or known rules engine is available.

Why Go defeated conventional search

Go combines a large branching factor with long tactical and strategic horizons. A move can create influence across the board, sacrifice local territory for global initiative, or change the value of a fight many moves later. Concepts such as sente, thickness, shape, and territorial balance are difficult to encode with a small set of reliable hand-written rules.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DeepMind describes Go’s state space as approximately 10170 board configurations. That is an order-of-magnitude scale estimate, not an exact count of all legal games. The practical problem is not only the number of positions: a locally attractive move can be globally wrong, and nonterminal positions are hard to evaluate.

Chess also has a vast search space, but decades of work produced comparatively effective handcrafted evaluation functions and alpha–beta techniques. For Go, evaluating a position well was itself a central research challenge. More raw search without a strong guide would spend most of its effort on unpromising variations.

How original AlphaGo combined MCTS and neural networks

The system described in the 2016 Nature paper used several learned components rather than a single generic “Go network.” Its search pipeline can be summarized as follows:

Current board
     ↓
Policy network → plausible candidate moves
     ↓
MCTS expands selected branches
     ↓
Value network and rollout evaluation assess leaves
     ↓
Results are backed up through the tree
     ↓
Choose the move with the strongest search statistics

The root node represented the position to play. AlphaGo did not expand every legal move equally. A policy model supplied probabilities that concentrated simulations on plausible actions. A value model estimated the chance of eventual victory from positions reached in the tree, while the original system also used a fast rollout policy for additional evaluation. MCTS aggregated these imperfect signals across many continuations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The policy network: learned priors for moves

AlphaGo’s supervised policy network learned to predict expert moves from approximately 30 million positions in human games. Google reported 57% expert-move prediction accuracy, compared with a 44% previous record cited in its contemporaneous account (Google’s AlphaGo explanation). That was a historical move-prediction benchmark, not a measure of match-winning strength.

The network had two jobs. First, it learned a useful prior over legal moves. Second, it reduced the effective branching factor during search by directing simulations toward moves likely to be meaningful. A policy prior is guidance, not a command: an unusual move can still gain support if deeper search shows that it works.

Rank #2
19 x19 Folding Go Game Set Board with and Bamboo Bowls and Imitation Jade Go Pieces。
  • Chess board - easy to fold in half, convenient for compact storage, easy to carry, can play chess with family and friends when traveling or camping, without worrying about the complex Go game set, the standard game size is 19x19, 22X24mm grid. The board size is 18.71 x 17.33 x 0.98 inches (47.5 x 44 x 2.5 cm). The folding size is 17.33 x 9.45 x 1.97 inches (44 x 24 x 5 cm).
  • Go pieces are made of imitation jade. The white chess pieces are smooth imitation white jade. The black chess pieces are smooth, round and tactile. The chess pieces are stronger and not easily damaged. The size of chess pieces is 2.2x2.2 cm (0.86 x 0.86 inches), 180 white chess pieces, 181 black chess pieces, 10 white chess pieces and 10 black chess pieces
  • Packaging - professionally designed printed packaging that can be used as an educational tool for children in the classic Go game or as a gift for children's elders.
  • We have presented a guide to the primary Go game for beginners to understand the rules of the game.

The value network: judging positions before the end

The value network estimated the expected outcome from a board position without requiring every continuation to reach a terminal result. This reduced the noise and expense of relying only on random or weakly guided rollouts.

A value estimate is not an oracle. It can be miscalibrated, extrapolate poorly outside its training distribution, or be contradicted by deeper tactical evidence. Search and the value network compensate for different weaknesses: the network provides a fast strategic estimate, while MCTS tests that estimate against concrete future sequences.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why the final move was not simply the network’s top prediction

The policy network’s highest-probability move was only a starting hypothesis. After repeated simulations, AlphaGo selected using the tree’s accumulated search statistics. This allowed search to discover moves that were less likely under the raw policy but stronger after their consequences had been examined.

The four stages inside AlphaGo-style search

1. Selection

Starting at the root, the search chose child edges according to a score balancing estimated value, exploration, and (in neural-guided variants) the policy prior. Well-performing moves received more visits, but insufficiently explored moves retained a chance to be tested.

2. Expansion

When the path reached an unexpanded position, AlphaGo added legal child moves, concentrating on candidates favored by its policy guidance rather than treating the full legal move set as equally urgent.

3. Evaluation

The original system combined neural value estimates with rollout-based components. A rollout was a fast approximate continuation; it was not a claim that every branch was random. Learned policies and evaluators made the simulations highly structured compared with textbook random-play MCTS.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Backup

The estimated outcome was propagated back through the path, updating each edge’s visit count and value statistics from the appropriate player perspective. Later simulations used those updated statistics, progressively allocating more computation to promising variations.

The following pseudocode captures the teaching idea, not every implementation detail of the proprietary 2016 system:

function MCTS(root_state, simulations):
    root = create_node(root_state)

    repeat simulations times:
        node = root
        path = []

        while node is fully expanded and not terminal:
            edge = select_child_using_search_score(node)
            path.append(edge)
            node = edge.child

        if node is terminal:
            value = terminal_outcome(node.state)
        else:
            policy, value = neural_network(node.state)
            expand(node, policy)

        backup(path, value)

    return action_with_highest_visit_count(root)

The selection score, network arrangement, and evaluation step vary by system. Production implementations also add batching, parallel workers, virtual losses, memory management, symmetry handling, and hardware-specific inference optimizations.

Rank #3
Go Game Set 11 inches - Foldable, Portable and Travel Set, (19 x 19) Strategic Magnetic Board Game (weiqi)
  • The Go game set (19 x 19) is a foldable travel Go game set with all plastic stones designed with magnetism.
  • The Go set includes 181 black and 180 white magnetic plastic stones, each placed in 2 separate bowls. The size of the chessboard is 11.6 x 11.2 x 0.59 inches (29.5 x 28.5 x 1.5 centimeters).
  • The magnetic Go set is made of high-quality plastic, convenient storage bowl, durable, smooth, and long-lasting, with sturdy hinges.
  • Chessboard - easy to fold, compact storage, easy to carry, can play chess with family and friends while traveling or camping.
  • The whole set weighs 1.5 pounds (0.68 kilograms).

Why the combination worked

  • Prioritization: policy probabilities focused computation on plausible moves.
  • Look-ahead: MCTS examined tactical and strategic consequences instead of trusting a single prediction.
  • Nonterminal evaluation: the value network supplied a stronger estimate than a shallow or random rollout alone.
  • Policy improvement: the distribution of search visits could be better than the network’s unsearched move distribution.
  • Adaptation: search could elevate an unconventional move when its continuations justified it.

The central lesson is complementarity. A neural network generalizes from positions but can miss a concrete refutation. Search exposes consequences but is too expensive to direct uniformly and needs a useful evaluator. AlphaGo made the two capabilities reinforce one another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How original AlphaGo was trained

Supervised learning from expert games

The initial policy model learned to imitate strong human play from expert records. This provided a practical starting distribution over moves instead of beginning with an uninformed search.

Reinforcement learning through self-play

A policy improved by playing against earlier versions of itself. The objective was no longer to reproduce a human move; it was to increase the probability of winning. This stage could discover moves that were rare or absent in the training games.

Training the value model

Positions from self-play were used to train a network to estimate eventual outcomes. During search, that estimate replaced the need to complete every branch and gave MCTS a strategically informed leaf evaluation.

Distributed inference and search

The algorithmic loop is short, but AlphaGo’s strength also depended on systems engineering: many simulations, parallel workers, efficient neural inference, memory management, and substantial compute. Reimplementing the four stages on a laptop demonstrates the method; it does not reproduce the scale of the 2016 production system.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Original AlphaGo, AlphaGo Zero, and AlphaZero

“AlphaGo” is often used as an umbrella term for different systems. The following distinctions matter when describing a network, training method, or search rule.

Feature Original AlphaGo AlphaGo Zero AlphaZero
Human game records Used expert games for initial supervised policy training Did not require human game records Did not require human game records for its self-play approach
Network design Separate policy and value networks, with rollout components One network with policy and value outputs One general framework with policy and value outputs
Learning Supervised learning followed by reinforcement learning Self-play reinforcement learning from the rules Self-play reinforcement learning across Go, chess, and shogi
Search Neural-guided MCTS supplemented by rollout evaluation Neural-guided MCTS using the network’s policy and value AlphaZero-style neural-guided MCTS, commonly described with PUCT
Main significance Showed the power of combining deep learning and search in Go Showed that self-play could surpass the earlier system without human examples Generalized the approach beyond Go

DeepMind’s AlphaGo Zero account and the Nature paper describe a single network producing both outputs. “Without human knowledge” is shorthand: AlphaGo Zero still received Go’s rules, legal-action structure, board representation, and win/loss reward.

AlphaZero extended the approach to chess, shogi, and Go with a common self-play, neural-network, and search framework. It is related to AlphaGo Zero, not identical to every AlphaGo implementation. MuZero is a further development that learns an internal model and searches in latent representations; it is outside the core AlphaGo explanation.

UCT, PUCT, and terminology that should not be blurred

Traditional UCT balances exploitation of high-value children against exploration of uncertain ones. AlphaGo-style systems add learned move priors, so the search no longer treats all legal actions as equally promising.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
YMI Magnetic 19x19 Go Game Set, 14.6 in
  • Large And Portable: Grab and go with this foldable travel Go game set that measures 14.6 x 14.6 x 1.1 inches (37.1 x 37.1 x 2.8 centimeters) with a 19 x 19 standard playing field
  • Perfect Beginner Set: High-quality plastic, durable hinges, and convenient storage bowls keep the Go Stones in great shape, and the board lays flat after unfolding
  • Magnetic Single Convex Stones: This Go board and stones set includes 181 black magnetic and 180 white magnetic stones for calculated moves that stay put until the very end; Stones measure 6 x 17 millimeters
  • Easy Does It: With everything you need (and nothing you don't weighing you down!) you're ready to play with this magnetic Go game set, anytime, anywhere.
  • Entire Set Weighs 3.3lbs (1.5kg)

OpenSpiel’s documentation distinguishes standard MCTS, which can use uniform priors and rollout values, from AlphaZero-style search, which uses neural policy priors and value estimates with a PUCT-style rule. Do not label every version of AlphaGo “PUCT” without specifying the system and implementation being discussed.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What AlphaGo demonstrated—and what it did not

Demonstrated

  • Deep neural networks can provide useful strategic policy and value estimates in Go.
  • Search can turn approximate learned predictions into stronger action selection.
  • Self-play and reinforcement learning can produce highly capable game policies.
  • Explicit look-ahead remains valuable even when a neural network is very strong.

The original 2016 AlphaGo defeated European champion Fan Hui 5–0 and recorded a reported 99.8% win rate against other Go programs in the benchmark reported by the paper and Google Research (Google Research summary; paper PDF). These are historical results for a particular system, opponents, and match conditions—not a universal win rate for every later AlphaGo-branded model.

Did not demonstrate

  • MCTS alone is generally intelligent.
  • Neural networks always need search.
  • High Go performance proves human-like understanding.
  • A Go-playing system directly solves open-ended real-world reasoning.
  • Success in Go automatically transfers to arbitrary domains.

DeepMind’s 2026 “AlphaGo at 10” retrospective presents AlphaGo as an influential precursor to later systems. That is DeepMind’s interpretation of its legacy, not proof that AlphaGo directly caused every modern AI capability.

Costs, limitations, and failure modes

  • Compute: neural inference and thousands or millions of simulations can be expensive.
  • Diminishing returns: additional simulations do not guarantee better decisions when the value model is systematically wrong.
  • Model dependence: results depend on policy quality, value calibration, exploration settings, simulation budget, and implementation details.
  • Environment assumptions: MCTS is most natural for discrete actions with a simulator or known transition rules. Continuous actions, partial observability, and unavailable dynamics require additional methods.
  • Engineering: batching, parallelism, synchronization, memory use, and evaluation infrastructure can dominate a basic algorithm’s runtime.

Common misconceptions

  1. AlphaGo was not merely playing random games and counting wins; learned policies and evaluators structured its simulations.
  2. AlphaGo Zero did not use two separate policy and value networks at its core; one network produced both outputs.
  3. AlphaGo Zero did not learn from literally nothing; it was given the rules and representation.
  4. Policy prediction accuracy is not the same as playing strength.
  5. The final move was selected from search statistics, not simply the raw top policy probability.

Can you reproduce AlphaGo-style MCTS?

A small implementation can reproduce the ideas, but not the proprietary system’s strength or infrastructure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Start with a known rules engine and a correct legal-move generator.
  2. Implement plain MCTS on Tic-Tac-Toe, Connect Four, or a small Go board.
  3. Add a policy prior so expansion favors plausible actions.
  4. Add a value evaluator and verify player-perspective handling during backup.
  5. Generate self-play games and train a policy/value model.
  6. Only then add batching, parallel workers, virtual losses, and accelerator inference.
  7. Evaluate against fixed opponents with controlled rules, komi, seeds, and search budgets.

OpenSpiel provides open-source game-theory and reinforcement-learning components, including MCTS and AlphaZero-related documentation. It is suitable for education and research experiments, not a drop-in reproduction of AlphaGo’s proprietary production stack.

Implementation checklist

  • Represent board state and player-to-move unambiguously.
  • Generate and validate legal actions, captures, passes, and terminal states.
  • Track whose perspective each backed-up value uses.
  • Choose and document an exploration constant or PUCT parameters.
  • Decide whether final action selection uses visit count, value, or a temperature schedule.
  • Batch neural evaluations when inference becomes the bottleneck.
  • Control random seeds and record ruleset, komi, board size, hardware, and training budget.
  • Compare against fixed opponents rather than relying on a single self-play score.

When cloud compute is useful

Most learners should begin locally. Cloud accelerators become relevant when self-play generation or neural training—not the basic tree algorithm—is the bottleneck.

  • OpenSpiel locally: best for learning and small experiments; no paid plan is indicated in its cited documentation.
  • Google Cloud: GPU and VM charges are usage-based, with storage and networking added separately. The pricing overview and GPU pricing page list current rates and eligibility-dependent new-customer credits; verify them before budgeting.
  • AWS: pricing varies by instance type, region, operating system, and on-demand, reserved, or Spot capacity.
  • Paperspace: its usage-based GPU pricing is oriented toward notebook and ML experimentation, with availability and rates varying by machine type.

Buying accelerator time makes experiments easier; it does not remove the need for a training pipeline, distributed self-play, evaluation design, and substantial engineering. Compare total cost, availability, storage, data transfer, and setup time rather than only an hourly GPU figure.

The enduring lesson

AlphaGo’s breakthrough was not “MCTS plus a neural network” as a slogan. It was a carefully engineered division of labor: learned models generalized from data and supplied priors and evaluations, while selective search checked consequences and improved decisions at inference time. The result showed why approximate intuition and explicit planning can be more powerful together than either one alone in a domain with enormous search and difficult evaluation.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Did AlphaGo search every possible Go move sequence?

No. It explored a selective partial tree, allocating simulations unevenly according to learned move priors, value estimates, and accumulated search statistics.

Was AlphaGo’s policy network the same as its final playing strategy?

No. The policy network guided candidate moves, but MCTS aggregated many continuations and the final action was chosen from search statistics.

Can a basic MCTS project reproduce AlphaGo?

It can demonstrate the core ideas, especially with OpenSpiel, but matching AlphaGo requires neural training, distributed self-play, efficient inference, and substantial compute.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.