[{"data":1,"prerenderedAt":29},["ShallowReactive",2],{"i-lucide:grip":3,"i-lucide:chevron-right":8,"i-lucide:moon":10,"i-lucide:sun":12,"i-lucide:languages":14,"i-lucide:chevron-down":16,"i-lucide:shield-check":18,"i-lucide:mail":20,"blog-body-go-mcts-algorithm-en":22,"i-lucide:cpu":23,"i-lucide:code":25,"i-lucide:star":27},{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":7},0,24,false,"\u003Cg fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\">\u003Ccircle cx=\"12\" cy=\"5\" r=\"1\"\u002F>\u003Ccircle cx=\"19\" cy=\"5\" r=\"1\"\u002F>\u003Ccircle cx=\"5\" cy=\"5\" r=\"1\"\u002F>\u003Ccircle cx=\"12\" cy=\"12\" r=\"1\"\u002F>\u003Ccircle cx=\"19\" cy=\"12\" r=\"1\"\u002F>\u003Ccircle cx=\"5\" cy=\"12\" r=\"1\"\u002F>\u003Ccircle cx=\"12\" cy=\"19\" r=\"1\"\u002F>\u003Ccircle cx=\"19\" cy=\"19\" r=\"1\"\u002F>\u003Ccircle cx=\"5\" cy=\"19\" r=\"1\"\u002F>\u003C\u002Fg>",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":9},"\u003Cpath fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\" d=\"m9 18l6-6l-6-6\"\u002F>",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":11},"\u003Cpath fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\" d=\"M20.985 12.486a9 9 0 1 1-9.473-9.472c.405-.022.617.46.402.803a6 6 0 0 0 8.268 8.268c.344-.215.825-.004.803.401\"\u002F>",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":13},"\u003Cg fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\">\u003Ccircle cx=\"12\" cy=\"12\" r=\"4\"\u002F>\u003Cpath d=\"M12 2v2m0 16v2M4.93 4.93l1.41 1.41m11.32 11.32l1.41 1.41M2 12h2m16 0h2M6.34 17.66l-1.41 1.41M19.07 4.93l-1.41 1.41\"\u002F>\u003C\u002Fg>",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":15},"\u003Cpath fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\" d=\"m5 8l6 6m-7 0l6-6l2-3M2 5h12M7 2h1m14 20l-5-10l-5 10m2-4h6\"\u002F>",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":17},"\u003Cpath fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\" d=\"m6 9l6 6l6-6\"\u002F>",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":19},"\u003Cg fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\">\u003Cpath d=\"M20 13c0 5-3.5 7.5-7.66 8.95a1 1 0 0 1-.67-.01C7.5 20.5 4 18 4 13V6a1 1 0 0 1 1-1c2 0 4.5-1.2 6.24-2.72a1.17 1.17 0 0 1 1.52 0C14.51 3.81 17 5 19 5a1 1 0 0 1 1 1z\"\u002F>\u003Cpath d=\"m9 12l2 2l4-4\"\u002F>\u003C\u002Fg>",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":21},"\u003Cg fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\">\u003Cpath d=\"m22 7l-8.991 5.727a2 2 0 0 1-2.009 0L2 7\"\u002F>\u003Crect width=\"20\" height=\"16\" x=\"2\" y=\"4\" rx=\"2\"\u002F>\u003C\u002Fg>","\u003Cblockquote>\n\u003Cp>Go is hard not only because the rules are intricate, but because \u003Cstrong>an empty board already offers hundreds of nearly interchangeable legal points\u003C\u002Fstrong>. Exhaustive search to the end is impossible, and hand-written evaluators are easy to get wrong. MCTS charts a practical middle path: \u003Cstrong>many random games, with compute steered toward more promising branches\u003C\u002Fstrong>.\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Cp>\u003Cimg src=\"\u002Fblog\u002Fgo-mcts-algorithm\u002Fcover.webp\" alt=\"On an empty Go board branching explodes; MCTS concentrates compute onto a few promising points via random playouts\">\u003C\u002Fp>\n\u003Ch2>Why Is Go Especially Suited to—and Dependent on—MCTS?\u003C\u002Fh2>\n\u003Cp>Go's search space is too large to &quot;calculate to the end&quot;: an empty 19×19 board has about 361 empty intersections, legal moves shift with the position, and games often run past a hundred moves. If each turn still has dozens of reasonable candidates, both width and depth explode. Hand-written heuristics (capture first, take corners, connect liberties) often degrade into aimless crawling along the edge in open openings.\u003C\u002Fp>\n\u003Cp>MCTS (Monte Carlo Tree Search) does not try to exhaust the tree, and it does not rely on a global scoring table. It repeatedly plays short random games, backs up wins and losses, and spends more simulation budget on branches that look better. The larger the branching factor and the harder it is to write an accurate evaluator, the more valuable this &quot;sample instead of exhaust&quot; idea becomes—which is why it caught on in Go and later in many imperfect-information games.\u003C\u002Fp>\n\u003Ch2>What Does One MCTS Iteration Do?\u003C\u002Fh2>\n\u003Cp>A full MCTS iteration usually has four steps, repeated until time or simulation count runs out:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Selection\u003C\u002Fstrong>: From the root, walk down expanded children with a formula such as UCB1, balancing &quot;known high win rate&quot; against &quot;not tried enough yet&quot;;\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Expansion\u003C\u002Fstrong>: When a node still has untried legal moves, pick one (randomly or with a prior) and grow a new child;\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Simulation \u002F Rollout\u003C\u002Fstrong>: From the new position, both sides play randomly (or with a weak heuristic) to the end and score under the rules;\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Backpropagation\u003C\u002Fstrong>: Write that game's result back along the path into each node's visits \u002F wins.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>The move finally played is usually the \u003Cstrong>root child with the most visits\u003C\u002Fstrong> (argmax visits), not the one with the highest instantaneous win rate—visit count already encodes &quot;tried many times and still standing,&quot; which is stabler than early, noisy win rates.\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Step\u003C\u002Fth>\n\u003Cth>Input\u003C\u002Fth>\n\u003Cth>Output\u003C\u002Fth>\n\u003Cth>Common pitfall\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>Selection\u003C\u002Ftd>\n\u003Ctd>Expanded tree + UCB1\u003C\u002Ftd>\n\u003Ctd>Path to a leaf \u002F node to expand\u003C\u002Ftd>\n\u003Ctd>Exploration constant too high → flat visits\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Expansion\u003C\u002Ftd>\n\u003Ctd>Untried legal moves\u003C\u002Ftd>\n\u003Ctd>One new child\u003C\u002Ftd>\n\u003Ctd>Suicide \u002F eye-fill at the root poisons the endgame\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Simulation\u003C\u002Ftd>\n\u003Ctd>Current position\u003C\u002Ftd>\n\u003Ctd>One game outcome\u003C\u002Ftd>\n\u003Ctd>Filling own eyes during rollouts corrupts value\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Backprop\u003C\u002Ftd>\n\u003Ctd>Outcome\u003C\u002Ftd>\n\u003Ctd>Updated visits\u002Fwins on the path\u003C\u002Ftd>\n\u003Ctd>Keep the perspective fixed (always relative to the same side)\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Ch2>How Does UCB1 Trade Off Exploitation and Exploration?\u003C\u002Fh2>\n\u003Cp>UCB1 scores each child roughly as:\u003C\u002Fp>\n\u003Cp>[\n\\frac{w_i}{n_i} + C \\sqrt{\\frac{\\ln N}{n_i}}\n]\u003C\u002Fp>\n\u003Cp>The first term is empirical win rate (exploitation); the second grows when a node has been visited little (exploration). Larger (C) means more willingness to try unfamiliar branches.\u003C\u002Fp>\n\u003Cp>Textbooks often set (C=\\sqrt{2}\\approx 1.41), derived for rewards in ([0,1]) with not too many children. A Go root often has dozens or hundreds of legal points, while a browser think may only run thousands to tens of thousands of simulations: the exploration term easily swamps win-rate gaps, so \u003Cstrong>visits become almost uniform and the top move can fall to about 2%\u003C\u002Fstrong>—poor move choice and unreadable &quot;tendency&quot; at the root. In practice (C) is often lowered (e.g. to 0.4) so a limited budget concentrates faster.\u003C\u002Fp>\n\u003Cp>A useful convergence signal is the root top-1 visit share. On the same 9×9 position, ~1200 simulations often leave top-1 in the single-digit percent range; ~20,000 can reach about 17%; ~40,000 about 40%. When the share is low, prefer &quot;not enough search&quot; over &quot;every point on the board is equally good.&quot;\u003C\u002Fp>\n\u003Ch2>How Do &quot;Tendency&quot; Scores Relate to the Move Played?\u003C\u002Fh2>\n\u003Cp>Each root child's \u003Ccode>visits \u002F Σvisits\u003C\u002Fcode> can be read as relative tendency: where search spent compute. A UI that shows top-N candidates usually sorts by visits and truncates; the move played is still the visits maximum.\u003C\u002Fp>\n\u003Cp>Boundaries that are easy to miss:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Displayed probabilities often sum to less than 1\u003C\u002Fstrong>: after filtering the long tail, the shown set need not sum to 1—on purpose. Top-1 share itself says how sure the search is;\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Too few simulations → flat distribution\u003C\u002Fstrong>: thousands of playouts spread across a hundred empties on 19×19 leave only dozens of visits per point—tendency is near noise;\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Do not weaken play with temperature sampling on an unconverged distribution\u003C\u002Fstrong>: sampling by visits when the distribution is still flat pushes it toward uniform and collapses into random play. A cleaner way to weaken is \u003Cstrong>fewer simulations\u003C\u002Fstrong>, still picking argmax.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Why Can't You Speed Up the Opening by &quot;Thinking Less&quot;?\u003C\u002Fh2>\n\u003Cp>An empty board has the most legal points and the fewest simulations per point—globally the hungriest stage for compute. A tempting reading: &quot;Opening top-1 is only a few percent, so points are similar—search less.&quot; Self-play disagrees: the side that cuts opening budget against the same-strength engine can lose about 2:8. The right reading is \u003Cstrong>low share = not converged yet\u003C\u002Fstrong>, not &quot;already equivalent.&quot;\u003C\u002Fp>\n\u003Cp>A safer speedup is \u003Cstrong>opening knowledge that narrows root candidates\u003C\u002Fstrong>: restrict untried moves at the root to a few well-established points (e.g. still-empty corner star points). Branching drops from hundreds to a handful, and each candidate can get hundreds of simulations. When corners are taken, contact starts near those points, or a group drops to one liberty, fall back to full-board search. On large boards this usually saves time and raises per-point signal; on 9×9, where branching is already modest, forced narrowing can be a net loss—gate by board size.\u003C\u002Fp>\n\u003Ch2>Why Explicitly Avoid Filling Your Own Eyes in the Endgame?\u003C\u002Fh2>\n\u003Cp>Playing in your own true eye is usually legal but terrible—it kills liberties. If random rollouts fill eyes freely, valuations are polluted by self-destruction. If the root does not exclude own eyes, a side with nothing left to play may fill an eye and kill a living group instead of passing.\u003C\u002Fp>\n\u003Cp>A more mature pure-MCTS Go engine handles both places: rollouts skip eye fills; root candidates filter own eyes. When only eye fills remain, the correct action is \u003Cstrong>pass\u003C\u002Fstrong>, so both sides can move into scoring.\u003C\u002Fp>\n\u003Ch2>Where Are the Limits of Pure MCTS?\u003C\u002Fh2>\n\u003Cp>Boundaries matter more than memorizing the formula:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Dimension\u003C\u002Fth>\n\u003Cth>Pure MCTS (random rollouts)\u003C\u002Fth>\n\u003Cth>MCTS with policy\u002Fvalue nets\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>Prior\u003C\u002Ftd>\n\u003Ctd>Almost none; concentrate via visits\u003C\u002Ftd>\n\u003Ctd>Network move priors + position value\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Strength at equal time\u003C\u002Ftd>\n\u003Ctd>Hard-capped by sim count; noisy on large boards\u003C\u002Ftd>\n\u003Ctd>Usually much stronger\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Readable &quot;tendency&quot;\u003C\u002Ftd>\n\u003Ctd>Directly from visit shares\u003C\u002Ftd>\n\u003Ctd>Mixed with network prior—read separately\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Implementation \u002F compute\u003C\u002Ftd>\n\u003Ctd>Feasible in a browser Web Worker\u003C\u002Ftd>\n\u003Ctd>Needs model weights and more compute\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Weakening difficulty\u003C\u002Ftd>\n\u003Ctd>Lower \u003Ccode>maxSims\u003C\u002Fcode> \u002F time\u003C\u002Ftd>\n\u003Ctd>Also tune temperature, noise, etc.\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>A few more engineering boundaries: without a network, short 19×19 searches yield tendency that is reference-only; Chinese area scoring and Japanese territory scoring define different terminals—playout scoring must match product rules; ko and suicide legality must be correct in move generation, or the tree learns illegal shortcuts.\u003C\u002Fp>\n\u003Ch2>Takeaway\u003C\u002Fh2>\n\u003Cp>MCTS spends limited compute on more promising branches via \u003Cstrong>selection → expansion → simulation → backprop\u003C\u002Fstrong>, and usually plays the most-visited move. On Go's wide root, the UCB1 exploration constant, whether to narrow opening candidates, and whether rollouts forbid eye fills decide strength and readability under real budgets. When you read visit shares, ask first how many simulations ran—before convergence, a flat distribution means too little compute, not that every empty point is a good move.\u003C\u002Fp>\n",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":24},"\u003Cg fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\">\u003Cpath d=\"M12 20v2m0-20v2m5 16v2m0-20v2M2 12h2m-2 5h2M2 7h2m16 5h2m-2 5h2M20 7h2M7 20v2M7 2v2\"\u002F>\u003Crect width=\"16\" height=\"16\" x=\"4\" y=\"4\" rx=\"2\"\u002F>\u003Crect width=\"8\" height=\"8\" x=\"8\" y=\"8\" rx=\"1\"\u002F>\u003C\u002Fg>",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":26},"\u003Cpath fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\" d=\"m16 18l6-6l-6-6M8 6l-6 6l6 6\"\u002F>",{"left":4,"top":4,"width":5,"height":5,"rotate":4,"vFlip":6,"hFlip":6,"body":28},"\u003Cpath fill=\"none\" stroke=\"currentColor\" stroke-linecap=\"round\" stroke-linejoin=\"round\" stroke-width=\"2\" d=\"M11.525 2.295a.53.53 0 0 1 .95 0l2.31 4.679a2.12 2.12 0 0 0 1.595 1.16l5.166.756a.53.53 0 0 1 .294.904l-3.736 3.638a2.12 2.12 0 0 0-.611 1.878l.882 5.14a.53.53 0 0 1-.771.56l-4.618-2.428a2.12 2.12 0 0 0-1.973 0L6.396 21.01a.53.53 0 0 1-.77-.56l.881-5.139a2.12 2.12 0 0 0-.611-1.879L2.16 9.795a.53.53 0 0 1 .294-.906l5.165-.755a2.12 2.12 0 0 0 1.597-1.16z\"\u002F>",1786455161836]