AIWiki
Malaysia
Back to all articles
ApplicationsGoDeepMindreinforcement learning

AlphaGo

4 min readUpdated August 2026
AlphaGo
Type
Computer program for the game Go
Developer
DeepMind (Google)
Announced
January 2016 (Nature paper)
Landmark result
Defeated Lee Sedol 4-1, March 2016
Successor
AlphaGo Zero, MuZero
Related
Reinforcement learning, MCTS, neural networks

AlphaGo is a computer program for the board game Go developed by DeepMind, the London-based artificial intelligence company acquired by Google in 2014. In October 2015 it became the first computer program to defeat a professional human Go player without handicap on a full-sized 19×19 board, and in March 2016 it defeated Lee Sedol, a 9-dan professional and winner of 18 international titles, by four games to one in a widely watched match in Seoul.[1][2]

History and Background

Go is an ancient board game originating in China in which two players alternately place black and white stones on a 19×19 grid. The game was long considered a grand challenge for artificial intelligence because a typical game has far more possible board positions than chess, making brute-force search infeasible.[1]

AlphaGo was developed by a DeepMind team led by David Silver and Demis Hassabis, and its results were first published in the journal Nature on 27 January 2016. The paper reported that AlphaGo had defeated Fan Hui, the reigning European Go champion, 5–0 in a private match played in October 2015, the first time a computer program had beaten a professional player without handicap on a full-sized board.[2][3]

In March 2016, AlphaGo played a five-game match against Lee Sedol of South Korea at the Four Seasons Hotel in Seoul. AlphaGo won the first three games, Lee Sedol won the fourth game — after playing move 78, which professional commentators dubbed the "divine move" — and AlphaGo won the final game to take the match 4–1. DeepMind donated the US$1 million prize money to charitable organisations, and the Korea Baduk Association awarded AlphaGo an honorary 9-dan rank.[3]

In December 2016, an upgraded version of the program, playing anonymously under the account name "Master" on online Go servers, won all 60 of its games against leading professionals. In May 2017, at the Future of Go Summit in Wuzhen, China, AlphaGo defeated Ke Jie, then the world number one, 3–0, after which DeepMind announced that AlphaGo would retire from competitive play.[3]

Technology

AlphaGo combined Monte Carlo tree search with two deep neural networks: a policy network that narrows the search space by predicting the most promising moves, and a value network that estimates the probability of winning from a given board position. The networks were trained first by supervised learning on a database of human expert games, and then refined by reinforcement learning through millions of games of self-play. The distributed version of AlphaGo that played Lee Sedol used approximately 1,202 CPUs and 176 GPUs.[1][2]

In October 2017, DeepMind published AlphaGo Zero, a version that learned to play Go entirely from self-play, starting from random play and using no human game data. AlphaGo Zero used a single neural network, trained with reinforcement learning, and defeated the version that had beaten Lee Sedol by 100 games to 0. Its successor program, MuZero, extended the same learning approach to other games without being given the rules.[4][5]

Applications and Impact

AlphaGo demonstrated that deep reinforcement learning could solve problems previously considered beyond the reach of artificial intelligence, and it is widely regarded as a milestone in the modern AI era. Its unconventional "Move 37" in the second game against Lee Sedol, a move that surprised professional commentators, became a widely discussed example of machine creativity. The match was the subject of a 2017 documentary film also titled AlphaGo, and the techniques developed for the program influenced later DeepMind systems such as AlphaZero and MuZero, as well as subsequent research in planning, search, and self-play.[3][5]

See Also

🇲🇾Malaysian Context

In Malaysia, the AlphaGo matches of 2016 are frequently cited in policy discussions, media coverage, and training materials as a turning point that raised public awareness of artificial intelligence capabilities. The National AI Office (NAIO) and the Malaysia Digital Economy Corporation (MDEC) have promoted AI literacy and adoption programmes in which milestones such as AlphaGo are used to illustrate the progress of machine learning and reinforcement learning.[6][8] Malaysian universities offering AI and data science programmes, including Universiti Malaya and Universiti Teknologi Malaysia, teach reinforcement learning and Monte Carlo tree search — the techniques underlying AlphaGo — as core topics, and local AI communities reference the system in meetups and hackathon materials as an introduction to modern AI.[7]

References

  1. Google Research Blog. (2016, January 27). AlphaGo: Mastering the ancient game of Go with Machine Learning. https://research.google/blog/alphago-mastering-the-ancient-game-of-go-with-machine-learning/
  2. Silver, D., Huang, A., Maddison, C. J., et al. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529, 484–489. https://www.nature.com/articles/nature16961
  3. Wikipedia. AlphaGo versus Lee Sedol. https://en.wikipedia.org/wiki/AlphaGo_versus_Lee_Sedol
  4. Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). Mastering the game of Go without human knowledge. Nature, 550, 354–359. https://www.nature.com/articles/nature24270
  5. DeepMind. AlphaGo Zero: Starting from scratch. https://deepmind.google/discover/blog/alphago-zero-starting-from-scratch/
  6. AIWiki Malaysia. National AI Office (Malaysia). https://aiwiki.com.my/wiki/national-ai-office-malaysia
  7. AIWiki Malaysia. Malaysia AI Talent. https://aiwiki.com.my/wiki/malaysia-ai-talent
  8. MDEC - Malaysia Digital Economy Corporation. https://www.mdec.my/