AIFEATURE

What Is Reinforcement Learning? How AI Agents Learn From Trial and Error, and What AlphaGo Proved

林紀旭 James LinEditor-in-Chief
Published · Updated
Reinforcement learning trains autonomous agents to make decisions by interacting directly with their environment through trial and error, without labeled examples. Agents follow a policy that maps observations to actions, with the sole objective of maximizing cumulative reward. Google DeepMind's AlphaGo demonstrated the approach concretely, winning 5-0 in its first match against a Go professional in 2015 and 4-1 in Seoul in March 2016, a match watched by over 200 million people worldwide.

What is reinforcement learning, and how do agents learn through interaction with an environment?

Reinforcement learning is a machine learning process in which autonomous agents learn to make decisions by interacting with their environmentCITE:E1. Rather than following instructions, the agent learns to perform a task by trial and error, with no guidance from a human user during the processCITE:E2.

How does reinforcement learning fundamentally differ from supervised learning?

Reinforcement learning does not rely on labeled examples of correct or incorrect behavior, which is what separates it from supervised learningCITE:E3. Where supervised learning trains on pre-tagged data, a reinforcement learning agent instead builds its understanding through its own trial-and-error interactions with the environmentCITE:E2.

What is the agent's objective and the core decision-making mechanism in reinforcement learning?

An agent's only objective is to maximize its cumulative rewards from the environmentCITE:E4. It pursues that objective through a policy — a function that maps the agent's observations of the environment to the action it will take nextCITE:E5.

How did AlphaGo demonstrate reinforcement learning's real-world breakthrough?

AlphaGo, developed by Google DeepMind, won the first-ever match between an AI system and a Go professional by a score of 5-0 in 2015CITE:E6. It followed that result with a 4-1 victory in Seoul, South Korea, in March 2016, a match watched by more than 200 million people worldwideCITE:E7.

MatchDateResultViewers
First AI vs. Go professional match2015AlphaGo won 5-0Not reported
Seoul, South Korea matchMarch 2016AlphaGo won 4-1Over 200 million worldwide

Taken together, the two results show a consistent pattern: an agent operating without labeled instruction, guided only by a policy that maps observations to actions and an objective to maximize cumulative reward, won both its first-ever professional match and its highest-profile public matchCITE:E4CITE:E5CITE:E6CITE:E7. The trial-and-error framework defined by IBM and the policy mechanism defined by Google describe, in the abstract, the same process that Google DeepMind's own results describe in practiceCITE:E1CITE:E5CITE:E7.

📊 Evidence

FAQ

What is reinforcement learning, and how do agents learn through interaction with an environment?

Reinforcement learning is a machine learning process in which autonomous agents learn to make decisions by interacting with their environmentCITE:E1.

How does reinforcement learning fundamentally differ from supervised learning?

Reinforcement learning does not rely on labeled examples of correct or incorrect behavior, which is what separates it from supervised learningCITE:E3.

What is the agent's objective and the core decision-making mechanism in reinforcement learning?

An agent's only objective is to maximize its cumulative rewards from the environmentCITE:E4.

How did AlphaGo demonstrate reinforcement learning's real-world breakthrough?

AlphaGo, developed by Google DeepMind, won the first-ever match between an AI system and a Go professional by a score of 5-0 in 2015CITE:E6.

📎 Sources

  1. ibm.com
  2. developers.google.com
  3. deepmind.google

Related data

Author's Take林紀旭 James Lin

The mechanism worth focusing on here is the policy: a function mapping observations directly to actions, paired with a single objective of maximizing cumulative reward, and no labeled examples to lean on. That framing explains why AlphaGo's two results are worth reading together rather than separately — a 5-0 sweep in the first-ever professional match and a 4-1 win in the far more scrutinized Seoul match, watched by over 200 million people, are two data points from the same trial-and-error process, not two different techniques. The indicator worth watching next is consistency: whether a policy trained this way keeps producing decisive, lopsided results as the scale and visibility of the contest increases, the way it did between 2015 and March 2016.

林紀旭 James LinEditor-in-Chief

Related

BRIEF

Google to Invest €13 Billion in Finnish AI Infrastructure Through 2028

Google announced a €13 billion ($15.1 billion) investment in Finnish AI infrastructure to be deployed in 2027–2028, its largest single investment in Europe. The plan covers new data centers in Kajaani, Muhos, and Vaala plus an expansion in Hamina, backed by a 22-year power deal for 50% of Loviisa nuclear plant output and projected to support 37,000 construction-phase jobs and 7,000 permanent roles.

EffectStory 編輯部 ·
BRIEF

Chang Hwa Bank's Jan–Aug Profit Hits Record NT$15.57 Billion, Up 23.21% YoY on Loan and Wealth Management Strength, EPS NT$1.29

Chang Hwa Bank (TWSE: 2801) reported cumulative after-tax profit of NT$15.567 billion for January through August, up 23.21% year-on-year with EPS of NT$1.29, a record for the period. The bank attributed the gain to loan, deposit, and wealth management momentum, and August alone delivered NT$2.916 billion in profit, up 29.16% year-on-year.

EffectStory 編輯部 ·