COMP 3200 - Intro to Artificial Intelligence - Lecture 02 - Agents, Actions, and Environments
Watch on YouTube →
Overview
Dave Churchill establishes the classical AI framework of agents acting in environments: agents perceive states, choose actions through policies, and are evaluated against explicit performance measures. He connects these definitions to examples including chess, League of Legends, cart-pole, UPS route optimization, and dice, then classifies environments by observability, randomness, time structure, dynamics, and information—properties that determine which algorithms can apply.
Key takeaways
- An AI agent is defined relative to a task: the agent observes an environment, chooses actions, and changes the environment's state in pursuit of a goal.
- A policy π maps states to actions, but whether those actions are rational depends on an explicit performance measure and the agent's available knowledge.
- Decision quality should be judged by expected value rather than a single outcome; a lucky result does not make a low-probability decision sound.
- State-space size and game-tree size are different: chess has roughly 10⁴³ possible configurations, while Go's 19×19 board permits about 10¹⁷² configurations before considering action histories.
- Environment properties—observability, stochasticity, episodic versus sequential structure, dynamics, and discrete versus continuous actions—constrain which AI algorithms are suitable.
- A large configuration space does not guarantee strategic complexity: changing Go's rules to eliminate captures leaves a huge state space but makes the winner trivial to determine.
Chapters
- Dave Churchill frames the lecture as a foundation of definitions for the rest of COMP 3200.
- The course uses a classical AI definition of agents, distinct from the newer idea of agentic LLMs.
- The terminology will recur in later lectures, assignments, and exam preparation.
- An agent is the entity within a problem that makes decisions and takes actions; the problem context determines which entity it is.
- For a robot vacuum, the house is the environment, keeping it clean is a possible goal, and the vacuum is the agent.
- Boston Dynamics' Spot, the cart-pole balancing task, and a chess player illustrate agents with very different capabilities.
- In video games, the agent might control Mario or an NPC dragon; it is not necessarily the human player.
- An agent observes the environment, selects an action, changes the environment's state, and repeats the cycle.
- Computational agents usually operate in discrete time steps, such as once per second or once per millisecond.
- Physical agents use sensors such as cameras, sonar, or laser rangefinding, and actuators such as wheels, arms, or limbs.
- In software environments such as Super Mario, an agent may read game memory directly rather than depend on noisy physical sensors.
- A percept is an instantaneous input snapshot; the environment state describes its configuration at a particular moment.
- A percept sequence is the full history of observations available to an agent, and some decisions depend on that history.
- A snapshot of a person falling does not reveal their velocity or trajectory, so a rescue agent may need observations from earlier moments.
- Chess is generally treated as fully observable and state-based, while a League of Legends Baron steal can require tracking damage, levels, items, and timing history.
- A state transition function maps a current state and an action to a next state: the notation is commonly S, A, and S′ or Sₜ and Sₜ₊₁.
- A transition function may be represented as a table, graph, neural network, or game engine.
- In Connect Four, dropping a yellow piece into a column changes the board state; the physical motion of the piece can be abstracted away.
- In Super Mario World, a button press affects the next frame, with the game running at 60 frames per second.
- A policy, written π, maps a state to an action; an optimal policy is denoted π*.
- A policy can be implemented as a lookup table, a flowchart of if-statements, a mathematical function, or a neural network.
- Blackjack policies specify choices based on the player's cards and the dealer's visible card.
- In pathfinding, arrows at each grid location can show which move the policy recommends toward a goal.
- A rational agent selects the action expected to succeed best according to a defined performance measure.
- Rationality depends on the performance measure, prior knowledge of the environment, available actions, and percept history.
- A rational decision is judged relative to the agent's knowledge: choosing an apparently clear route can be reasonable even if an unseen hazard exists.
- Performance measures may reward higher values or require minimizing a cost, such as time, distance, or money.
- For Flappy Bird, horizontal distance provides a simple score: reaching x = 100 performs better than reaching x = 50.
- In cart-pole balancing, the performance measure can be the number of seconds the pole stays upright.
- Dave Churchill describes UPS routing as an example where reducing left turns—not simply minimizing distance—reportedly saved about $50 million per year.
- Chess piece values such as pawn = 1, bishop = 3, rook = 5, and queen = 9 are useful intermediate guidance, but can mislead when a tactical move leads to checkmate.
- Expected value is the sum of each possible numerical outcome multiplied by its probability.
- A fair six-sided die has expected value 3.5, even though no single roll produces 3.5.
- A bet on rolling 5 or 6 has a one-in-three chance; repeated decisions should be evaluated by their expected payoff rather than one result.
- A rational agent chooses using expected outcomes, whereas an omniscient agent would know the exact result of randomness in advance.
- A task environment defines the problem; in chess it is the board, rules, pieces, legal moves, and game conditions—not the room containing the board.
- A chess environment includes the initial arrangement of 32 pieces and movement rules such as diagonal bishop moves and forward pawn moves.
- Some tasks have a specific goal state, such as reaching the Avalon Mall on a map; chess instead has a goal condition, such as checkmating the opponent.
- The environment's specifications determine which agent design and algorithms are appropriate.
- The state space is the set of possible environment configurations; the action space is the set of actions available from a state.
- Tic-tac-toe has at most 3⁹, or roughly 10⁴, board configurations, making exhaustive enumeration feasible.
- Chess has an estimated state space around 10⁴³, while Go's 19×19 board has 3³⁶¹, approximately 10¹⁷², possible configurations.
- Action-space size can vary by state: chess has 20 legal opening moves, but a midgame position may offer 50–70.
- Game-tree complexity counts action sequences, so multiple histories can reach the same board configuration and make the tree larger than the state space.
- The scale of exponential growth is difficult to intuit: the number of seconds in a century is about 10⁹, while the number of atoms in the universe is about 10⁸⁰.
- Large state spaces do not alone prove that a game is strategically complex; Go without captures has a similarly large space but the first player can always retain a one-piece lead.
- Complexity depends on both possible configurations and the game's rules, not just how many board arrangements exist.
- Fully observable environments expose all relevant state information, as in chess; partially observable environments hide information, as in card games or unexplored StarCraft territory.
- Deterministic transitions produce the same next state from the same state and action; stochastic environments include uncertainty such as dice rolls or card draws.
- Episodic tasks, such as classifying independent handwritten digits, do not depend on prior instances; sequential tasks such as chess have long-term action consequences.
- Dynamic environments can change while an agent deliberates, as in StarCraft; static turn-based chess waits for the player to act.
- Discrete problems use distinct states or actions, while real-world robotics often involves continuous time, positions, forces, and control values.
- Single-agent pathfinding differs from multi-agent settings such as cooperative robots or competitive games, where other agents' actions matter.
- Incomplete information means the agent does not know all rules or physics; this is distinct from partial observability, which concerns hidden state.
- A chessboard can be fully visible yet involve incomplete information if the agent has not learned the rules.
- The lecture's difficulty progression moves from a crossword-like single-agent task through chess, backgammon, poker, StarCraft, and continuous robotics.
- Potential exam tasks include identifying environment properties from a game, naming a game from its properties, and estimating a state or action space.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Dave Churchill.