Agents
| Theme | Objectives (after the course, you ...) |
| Agents |
|
Episode 12: Agents, reinforcement learning, and robots
An agent is a system that perceives an environment and acts in it. This definition covers a thermostat, a chess program, a delivery robot, and an LLM that chooses tools. The important shift is from predicting an answer to selecting an action whose consequences change what happens next. An agent therefore operates in a loop: observe, update an internal state, choose an action, and observe the result.
Rational action under limited information
A rational agent chooses the action expected to perform best according to its performance measure, given what it has perceived and what it knows. Rational does not mean omniscient or morally good. A route planner can be rational while choosing a road that later becomes blocked; it acted on the information available. Nor does rationality imply unlimited computation. Real agents have deadlines, finite memory, incomplete models, and limited energy. Bounded rationality asks what good action is possible within those limits.
Specifying the performance measure is therefore a design decision, not bookkeeping. A delivery system rewarded only for speed may drive dangerously; a chatbot rewarded for engagement may prolong an unhelpful conversation. Proxy measures are useful because goals such as “help the user” are difficult to calculate, but optimizing a proxy can produce unintended behaviour. This is the beginning of the alignment problem: how can the behaviour produced by an objective remain consistent with human intentions and constraints?
Specifying a task with PEAS
Before choosing an algorithm, an agent designer should specify the task. A useful checklist is PEAS: Performance measure, Environment, Actuators, and Sensors. Consider an autonomous campus delivery robot. Its performance measure might combine successful deliveries, travel time, energy consumption, comfort around pedestrians, and avoidance of collisions. Its environment includes paths, doors, lifts, people, weather, and other vehicles. Its actuators include wheels, brakes, a compartment lock, lights, and perhaps speech. Its sensors may include cameras, lidar, wheel encoders, GPS, inertial sensing, and signals from the delivery compartment.
PEAS exposes assumptions that are otherwise easy to hide. “Deliver quickly” is not yet a sufficient objective: how are failed deliveries counted, how much delay is acceptable near pedestrians, and who may open the compartment? Similarly, a tool-using research assistant needs more than a language model. Its environment includes documents, databases, websites, user instructions, and institutional policies; its actuators may include search, file creation, code execution, and messaging; and its sensors include prompts, tool outputs, permissions, and error messages.
| Agent | Performance | Environment | Actuators | Sensors |
|---|---|---|---|---|
| Campus robot | Safe, correct, timely delivery | Paths, buildings, people, weather | Drive, stop, signal, unlock | Camera, lidar, GPS, encoders |
| Tool-using assistant | Correct and useful completion within limits | User, files, services, policies | Respond, retrieve, calculate, call API | Prompt, memory, results, errors |
| Game agent | Expected score or probability of winning | Board, rules, opponent | Legal moves | Current state or observation history |
What kind of environment?
Task environments differ along dimensions that strongly affect architecture. An environment is fully observable when the agent can access all state relevant to its decision, and partially observable when observations are incomplete or noisy. It is deterministic when an action and state fix the next state, and stochastic when several outcomes remain possible. In an episodic task, one decision does not affect later cases; in a sequential task, current action changes future choices. Environments may also be static or dynamic, discrete or continuous, known or initially unknown, and single-agent or multi-agent.
Chess is fully observable and discrete, but adversarial and sequential. Driving is partially observable, stochastic, continuous, dynamic, and multi-agent. Document classification is often treated as episodic and static. An LLM agent working on a live software repository is sequential and partially observable: a file may change between observations, tools may fail, and another developer may edit the same code. Recognizing these properties prevents us from applying a solution designed for a clean puzzle to a changing real environment.
A language model maps context to likely tokens. An agentic system places such a model inside a loop with memory, tools, an action policy, observations, and stopping conditions. The model may propose an action; the surrounding system validates and executes it. Calling an API is not magically intelligent: it expands what the system can affect and therefore increases the need for permissions, monitoring, and recovery.
Agent architectures
A simple reflex agent uses condition–action rules: if the temperature is low, turn on heating. It is fast but fails when the current observation does not reveal everything relevant. A model-based agent maintains an internal state—such as an estimate of where a robot is—and updates it as observations arrive. A goal-based agent evaluates actions by whether they reach a desired state, often using search or planning. A utility-based agent compares trade-offs among outcomes, for example speed, safety, energy, and comfort. A learning agent changes its behaviour from data or experience. Practical systems combine these ideas rather than fitting one box.
| Architecture | Question it answers | Typical limitation |
|---|---|---|
| Reflex | What rule matches now? | No memory or foresight |
| Model-based | What state am I probably in? | The model may be wrong |
| Goal/utility-based | Which future outcome is preferable? | Planning can be expensive |
| Learning | How should behaviour change with experience? | Needs suitable feedback and safe exploration |
Modern LLM agents often implement a repeated plan–act–observe cycle. The model receives a goal and current context, proposes a tool call, receives the tool result, and continues until it returns an answer or reaches a stopping rule. Useful components include a tool registry, working memory, longer-term retrieval, a planner, and an evaluator. Reliability comes from engineering the whole loop: typed tool interfaces, least-privilege permissions, time and cost limits, validation of outputs, logging, and human approval for consequential actions.
Memory, planning, and tools in LLM agents
An LLM has a finite context window, not a durable autobiographical memory. Agent systems therefore distinguish several forms of memory. Working memory holds the current goal, intermediate results, and recent observations. Episodic memory stores previous trajectories or task outcomes. Semantic memory retrieves facts and documents, often through embeddings. Procedural memory is represented by instructions, tools, code, or learned policies. More memory is not automatically better: irrelevant retrieval can distract the model, stale information can conflict with the present environment, and stored material may contain sensitive data.
Planning can also take several forms. A model may produce a complete plan before acting, alternate one reasoning step with one tool call, generate several candidate plans and compare them, or use a separate evaluator to decide whether progress is adequate. Long plans can become obsolete after the first unexpected observation, while purely reactive action may lose the global goal. Effective systems often combine a coarse plan with frequent replanning.
A tool converts a generated symbol into an external action. Its interface should therefore specify required arguments, valid types, possible errors, side effects, and permissions. Reading a calendar and deleting an event must not receive the same authority. A robust execution layer validates arguments, asks for confirmation when consequences are difficult to reverse, records what happened, and returns structured observations. The model may be fluent, but the tool layer determines what the system can actually change.
Why agentic loops fail
Agent failures are often failures of the loop rather than failures of language generation alone. The model may select the wrong tool, use a correct tool with incorrect arguments, misread a result, repeat an action indefinitely, stop before verifying success, or continue after the goal has been achieved. Retrieved documents or web pages can also contain instructions that conflict with the user’s goal—a form of prompt injection. Because the observation is text, the model may not distinguish trusted system instructions from untrusted content unless the architecture makes that distinction explicit.
| Failure | Example | Engineering response |
|---|---|---|
| Wrong action | Searches when calculation is required | Tool descriptions, routing tests, evaluator |
| Invalid action | Malformed API arguments | Typed schemas and validation |
| Unsafe authority | Sends or deletes without confirmation | Least privilege and approval gates |
| Looping | Repeats the same unsuccessful call | Step, time, and cost budgets |
| False completion | Claims success without checking result | Explicit success tests and verification |
| Untrusted instruction | Web content redirects the agent’s goal | Provenance, isolation, and policy checks |
Reinforcement learning: learning from consequences
In supervised learning, a training example normally supplies a target answer. In reinforcement learning (RL), an agent instead interacts with an environment. At time t it observes state s, chooses action a, receives reward r, and moves to a new state. A policy specifies how actions are selected. The objective is not necessarily the next reward but the expected return: the accumulated future reward, often discounted so that distant rewards count less.
The central difficulty is delayed consequence. A chess move receives no immediate label saying whether it was good; its value may become visible many moves later. RL must assign credit across a trajectory. In a Markov decision process, the state is assumed to contain the information needed to predict the next-state and reward distributions. This abstraction is powerful, although real observations are often partial, making memory or belief-state estimation necessary.
The Markov decision process
A finite Markov decision process can be described by states, actions, transition probabilities, rewards, and a discount factor. The transition model answers how likely each next state is after an action. The reward model assigns immediate numerical feedback. The discount factor determines how strongly future rewards matter. A policy maps a state—or an observation history—to a distribution over actions. Learning seeks a policy with high expected return, while planning computes a policy when a sufficiently accurate environment model is already known.
The Markov assumption does not mean that history never matters. It means that the chosen state representation summarizes the relevant history. The physical location of a delivery robot is not enough if battery level, payload, traffic direction, and current reservation also affect what happens next. If important information is hidden, the agent may maintain a belief state: a probability distribution over possible underlying states. This connects RL with the probabilistic reasoning introduced earlier in the course.
Value functions estimate long-term return. The action-value function Q(s,a) asks: if I take action a in state s and then continue according to a policy, how much return should I expect? Q-learning updates an estimate toward the observed reward plus the best estimated value of the next state. In words:
new estimate ← old estimate + learning rate × (better target − old estimate)
The “better target” contains two pieces: the reward just observed and the estimated value of the best action available next. This is the Bellman idea: the value of a decision can be decomposed into immediate consequence plus the value of what becomes possible afterward. Repeated updates propagate information about distant rewards backward through previously visited states. Q-learning is model-free because it does not first estimate the complete transition model.
By contrast, model-based RL learns or uses a model of how the environment changes, then plans through that model. It can be more data-efficient because imagined experience supplements real interaction, but errors in the model may be exploited by the planner. Policy-gradient methods optimize a parameterized policy directly, which is useful for continuous actions. Actor–critic methods combine an actor that selects actions with a critic that estimates value. These families differ technically, but they address the same core problem: improving sequential decisions from evaluative feedback.
The discount factor controls the weight of future rewards. The learning rate controls how quickly new evidence replaces old estimates. With neural networks, deep RL approximates values or policies for large state spaces, but training can be unstable and data-hungry.
A worked Q-learning update
Suppose the agent is one step from the goal. Moving right gives reward 10 and ends the episode. The present estimate is Q(state, right) = 2, and the learning rate is 0.5. Because the episode ends, there is no future value to add. The target is therefore 10, the prediction error is 10 − 2 = 8, and the updated estimate is 2 + 0.5 × 8 = 6. The estimate moves halfway toward the newly observed outcome.
Now consider an earlier state. Moving right gives immediate reward 0 and reaches the state whose best estimated action has value 6. With discount factor 0.9, the target is 0 + 0.9 × 6 = 5.4. If the old estimate was 1 and the learning rate remains 0.5, the new estimate becomes 1 + 0.5 × (5.4 − 1) = 3.2. In this way, information about the goal propagates backward one transition at a time. Repeated experience gradually distinguishes actions that merely look promising from actions that reliably lead to long-term reward.
| Quantity | Meaning | Worked value |
|---|---|---|
| Q(s,a) | Current expected return for the chosen action | 1.0 |
| r | Reward observed after acting | 0 |
| γ max Q(s′,a′) | Discounted value available next | 0.9 × 6 = 5.4 |
| Target | Immediate reward plus estimated future value | 5.4 |
| TD error | Target minus current estimate | 4.4 |
| Updated Q | Old value plus learning rate times error | 1 + 0.5 × 4.4 = 3.2 |
How the main RL families differ
| Approach | What is learned? | Strength | Typical difficulty |
|---|---|---|---|
| Value-based | Value of states or state–action pairs | Clear action comparison; effective for discrete actions | Awkward for large continuous action spaces |
| Policy gradient | Policy parameters directly | Handles stochastic and continuous policies | Gradient estimates can have high variance |
| Actor–critic | Policy plus a value estimator | Critic can make policy learning more efficient | Two coupled learners can become unstable |
| Model-based | Environment dynamics and/or reward model | Can plan and reuse experience efficiently | Planning may exploit model errors |
The distinction between on-policy and off-policy learning is also important. An on-policy method evaluates and improves the policy that currently generates behaviour. An off-policy method can learn about one target policy from experience produced by another behaviour policy. Q-learning is off-policy: exploratory behaviour may take a random action, while the update uses the estimated value of the best next action. This permits experience replay and learning from logged data, but the difference between the data-generating policy and the learned policy can also create instability.
Why deep reinforcement learning is difficult
Replacing a Q-table with a neural network allows generalization across high-dimensional observations, but it removes the simplicity of updating one independent table entry. A change intended to improve one state can alter predictions for many other states. Consecutive observations are correlated, the target itself changes as the network learns, and rare successful trajectories may be overwhelmed by ordinary failures. Common techniques include experience replay, slowly updated target networks, normalized observations, reward scaling, and parallel environments.
Sample efficiency matters because celebrated game-playing results may require millions or billions of simulated interactions. A physical robot cannot safely repeat that amount of trial and error. Offline RL learns from an existing dataset; imitation learning begins from demonstrations; model-based learning generates imagined transitions; and curriculum learning introduces harder situations gradually. Each reduces some cost but introduces assumptions about coverage, model accuracy, or demonstrator quality.
Evaluation should separate training return from genuine capability. A policy may memorize a fixed level, exploit an emulator bug, depend on an exact camera angle, or collapse after a small change in dynamics. Strong evaluation uses unseen initial states, altered environments, several random seeds, explicit safety measures, and comparisons with simple baselines. The question is not only whether reward increased, but what behaviour produced that reward and whether it transfers.
Exploration, exploitation, and reward design
An agent faces an exploration–exploitation dilemma: use the action currently believed best, or try alternatives that may teach it something. An epsilon-greedy policy usually chooses the best-known action but sometimes explores randomly. Too little exploration locks in early mistakes; too much prevents consistent performance. In real systems, random exploration may be unsafe, so learning may begin in simulation, from logged data, or under explicit safety constraints.
Reward is not the same as the true goal. If a cleaning robot earns points for detecting dirt, it might avoid finishing the job so that dirt remains detectable. This illustrates reward hacking: the agent finds a high-scoring behaviour that violates the designer’s intent. Good design combines reward shaping with constraints, evaluation on varied situations, monitoring for distribution shift, and the ability to interrupt or correct behaviour.
From human feedback to language-model behaviour
Reinforcement learning also appears in the development of language models. After pretraining, people may compare alternative responses. A reward model learns to predict these preferences, and the language model is then optimized to produce responses receiving higher predicted reward. This family of methods is commonly called reinforcement learning from human feedback (RLHF). Related approaches optimize directly from preference pairs or use AI-generated feedback.
The connection to an embodied RL agent is real but should not be overstated. During ordinary conversation, a deployed language model usually does not update its parameters after every user response. Much of the reinforcement learning occurred during training. A tool-using agent, however, does execute a sequential policy at deployment time: it observes results, changes its next action, and may store information in external memory. Training-time alignment and deployment-time agency are therefore different layers of the system.
Human preferences are also an imperfect proxy. Annotators may disagree, reward models can be confidently wrong outside their training distribution, and optimizing an average preference may suppress legitimate minority needs. These issues motivate diverse evaluation, uncertainty estimates, red-team testing, and limits on actions—not merely a larger reward model. Part 7 examines the broader ethical and governance questions; here the engineering lesson is that feedback must be interpreted in relation to the task and the people affected.
Multi-agent systems
When several agents share an environment, each agent’s outcome depends on the others. Interactions may be cooperative, competitive, or mixed. Traffic participants coordinate implicitly; auction bidders compete; a robot team may divide tasks. The environment becomes non-stationary from one learner’s perspective because other agents are also changing their policies.
Game theory provides concepts such as best response and equilibrium, while multi-agent RL studies how policies can be learned through interaction. Communication can help agents coordinate, but messages can be incomplete, strategic, or deceptive. A system of individually capable agents is not automatically collectively effective: congestion, free riding, duplicated work, and cascades of mistaken assumptions can emerge. Evaluation should therefore examine system-level outcomes—fairness, robustness, resource use, and failure propagation—not only each agent’s score.
Coordination, competition, and emergent behaviour
In a cooperative task, agents share a team reward, but assigning credit remains difficult: which agent’s action caused success? In a competitive task, improvement by one agent changes the challenge faced by others. Mixed settings contain both, as in traffic: drivers share an interest in avoiding collisions but compete for space and time. Centralized training may use information about all agents, while decentralized execution gives each agent only local observations.
Communication is an action with costs and consequences. Agents must decide what to communicate, to whom, and when. A message may reduce uncertainty, coordinate complementary roles, or create common knowledge; it may also consume bandwidth or reveal strategy. In teams of LLM agents, assigning “researcher,” “critic,” and “writer” roles can improve coverage, but it can also produce duplicated work or amplify the same misconception when all agents rely on similar models and sources.
Emergence means that system-level patterns arise from local interactions without being explicitly programmed as a global plan. Flocking can emerge from simple separation, alignment, and cohesion rules. Congestion can emerge even when every driver tries to minimize travel time. Emergence is not automatically intelligent or beneficial. Multi-agent evaluation therefore needs counterfactual tests, varied group composition, communication failures, and measurement of collective rather than merely individual outcomes.
Robots: agents with bodies
A robot closes the agent loop in the physical world. Sensors translate light, force, sound, distance, or joint position into measurements. Actuators produce motion. Perception estimates relevant state; planning selects a route or sequence of actions; control turns that plan into motor commands. All stages are uncertain. A wheel slips, a camera is occluded, an object moves, or the environment differs from training.
This embodiment creates several gaps. The reality gap separates simulation from physical deployment. The semantic gap separates sensor values from meaningful objects and situations. The control gap separates a planned motion from what imperfect hardware actually does. Robust robots repeatedly estimate, act, measure the consequence, and correct. Autonomy is therefore usually layered: low-level controllers stabilize motion, planners choose actions, and people provide goals, supervise exceptions, or share control.
Perception and state estimation
Raw sensors do not deliver a ready-made symbolic world. A camera produces arrays of intensity values; lidar produces distances; an inertial measurement unit produces accelerations and rotation rates. Perception detects objects, surfaces, people, and free space. State estimation combines uncertain measurements over time. For a mobile robot, localization estimates where the robot is; mapping estimates the environment; simultaneous localization and mapping (SLAM) tackles both together.
Combining sensors can make estimates more robust. GPS may drift or disappear indoors, cameras may fail in darkness, and wheel odometry accumulates error when wheels slip. Sensor fusion uses the different error characteristics of several measurements. The result is still an estimate with uncertainty, not ground truth. A planner should behave differently when localization is precise than when several locations remain plausible.
Planning and control
Planning operates at multiple levels. A task planner may decide to collect a parcel before going to the recipient. A motion planner finds a collision-free trajectory through configuration space. A local planner reacts to nearby obstacles. A controller continuously converts a desired trajectory into motor commands and corrects deviations. These layers run at different timescales: strategic plans may update every few seconds, while balance control may update hundreds of times per second.
Classical search remains important here. A* can plan paths on a map; sampling-based methods such as rapidly exploring random trees handle high-dimensional robot configurations; model-predictive control repeatedly optimizes a short future horizon and replans after new measurements. Learned components can improve perception, predict dynamics, or propose actions, but geometric constraints and feedback control continue to provide valuable structure.
Simulation, transfer, and shared autonomy
Physical interaction is slow, costly, and sometimes dangerous, so robots often learn or are tested in simulation. Unfortunately, a policy may exploit details of the simulator that do not hold in reality. Domain randomization varies textures, lighting, friction, mass, delays, and sensor noise during training so that no single simulated world is sufficient. System identification instead tries to make the simulator match the physical system. Real-world fine-tuning and conservative validation remain necessary.
Full autonomy is not always the most useful objective. In shared autonomy, a person supplies intent while the robot assists with execution. A wheelchair may infer the intended doorway while preserving the user’s authority; a surgical training system may guide movement without replacing the learner’s action. Assistance that is too weak provides little benefit, but assistance that is too strong can hide errors and prevent skill development. The design question is therefore not simply “human or robot?” but how initiative, information, and control should be distributed over time.
Robots also make stakes tangible. An LLM’s incorrect sentence can be checked before use; a robot’s incorrect movement can cause immediate harm. Safe deployment relies on redundant sensing, conservative operating envelopes, collision avoidance, emergency stops, testing, and clear allocation of responsibility. Human oversight must be actionable: a nominal supervisor who cannot understand or interrupt the system is not meaningful oversight.
Putting the pieces together
| Example | State/observation | Actions | Feedback |
|---|---|---|---|
| Tool-using LLM | Prompt, memory, tool results | Respond, search, call API | Task success, evaluator or user |
| Game-playing agent | Board or observation history | Legal moves | Score or win/loss |
| Delivery robot | Map and noisy sensors | Move, stop, manipulate | Progress, safety and energy |
The same vocabulary—environment, observation, action, policy, objective, and feedback—connects all three. Yet their engineering requirements differ because tools, other agents, and physical bodies introduce different consequences. When analysing any agent, ask: What can it observe? What can it change? What objective is actually optimized? How does it learn or plan? Where can uncertainty enter? Who can intervene? These questions are more informative than simply asking whether the system is “autonomous.”
Agency is a relationship between a system and an environment. Reinforcement learning provides one way to improve action through consequences; agentic frameworks connect models to tools and memory; robotics grounds the loop in uncertain physical reality. Capability depends on the entire loop, while safe and intended behaviour depends on objectives, constraints, monitoring, and human authority.
Before we can build an agent, we need a precise way to describe the job we're asking it to do. Russell and Norvig's classic recipe is PEAS: Performance measure (what counts as doing well), Environment (the world the agent operates in), Actuators (how it acts on the world), and Sensors (how it perceives the world).
As a worked example, here's the PEAS description for a robot vacuum cleaner:
| PEAS | Description |
|---|---|
| Performance measure | Amount of dirt cleaned, energy used, noise generated, time taken |
| Environment | A room (or house) with floor, furniture, dirt, stairs |
| Actuators | Wheels (steering, driving), brushes, vacuum suction |
| Sensors | Dirt sensor, bump sensor, cliff sensor, camera |
Once we have PEAS, we can also classify the environment itself along six standard dimensions:
- Fully vs. partially observable: can the agent's sensors see the complete state of the environment at every step?
- Deterministic vs. stochastic: does the next state depend completely on the current state and the agent's action, or is there uncertainty/randomness?
- Episodic vs. sequential: is each percept-action pair independent of the others, or do earlier actions affect later ones?
- Static vs. dynamic: can the environment change while the agent is deliberating?
- Discrete vs. continuous: are there a finite number of distinct states/actions, or a continuum?
- Single-agent vs. multi-agent: is our agent alone in the environment, or are there other agents (cooperative or competitive)?
For the vacuum cleaner above: partially observable (can't see dirt in the next room), stochastic (bumps and slips happen), sequential (where it already cleaned matters), dynamic (people walk through and make new mess), discrete (a finite grid of rooms and a finite action set), and single-agent (usually).
Now it's your turn. For each of the following three tasks, (a) write out a PEAS description in a table like the one above, and (b) classify the environment along all six dimensions, with one sentence justifying each choice:
- An autonomous drone delivering parcels across a city.
- A poker-playing agent at an online table with human opponents.
- A customer-service chatbot answering questions on a retailer's website.
Bonus (no extra points, just for fun): pick a job or hobby of your own and write its PEAS description. Is it harder or easier to pin down than you expected?
One reasonable PEAS description for the delivery drone:
| PEAS | Description |
|---|---|
| Performance measure | Parcels delivered on time, battery/fuel used, airspace violations avoided, no collisions |
| Environment | City airspace: buildings, other drones, weather, no-fly zones, delivery addresses |
| Actuators | Rotors (altitude, direction, speed), parcel-release mechanism |
| Sensors | GPS, camera, altimeter, obstacle-proximity sensors, wind sensor |
Environment classification: partially observable (can't see around buildings or predict other drones' exact paths), stochastic (wind, GPS drift, other traffic), sequential (route choices affect battery left for later deliveries), dynamic (weather and traffic change mid-flight), continuous (position, altitude, and speed are continuous quantities, unlike the vacuum world's discrete grid), and multi-agent (shares airspace with other drones and aircraft).
For the poker agent: performance measure is (expected) winnings; environment is the table, cards, and other players' visible actions/chip counts; actuators are bet/call/fold/raise actions; sensors are the visible cards and the game log. Classification: partially observable (opponents' hole cards are hidden — this is the whole game), stochastic (card deals), sequential (betting history within a hand matters, and bankroll carries across hands), dynamic only in the loose sense that other players act while you're "thinking" in real-time formats, discrete (finite cards, finite bet sizes in most formats), and multi-agent, specifically adversarial.
For the customer-service chatbot: performance measure is something like query resolution rate and customer satisfaction; environment is the website's chat interface and knowledge base/order database; actuators are the text (and maybe UI actions like "issue refund") it can send back; sensors are the customer's typed messages and account/order data. Classification: partially observable (can't see the customer's actual intent, only their words), stochastic (customers phrase things unpredictably), mostly episodic at the level of a single ticket but sequential within a conversation, dynamic in the sense that order status can change mid-chat, discrete (a finite vocabulary and action set, in practice), and single-agent from the chatbot's own point of view.
Descriptions will vary — what matters is that the performance measure is actually measurable, the environment/actuators/sensors are mutually consistent, and each of the six classifications is justified rather than just asserted.
A learning agent improves its behavior from experience rather than following a fixed rule. We'll train a real one — no theory derivation required, the algorithm is already correctly implemented for you. Follow Gymnasium's official "Training an Agent" tutorial (Blackjack environment).
The core training loop, already provided:
for episode in tqdm(range(n_episodes)):
obs, info = env.reset()
done = False
while not done:
action = agent.get_action(obs)
next_obs, reward, terminated, truncated, info = env.step(action)
agent.update(obs, action, reward, terminated, next_obs)
done = terminated or truncated
obs = next_obs
agent.decay_epsilon()
and the agent's hyperparameters are exposed right at the top of the script:
learning_rate, n_episodes, start_epsilon,
epsilon_decay, final_epsilon.
- Run the tutorial's full script as-is (it includes built-in plotting of the learning curve — episode reward and episode length over time).
- Change
learning_rateto a much smaller value (e.g.0.001) and a much larger one (e.g.0.5). Rerun with each and compare the learning curves. - Change
n_episodesdown to a small number (e.g.1000, from the tutorial's default which is much larger). Does the agent still learn a reasonable policy? - In 4–5 sentences: what did you observe about the learning-rate extremes — does a bigger learning rate always mean faster learning? Connect this to anything similar you've seen in this course (e.g. the trade-off you saw with learning rate in gradient descent for logistic regression).
agent.update() actually doing?
It's applying the Q-learning update rule: adjust the estimated value of the (state, action) pair you just took, using the reward you got plus the best value you think you can get from here on. You don't need to derive this rule yourself for this exercise — just observe how the agent's behavior changes as you change the parameters that control how aggressively/cautiously it learns and explores.
We've talked about agentic frameworks conceptually — now let's build one you can
actually run. smolagents is a lightweight, purpose-built-for-
teaching agent library from Hugging Face. Getting a first agent running is
genuinely four lines:
from smolagents import CodeAgent, InferenceClientModel
agent = CodeAgent(tools=[], model=InferenceClientModel(), add_base_tools=True)
agent.run("Could you give me the 118th number in the Fibonacci sequence?")
Setup note: this needs a free Hugging Face account and an access token (hf.co/settings/tokens) — no payment required, a free account includes enough inference credit for this exercise.
- Install with
pip install smolagents, get a free HF token, and run the code above. Read the printed trace — the agent will show its reasoning (its "Thoughts") and the code it decides to run. - Add your own tool — a Python function decorated with
@toolthat does something simple, e.g. converts a temperature between Celsius and Fahrenheit, or counts vowels in a string:from smolagents import tool @tool def celsius_to_fahrenheit(celsius: float) -> float: """Converts a temperature from Celsius to Fahrenheit. Args: celsius: the temperature in degrees Celsius """ return celsius * 9 / 5 + 32 agent = CodeAgent(tools=[celsius_to_fahrenheit], model=InferenceClientModel()) agent.run("What is 24 degrees Celsius in Fahrenheit?") - Ask the agent a question that requires it to use your tool, and one that doesn't. Confirm from the printed trace that it only calls your tool when it's actually relevant.
- In 3–4 sentences: describe, in your own words, the loop the agent went through — how did it decide when to use your tool vs. when to answer directly?