AStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 8: LQR-Style Algorithms
Watch on YouTube →
Overview
This lecture from Stanford Online's AA203 course delves into Linear Quadratic Regulator (LQR)-style algorithms for optimal control, extending beyond basic state regulation to trajectory tracking and optimization. It explains how LQR can be adapted for nonlinear systems through linearization, leading to iterative LQR (iLQR) and differential dynamic programming (DDP) for trajectory optimization, and highlights the complementary nature of LQR with PID controllers in control system hierarchies.
Key takeaways
- LQR, typically for state regulation, can be adapted for trajectory tracking by reformulating the problem in terms of deviation variables (delta_x, delta_u).
- Nonlinear system tracking with LQR is achieved by linearizing the dynamics around a nominal trajectory, creating an auxiliary LQR problem for the deviation variables.
- Iterative LQR (iLQR) uses successive linearizations and quadratic approximations of dynamics and cost to optimize open-loop trajectories.
- Differential Dynamic Programming (DDP) is a more advanced method that directly approximates the Bellman equation, using second-order derivatives of the dynamics.
- LQR-based tracking controllers combine a feedforward term (nominal control) with a closed-loop feedback term that corrects deviations from the reference trajectory.
- LQR is complementary to PID control, with LQR typically handling trajectory planning and optimization, while PID manages low-level actuator control.
Chapters
- Optimal control problem formulation focused on closed-loop policies.
- Introduction to dynamic programming as the fundamental algorithm for computing optimal policies.
- Mention of LQR as an example application of dynamic programming.
- LQR is a widely used theoretical algorithm, second only to PID.
- LQR allows explicit specification of objectives, unlike PID's indirect shaping.
- LQR is flexible for stabilization, trajectory tracking, and optimal control trajectory computation.
- PID lacks an explicit objective function, while LQR allows intentional objective specification.
- PID's response shaping is indirect (e.g., pole placement, phase/gain margins).
- LQR offers an explicit handle on optimizing operations and is flexible for tracking and trajectory computation.
- PID is typically used at the lowest level, manipulating actuator signals.
- LQR sits above PID, focusing on trajectories and optimization.
- Control systems often use a hierarchy: LQR-derived optimizers above PID.
- LQR assumes linear dynamics (x_k+1 = A x_k + B u_k) and no state/control constraints.
- Dynamics can be time-varying (A_k, B_k).
- Quadratic cost function includes terminal state penalty and running costs on state and control, potentially with cross-terms (x' Q x + u' R u + x' H u).
- Inclusion of cross-terms (x' H u) in the cost function is useful for linearization in tracking algorithms.
- Optimal control law remains linear feedback (u_k = -L_k x_k).
- Optimal cost remains quadratic in the state (x' P x).
- Cost function can include linear terms (state, control) and constant terms.
- Dynamics can be affine (x_k+1 = A x_k + B u_k + d_k), incorporating known forcing terms or disturbances.
- Optimal control policy becomes linear feedback plus a feedforward term (u_k = -L_k x_k + K).
- The designer's role is to set up the LQR problem by choosing matrices Q, R, H, and Q_N.
- LQR tuning is crucial for achieving desired system response and can be time-consuming.
- Matrices like Q, R, H, and Q_N are design parameters, not fixed values.
- Objective: drive the system state towards a desired reference trajectory (x_bar, u_bar).
- Problem reformulated in terms of deviation variables (delta_x = x - x_bar, delta_u = u - u_bar).
- Tracking becomes an LQR problem on deviation variables, aiming to drive delta_x to zero.
- Assumes linear time-varying dynamics (x_k+1 = A_k x_k + B_k u_k).
- Reference trajectory must be feasible (satisfy system dynamics).
- Deviation dynamics are linear: delta_x_k+1 = A_k delta_x_k + B_k delta_u_k.
- An auxiliary LQR problem is set up for deviation variables, with cost quadratic in delta_x and delta_u.
- Controller has a feedforward term (nominal control u_bar) and a feedback term (L_k * delta_x_k).
- This 'two-step design' balances computational efficiency of open-loop trajectories with robustness of closed-loop control.
- The structure is u_k = u_bar_k + delta_u_k, where delta_u_k = -L_k * delta_x_k.
- Nonlinear dynamics (x_k+1 = f(x_k, u_k)) are linearized around the nominal trajectory using Jacobians (A_k, B_k).
- The problem is then reformulated in deviation variables, leading to a linear system for delta_x and delta_u.
- An auxiliary LQR problem is solved for delta_u, and the full control is u_k = u_bar_k + delta_u_k.
- Idea: Use LQR to perturb a feasible nominal trajectory to find one with lower cost.
- Linearize dynamics around the nominal trajectory to get linear dynamics for deviation variables.
- Approximate the cost function quadratically around the nominal trajectory.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Stanford Online.