Dr Libo Li, UNSW Sydney
Dr Ruyi Liu, UNSW Sydney
This course provides a unified mathematical framework for decision-making under uncertainty, bridging the gap between classical Optimal Control and modern Reinforcement Learning (RL). Over four weeks, students progress from the foundations of Dynamic Programming and the Linear Quadratic Regulator (LQR) to Stochastic Differential Equations (SDEs) and the Stochastic Hamilton-Jacobi-Bellman (HJB) equation. The course will then go into Markov Decision Processes (MDPs) to establish the logic of discrete-state stochasticity, culminating in model-free Reinforcement Learning (Q-learning and Policy Function Approximation).
Week 1: Deterministic (Discrete and Continuous Time) Optimal Control
Week 2: Stochastic Control & The HJB
Week 3: Markov Decision Processes (MDPs)
Week 4: Reinforcement Learning (RL)
Take this pre-enrolment QUIZ to self evaluate and get a measure of the key foundational knowledge required.

Dr Libo Li is a Senior Lecturer in Statistics in the School of Mathematics and Statistics at the University of New South Wales (UNSW). His research lies at the intersection of probability theory, stochastic analysis, and mathematical finance. His work focuses on theory of stochastic processes, stochastic control, stochastic differential equations, backward stochastic differential equations (BSDEs), optimal stopping, and numerical methods for stochastic systems. More recently, his research has expanded to reinforcement learning and machine learning approaches to stochastic optimisation and mathematical finance.

Dr Ruyi Liu is a Lecturer in the School of Mathematics and Statistics at UNSW Sydney. He works on stochastic control, optimal stopping, and backward stochastic differential equations (BSDEs/FBSDEs), with a particular interest in how these tools apply to real-world markets — from derivatives pricing to trading strategies and electricity systems. A common thread in his work is turning these problems into explicit, implementable solutions, such as threshold trading rules, closed-form optimality, and related BSDEs for pricing and hedging.