Fall 2026, University of Wisconsin-Madison
| Instructor | Tengyang Xie |
| Teaching Assistants | Sarina Xi, Yurun Yuan |
| Time | Tuesday & Thursday, 11:00am - 12:15pm |
| Location | Morgridge Hall B2590 |
| Office Hours | Tengyang Xie: Wednesday 1-2pm, Morgridge Hall 5530 Sarina Xi: Monday 9-11am, Morgridge Hall 5683 Yurun Yuan: Friday 9-11am, Morgridge Hall 5683 |
| Announcements | Canvas |
| Homework Submission | Gradescope |
| Q&A | Piazza (please use Piazza first; email only for personal matters) |
This is a senior/master's-level introduction to reinforcement learning (RL): making good decisions in sequence, when each decision changes what happens next. The course covers the framework (Markov decision processes) and the core algorithms, from value iteration to RLHF: how each is designed, why it works, and when it fails. Algorithms are derived on the board, a few central theorems are proved in full, and the homework has you implement the algorithms and run them on small problems. Topics include (tentative):
| Date | Topic | Materials | Reading |
|---|---|---|---|
| Sep 03 | Introduction | Slides, Note1. HW0 out (due Sep 11). | Sutton & Barto, Ch. 1 (1.7 optional), Sec. 2.1, 2.9, 3.1-3.3 |
There is no project and no final exam. Three kinds of graded work:
Quizzes. About 15 minutes at the start of class, closed book, announced one or two lectures in advance. The best 5 of 6 count, so there are no make-ups.
Midterm. In class, 75 minutes, covering the first half of the course. More details before the midterm.
Homework. HW0 is a readiness self-check on the prerequisites, out Sep 03 and due Sep 11: everyone enrolled must submit it, but it does not count toward the grade. The four homeworks are programming-focused, in Python (NumPy first, PyTorch later). Submit on Gradescope.
HW1-HW4 share 4 slip days, with at most 2 used on any one assignment. Slip days are applied automatically from the submission timestamp; no request and no reason are needed.
Every assignment closes 48 hours after its deadline, and solutions are released then. Late time not covered by slip days costs 10% of the maximum score per started 24-hour period, counted the same way. After the hard close the score is 0.
Slip days do not apply to HW0 (required but not graded), to quizzes (the best 5 of 6 count instead), or to the midterm. Documented emergencies, disability accommodations, and religious observance are handled by the instructor outside the slip-day budget; please reach out.
Students should be comfortable with:
Deep learning is not assumed at the start: the second half of the course uses PyTorch at a basic level, and warm-up material will be released before HW3. Prior exposure to reinforcement learning is not required. HW0 is a readiness check on this background. If parts of it feel unfamiliar, start with the review material linked above and reach out early.
Homework must be your own work. You may discuss ideas and debugging approaches with at most two named classmates, but you must write every proof, every line of code, and every plot yourself, and every submission lists its collaborators and sources. Quizzes and the midterm are individual work.
Cheating and plagiarism will be dealt with in accordance with University procedures (see the UW-Madison Academic Misconduct Rules and Procedures). If you have any questions about this policy, please ask the instructor before you act.
Each assignment carries a traffic-light label for the use of generative AI tools (ChatGPT, Claude, Gemini, Copilot, and others):
Every submission includes an Assistance Statement, and you must be able to explain everything you submit. Use beyond an assignment's label is a violation of the course's expectations and will be addressed through UW-Madison's academic misconduct policy.