Back close

Course Detail

Course Name Reinforcement Learning
Course Code 26CSC403
Program 5 Year Integrated M.Sc in Data Science
Semester 7
Credits 4
Campus Coimbatore

Syllabus

Unit 1

Introduction: Reinforcement Learning, Elements of Reinforcement Learning, Limitations and Scope, An Extended Example- Tic-Tac-Toe.

Unit 2

Multi-armed Bandits: A k-armed Bandit Problem, Action-value Methods, the 10-armed Testbed, Incremental Implementation, tracking a Nonstationary Problem, Optimistic Initial Values, Upper-Confidence-Bound Action Selection, Gradient Bandit Algorithms

Unit 3

Finite Markov Decision Processes: The AgentEnvironment Interface, Goals and Rewards, Returns and Episodes, Unified Notation for Episodic and Continuing Tasks, Policies and Value Functions, Optimal Policies and Optimal Value Functions, Optimality and Approximation. Review of Markov process and Dynamic Programming.

Unit 4

Temporal-Difference Learning: TD Prediction, Advantages of TD Prediction Methods, Optimality of TD, Sarsa: On-policy TD Control, Q-learning: Policy TD Control. Expected Sarsa. Maximization Bias and Double Learning.Eligibility Traces, Functional Approximation, Fitted Q, DQN & Policy Gradient for Full RL and Hierarchical RL.

Text Books / References

Text Books

  1. Richard S. Sutton and Andrew G. Barto, Reinforcement Learning:An Introduction, second edition, MIT Press, 2019.

References

  1. Phil Winder, Reinforcement Learning, O’Reilly Media Publisher, 2020.
  2. Sudharsan Ravichandiran, Hand-on Reinforcement Learning with Python, Packt Publications, 2018.
  3. Sayon Dutta, Reinforcement Learning with Tensor Flow: A beginner’s guide, Packt Publications, 2018.

Introduction

Reinforcement learning (RL) is a paradigm that aims to model the trial-and-error learning process that is needed in many problem situations where explicit instructive signals are not available. It has roots in operations research, behavioral psychology and AI. The goal of the course is to introduce the basic mathematical foundations of reinforcement learning, as well as highlight some of the recent directions of research.

Objectives and Outcomes

Course Outcomes: After successful completion of the course, students will be able to

  • CO1:Understand the basic building blocks of reinforcement learning and its merits &
    limitations.
  • CO2: Apply the RL algorithms to solve Multi-Armed Bandit problem
  • CO3:Formulate your task as a Reinforcement Learning problem, and how to begin implementing a solution.
  • CO4: Relate Markov decision process with Dynamic Programming
  • CO5:Understand the space of RL algorithms (Temporal- Difference learning, Q-learning and Sarsa)

CO-PO Mapping:

  PO1 PO2 PO3 PO4 PO5 PO6 PO7 PO8 PO9 PO10 PO11 PO12
CO1 3 3 2   1       1 1   1
CO2 2 2 2   2       2 2   2
CO3 2 2 2   2       2 2   2
CO4 2 2 2   2       2 2   2
CO5 3 2 2   2       2 2   2

DISCLAIMER: The appearance of external links on this web site does not constitute endorsement by the School of Biotechnology/Amrita Vishwa Vidyapeetham or the information, products or services contained therein. For other than authorized activities, the Amrita Vishwa Vidyapeetham does not exercise any editorial control over the information you may find at these locations. These links are provided consistent with the stated purpose of this web site.

Admissions Apply Now