Âé¶¹´«Ã½

Skip to main content

Module

DSC8053 : Reinforcement Learning (Inactive)

  • Inactive for Year: 2026/27
  • Module Leader(s): Professor Nick Wright
  • Lecturer: Professor Damian Giaouris
  • Owning School: Engineering
  • Teaching Location: Âé¶¹´«Ã½ City Campus
Semesters

Your programme is made up of credits, the total differs on programme to programme.

Semester 2 Credit Value: 20
ECTS Credits: 10.0
European Credit Transfer System

Aims

Introducing the fundamentals of Reinforcement Learning:

• The core concepts behind Reinforcement Learning.
• Applications of Reinforcement Learning in real world scenarios.
• The common computer tools for Reinforcement Learning.
• How to appropriately apply these toolkits to real-world problems.
• Make students fluent in the approaches of Reinforcement Learning which will allow them:
o to build up the skills to enable them to apply these skills for applications in companies,
academia or the third sector.
o Know the limits and ethical impacts of using such tools to enable them to make appropriate
decisions when working.
o Make students aware of good practices and how to ensure their results are reported in a fair
and honest manner

Outline Of Syllabus

Material will cover such topics as, whilst also including new and emerging areas of research:

Core foundations include understanding the agent-environment interaction framework, where agents learn by taking actions and receiving rewards. This involves Markov Decision Processes (MDPs), which formalize sequential decision-making problems with states, actions, transition probabilities, and reward functions.

Value-based methods focus on learning value functions that estimate how good states or state-action pairs are. This includes techniques like Q-learning, SARSA, and Deep Q-Networks (DQN), which have been particularly successful in domains like game playing.

Policy-based methods directly learn a policy that maps states to actions, using approaches like policy gradient methods, REINFORCE, and actor-critic algorithms that combine value and policy learning.

Exploration versus exploitation addresses the fundamental trade-off between trying new actions to discover better strategies versus using known good actions to maximize rewards. This includes epsilon-greedy strategies, Upper Confidence Bound methods, and more sophisticated exploration techniques.

Function approximation and deep RL deal with scaling RL to complex, high-dimensional problems using neural networks. This has enabled breakthroughs in areas like Atari games, robotics, and Go.

Model-based RL involves learning a model of the environment's dynamics to enable planning and more sample-efficient learning, contrasting with model-free approaches that learn directly from experience.

Core software systems for developing Reinforcement Learning

Teaching Methods

Teaching Activities
Category Activity Number Length Student Hours Comment
Guided Independent StudyAssessment preparation and completion82:0016:00Examples and Tutorial sheets on topics covered (approximately 2 hours per section of course).
Scheduled Learning And Teaching ActivitiesLecture251:0025:00In-person lectures
Guided Independent StudyAssessment preparation and completion12:002:00Completion of final exam
Guided Independent StudyAssessment preparation and completion112:0022:00Revision for final exam
Scheduled Learning And Teaching ActivitiesPractical53:0015:00Practical sessions
Guided Independent StudyIndependent study1100:00100:00Reviewing lecture notes; general reading, self-study
Guided Independent StudyIndependent study102:0020:00Reading activity to supplement knowledge of material taught in each week.
Total200:00
Teaching Rationale And Relationship

Lectures provide the core material and give students the opportunity to engage with set questions and query material covered in the lecture.

Problem solving is introduced through tutorial sheets and class examples will help students’ understanding of each topic.

Practical sessions provide an opportunity to gain practical experience with a variety of systems and validate the theory introduced in lectures.

Assessment Methods

The format of resits will be determined by the Board of Examiners

Exams
Description Length Semester When Set Percentage Comment
Written Examination1202A1002 hour In-Person closed-book exam
Formative Assessments

Formative Assessment is an assessment which develops your skills in being assessed, allows for you to receive feedback, and prepares you for being assessed. However, it does not count to your final mark.

Description Semester When Set Comment
Prob solv exercises2MTutorial sheets and Class Examples Released after a topic is complete (total of 8 tutorial example sheets – each expected to take 2 hours to complete)
Assessment Rationale And Relationship

The examination allows students to demonstrate their understanding of reinforcement learning and to apply this to realistic problems

The formatively assessed tutorial sheets will be released throughout the semester after each topic is completed.

Reading Lists

Timetable

  • Timetable Website: