Reinforcement Learning with Declarative Specifications
Lecture slides:
Some reading materials:
- Background on logic (if needed): Wojciech Jamroga. Logical methods for specification and verification of multi-agent systems. pdf
- For lecture 2 on reward machines:
- The main reference is:
Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Anthony Valenzano, Sheila A. McIlraith:
Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning. J. Artif. Intell. Res. 73: 173-208 (2022) pdf Section 5.5 contains a link to a github repository with examples from the paper which illustrate how to incorporate reward machines in training an RL agent.
- L. Illanes, X. Yan, R. Toro Icarte, and S. McIlraith.
Symbolic planning and model-free reinforcement learning: Training taskable agents.
In Proceedings of 4th Multidisciplinary Conference on Reinforcement
Learning and Decision Making (RLDM), pages 191–195, 2019. pdf
- Giovanni Varricchione, Natasha Alechina, Mehdi Dastani, Brian Logan:
Maximally Permissive Reward Machines. ECAI 2024: 1181-1188. pdf
- Giovanni Varricchione, Toryn Q. Klassen, Natasha Alechina, Mehdi Dastani, Brian Logan, Sheila A. McIlraith:
Pushdown Reward Machines for Reinforcement Learning. KR 2025 arxiv pdf
- Rajeev Alur, Suguman Bansal, Osbert Bastani, Kishor Jothimurugan:
A Framework for Transforming Specifications in Reinforcement Learning. Principles of Systems Design 2022: 604-624.
arxiv pdf
- For lecture 3 on multi-agent reward machines:
- Cyrus Neary, Zhe Xu, Bo Wu, Ufuk Topcu:
Reward Machines for Cooperative Multi-Agent Reinforcement Learning. AAMAS 2021: 934-942 pdf
- Giovanni Varricchione, Natasha Alechina, Mehdi Dastani, Brian Logan:
Synthesising Reward Machines for Cooperative Multi-Agent Reinforcement Learning. J. Artif. Intell. Res. 85 (2026)
pdf
- For lecture 4 on shields:
-
Krasowski, H.; Thumm, J.; M¨ uller, M.; Sch¨ afer, L.; Wang, X.; and Althoff, M.
2023. Provably Safe Reinforcement Learning: A Theoretical and Experimental
Comparison. arXiv:2205.06750.
pdf
- Mohammed Alshiekh, Roderick Bloem, Rüdiger Ehlers, Bettina Könighofer, Scott Niekum, Ufuk Topcu:
Safe Reinforcement Learning via Shielding. AAAI 2018: 2669-2678 pdf
-
Ingy Elsayed-Aly, Suda Bharadwaj, Christopher Amato, Rüdiger Ehlers, Ufuk Topcu, Lu Feng:
Safe Multi-Agent Reinforcement Learning via Shielding. AAMAS 2021: 483-491. pdf
- Giovanni Varricchione, Natasha Alechina, Mehdi Dastani, Giuseppe De Giacomo, Brian Logan, Giuseppe Perelli:
Pure-Past Action Masking. AAAI 2024: 21646-21655 pdf
- Bettina Könighofer, Roderick Bloem, Nils Jansen, Sebastian Junges, Stefan Pranger:
Shields for Safe Reinforcement Learning. Commun. ACM 68(11): 80-90 (2025). pdf
- Integration of shields into standard RL libraries
Stefan Pranger, Bettina Könighofer:
Easy-to-Use Shielding for Reinforcement Learning. CoRR abs/2606.03804 (2026)
arxiv pdf
Last updated 10 July 2026