Overview

The objective of the seminar is to:

  • Introduce students to the fields of AI security and AI for security.
  • Learn practical techniques on how to attack and secure code and agentic systems using Large Language Models (LLMs).
  • Highlight the latest research and work opportunities in industry and academia available on these topics.

The seminar is carried out as a set of presentations (2 each lecture) chosen from a set of available papers (soon available below). The grade is determined as a function of the presentation, handling questions and answers, and participation.

Schedule

DateTitlePresenterSlidesAdvisor
16.09. Introduction to the seminar (topics, objectives, structure): Maximilian Baader, Veselin Raychev
07.10. CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution Max
Estimating Tail Risks in Language Model Output Distributions Max
14.10. Towards Trustworthy Smart Contract Synthesis: A Multi-Agent Framework with Lean-Based Verification Yuhao
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs Yuhao
21.10. Stealing Reasoning Traces from Proprietary LLM APIs Mark
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? Mark
28.10. Quantifying Frontier LLM Capabilities for Container Sandbox Escape Aleksei
Coding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps Instructions Aleksei
04.11. Skill-inject: Measuring agent vulnerability to skill file attacks Robin
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections Robin
11.11. Prompt Injection as Role Confusion David
Natural Emergent Misalignment from Reward Hacking in Production RL David
18.11. GPT-Red: Automated Red Teaming via Self-Play at Scale Chenhao
CyberGym-E2E: Scalable Real-World Benchmark for AI Agents’ End-to-End Cybersecurity Capabilities Chenhao
25.11. CaMeL: Defeating Prompt Injections by Design Marco
FIDES: Securing AI Agents with Information-Flow Control Marco
02.12. Quantamination: Dynamic Quantization Leaks Your Data Across the Batch Kazuki
Do Thinking Tokens Help with Safety? Kazuki
09.12. Large-scale online deanonymization with LLMs Kári
When ‘Correct’ Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents? Kári
16.12. Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Jasper
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data Jasper