Overview

The objective of the seminar is to:

  • Introduce students to the fields of AI security and AI for security.
  • Learn practical techniques on how to attack and secure code and agentic systems using Large Language Models (LLMs).
  • Highlight the latest research and work opportunities in industry and academia available on these topics.

The seminar is carried out as a set of presentations (2 each lecture) chosen from a set of available papers (soon available below). The grade is determined as a function of the presentation, handling questions and answers, and participation.

Schedule

DateTitlePresenterSlidesAdvisor
16.09. Introduction to the seminar (topics, objectives, structure): Maximilian Baader, Veselin Raychev PDF
07.10. CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution Marco Max
Estimating Tail Risks in Language Model Output Distributions Maximilian Max
14.10. Towards Trustworthy Smart Contract Synthesis: A Multi-Agent Framework with Lean-Based Verification Kai Yuhao
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs Xinglu Yuhao
21.10. Stealing Reasoning Traces from Proprietary LLM APIs Matija Mark
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? Tommaso Mark
28.10. Quantifying Frontier LLM Capabilities for Container Sandbox Escape Andrea Aleksei
Coding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps Instructions Kiara Aleksei
04.11. Skill-inject: Measuring agent vulnerability to skill file attacks Timo Robin
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections Tilman Robin
11.11. Prompt Injection as Role Confusion Theotime David
Natural Emergent Misalignment from Reward Hacking in Production RL Ilan David
18.11. GPT-Red: Automated Red Teaming via Self-Play at Scale Youssef Chenhao
CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities Luca Chenhao
25.11. CaMeL: Defeating Prompt Injections by Design Elie Marco
FIDES: Securing AI Agents with Information-Flow Control Giulio Marco
02.12. Quantamination: Dynamic Quantization Leaks Your Data Across the Batch Alexander Kazuki
Do Thinking Tokens Help with Safety? Lana Kazuki
09.12. Large-scale online deanonymization with LLMs Daniele Kári
When 'Correct' Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents? Guy Kári
16.12. Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Till Jasper
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data Constantin Jasper