< Agenda

The 'Enlightenment' of LLMs and How to Exploit Them

15:00 (10 minutes) · Room 1B · Lightning · Introductory · Security

Corey Duffy
Corey DuffySoftware Engineer/Tech Lead

Your system prompt, a user's message, and the contents of a scraped webpage all arrive at a language model as one flat stream of tokens. There are tags marking who said what, but models largely ignore them, inferring roles from what text sounds like instead. When tone and tag disagree, it looks like tone wins. This talk walks through why the most interesting attacks don't come from users typing "ignore previous instructions": they come from a Jira ticket, a PR description, or a tool response that quietly tells your agent to do something else. I'll also look at why writing a stern system prompt just isn't enough to protect your agentic system. Key takeaways: - Why LLMs can't distinguish instructions from data. - How indirect prompt injection reaches your agent through the content it retrieves, not necessarily just the input that it's given. - An overview of some practical mitigation strategies.

Corey Duffy

Hi, I'm Corey. I'm a senior software engineer and tech lead, with over a decade of experience building software. These days I'm mostly working on AI products and digital intelligence solutions for enterprise clients, while trying to keep the robots honest. Over the years I've worked across application security automation, fraud prevention, usage-based billing, AI knowledge agents, computer vision projects and more, taking things from early prototype to full-scale systems.