Olivia Grace Watkins

oliviawatkins @ openai . com

Who am I?

I am a researcher at OpenAI. I want to make sure AI is developed in a safe, accountable way that benefits all of humanity.

Most of my recent work has focused on measuring cybersecurity risks posed by frontier AI systems and helping accelerate cyber defense. More broadly, I am interested in AI policy, forecasting AI progress and risks, external transparency and accountability for AI development, alignment, defensive acceleration, and promoting democratic governance.

If you are working on these problems, please reach out. I am especially excited to chat with researchers, builders, policy people, and unusually determined nerds with big ideas. I may also be able to support small grants for promising projects in these areas.

Do you have a life outside of research?

In my spare time I play Quidditch with the Silicon Valley Vipers and D&D, hang out with friends, make mediocre puns, and procrastinate on keeping my website up to date. I remain robust to most adversarial inputs except chocolate, elaborate character backstories, and someone saying "one quick side quest."

publications

  1. arXiv 2026
    EVMbench: Evaluating AI Agents on Smart Contract Security
    Justin Wang, Andreas Bigger, Xiaohai Xu, Justin W. Lin, Andy Applebaum, Tejal Patwardhan, Alpin Yukseloglu, and Olivia Watkins
    arXiv 2026
  2. ICLR 2026
    GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
    Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, Simon Posada Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Samuel Miserendino, Gildas Chabot, David Li, Patrick Chao, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek
    ICLR Poster 2026
  3. ICLR 2026
    Estimating Worst-Case Frontier Risks of Open-Weight LLMs
    Eric Wallace, Olivia Watkins, Miles Wang, Kai Chen, and Chris Koch
    ICLR Poster 2026
  4. ICLR 2026
    Persona Features Control Emergent Misalignment
    Miles Wang, Tom Dupré la Tour, Olivia Watkins, Alex Makelov, Ryan A. Chi, Samuel Miserendino, Jeffrey Wang, Achyuta Rajaram, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing
    ICLR Poster 2026
  5. NeurIPS 2024
    A StrongREJECT for Empty Jailbreaks
    Alexandra Souly*, Qingyuan Lu*, Dillon Bowen*, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons*, Olivia Watkins*, Sam Toyer*
    NeurIPS Datasets and Benchmarks 2024
  6. ICML 2024 Oral
    Learning to Model the World with Language
    Jessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner, Pieter Abbeel, Dan Klein, and Anca Dragan
    ICML Oral 2024
  7. ICLR 2024 Spotlight
    Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
    Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell
    ICLR Spotlight 2024
  8. ICML
    Guiding Pretraining in Reinforcement Learning with Large Language Models
    Du*, Yuqing Watkins*, Olivia, Wang, Zihan, Colas, Cédric, Darrel, Trevor, Abbeel, Pieter, Gupta, Abhishek, and Andreas, Jacob,
    ICML 2023
  9. NeurIPS
    DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models
    Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Kangwook Lee, Kimin Lee
    NeurIPS 2023
  10. arXiv
    Aligning Text-to-Image Models using Human Feedback
    Lee, Kimin; Liu Hao; Ryu, Moonkyung; Watkins, Olivia; Du, Yuqing; Boutilier, Craig; Abbeel, Pieter; Ghavamzadeh, Mohammad; Gu, Shixiang Shane
    arXiv 2023
  11. NeurIPS
    Teachable Reinforcement Learning via Advice Distillation
    Watkins, Olivia, Darrel, Trevor, Abbeel, Pieter, Andreas, Jacob, and Gupta, Abhishek
    NeurIPS 2021 2021
  12. ICRA
    Auto-Tuned Sim-to-Real Transfer
    Du *, Yuqing; Watkins*, Olivia; Darrell, Trevor; Abbeel, Pieter; and Pathak, Deepak
    ICRA 2021
  13. ICML Workshop
    Explaining Reinforcement Learning Policies through Counterfactual Trajectories
    Frost, Julius; Watkins, Olivia; Weiner, Eric; Abbeel, Pieter; Darrell, Trevor; Plummer, Bryan; and Saenko, Kate
    ICML Workshop on Human in the Loop Learning 2021
  14. ICNLP
    Hierarchical text generation using an outline
    Drissi, Mehdi; Watkins, Olivia; and Kalita, Jugal
    International Conference on Natural Language Processing 2018
  15. ICML Workshop
    Program language translation using a grammar-driven tree-to-tree model
    Drissi*, Mehdi; Watkins*, Olivia; Khant, Aditya; Ojha, Vivaswat; Sandoval, Pedro; Segev, Rakia; Weiner, Eric; and Keller, Robert
    ICML Workshop on Neural Abstract Machines & Program Induction 2018