AI safety and information operations

I build evaluation systems that test frontier language models against real-world harm data instead of synthetic prompts.

  • InfoOps Bench is a live benchmark. It measures whether frontier models can be co-opted into authoritarian information operations. It draws on more than 2,100 documented operations from a live monitoring pipeline, and scores 17 models from 8 providers across four prompt framings. Integrity scores range from 9.3 percent to 91 percent. The benchmark refreshes weekly, so it resists saturation. arXiv:2607.28503
  • A live AI-harms intelligence system ingests documented model failures from the open web, converts each into an automated evaluation, replays it against target models, and tracks how failures reproduce across model versions.
  • I also red-team pre-release frontier models, designing adversarial evaluation suites that surface jailbreak, safety, and misuse failure modes before deployment.

Political attitudes and polarization in online networks

This is the core of my PhD at the University of Zurich, supervised by Prof. Alexandre Bovet.

  • Affective polarization on Bluesky asks how hostility toward an out-group survives in a space where that out-group is barely present. Work in progress.
  • Bluesky: network topology, polarization, and algorithmic curation maps the follow graph and the effect of custom feeds. PLoS ONE
  • Anatomy of politically homogeneous spaces studies how and why politically uniform spaces still fragment. arXiv:2506.03443

Multilingual misinformation

  • Lost in translation uses global fact-checks to measure how misinformation spreads and mutates across languages. EPJ Data Science
  • Language mutations shows how COVID-19 conspiracy theories persist by changing form as they cross languages. Computers in Human Behavior
  • The perils and promises of fact-checking with large language models tests whether LLMs can act as fact-checkers. Frontiers in AI

Platform migration

  • Simple contagion drives population-scale platform migration follows around 300,000 academics moving from Twitter to Bluesky. It won the Best Student Paper Award at the Networks in Science of Science satellite, NetSci 2025. arXiv:2505.24801

State media monitoring

At Pattrn AI I conceived and built a monitoring system for Russian, Chinese, and Iranian state-backed and proxy media. It spans more than 186 billion views and 1.9 billion engagements, across 219 languages and 247 countries and territories. Machine learning models map influence networks and flag newly emerging outlets.