Jobs and Careers
AN

Anthropic AI Safety Fellow, UK

Anthropic
United Kingdomfull_timeVerifiedPosted 29 Jul 2025
💰 £67,600/yr

About the role

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

Note: this is our UK job posting. You can find our US and Canada job postings on our careers page.

Please apply by August 17!

Responsibilities:

The Anthropic Fellows Program is an external collaboration program focused on accelerating progress in AI safety research by providing promising talent with an opportunity to gain research experience. The program will run for about 2 months, with the possibility of extension for another 4 months, based on how well the collaboration is going. Our goal is to bridge the gap between industry engineering expertise and the research skills needed for impactful work in AI safety.

  • Fellows will use external infrastructure (e.g. open-source models, public APIs) to work on an empirical project aligned with our research priorities, with the goal of producing a public output (e.g. a paper submission). Fellows will receive substantial support - including mentorship from Anthropic researchers, funding, compute resources, and access to a shared workspace - enabling them to develop the skills to contribute meaningfully to critical AI safety research.
  • We aim to onboard our next cohort of Fellows in October 2025, with later start dates being possible as well.

What To Expect

  • Direct mentorship from Anthropic researchers
  • Connection to the broader AI safety research community
  • Weekly stipend of 1300 GBP & access to benefits (benefits vary by country but include medical, dental, and vision insurance)
  • Funding for compute and other research expenses
  • Shared workspaces in Berkeley, California and London, UK
  • This role will be employed by our third-party talent partner, and may be eligible for benefits through the employer of record.

Mentors & Research Areas

Fellows will undergo a project selection & mentor matching process. Potential mentors include

Our mentors will lead projects in select AI safety research areas, such as:

  • Scalable Oversight: Developing techniques to keep highly capable models helpful and honest, even as they surpass human-level intelligence in various domains. 
  • Adversarial Robustness and AI Control: Creating methods to ensure advanced AI systems remain safe and harmless in unfamiliar or adversarial scenarios.
  • Model Organisms: Creating model organisms of misalignment to improve our empirical understanding of how alignment failures might arise.
  • Model Internals / Mechanistic Interpretability: Advancing our understanding of the internal workings of large language models to enable more targeted interventions and safety measures.
  • AI Welfare: Improving our understanding of potential AI welfare and developing related evaluations and mitigations.

For a full list of representative projects for each area, please see these blog posts: Introducing the Anthropic Fellows Program for AI Safety Research, Recommendations for Technical AI Safety Research Directions.

You may be a good fit if you:

  • Are motivated by reducing c

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Anthropic

View company profile →