Research Lead
Far Ai · Berkeley Office · 2026-04-24
About this role
FAR.AI http://FAR.AI is hiring a Research Lead to develop and lead a research agenda that reduces catastrophic risks from advanced AI. You'll build and lead a team executing this agenda — setting research direction, mentoring Members of Technical Staff to scale your vision, and remaining hands-on enough to write code and run experiments yourself. What counts is whether AI labs and governments actually change how they act; publications are useful but aren't the measure. Beyond your team, you can shape FAR.AI http://FAR.AI's broader work by directing millions of dollars in grants to external researchers extending your agenda, convening the people who can act on it, and influencing our independent testing and advising of AI companies and governments. This role suits you if you want high autonomy in an impact-driven environment, pursuing empirically grounded, scalable ML safety work.
ABOUT US
FAR.AI http://FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.
Since our founding in July 2022, we've grown to 40+ staff https://www.far.ai/about/team, published 40+ academic papers https://scholar.google.com/citations?user=FVJ24k8AAAAJ, and convened leading AI safety events https://far.ai/events/. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML, and ICLR, and features in the Financial Times https://www.ft.com/content/175e5314-a7f7-4741-a786-273219f433a1, Nature News https://www.nature.com/articles/d41586-024-02218-7 and MIT Technology Review https://www.technologyreview.com/2020/02/28/905615/reinforcement-learning-adversarial-attack-gaming-ai-deepmind-alphazero-selfdriving-cars/. We conduct pre-deployment testing on behalf of frontier developers such as OpenAI and independent evaluations for governments including the EU AI Office https://www.far.ai/news/far-ai-selected-to-lead-eu-ai-act-cbrn-risk-consortium. We help steer and grow the AI safety field through developing https://arxiv.org/abs/2405.06624 research https://arxiv.org/abs/2506.20702 roadmaps https://www.researchgate.net/publication/396910034_Open_Technical_Problems_in_Open-Weight_AI_Model_Risk_Management with renowned researchers such as Yoshua Bengio; running FAR.Labs https://www.far.ai/programs/far-labs, an AI safety-focused co-working space in Berkeley housing 40 members; and supporting the community through targeted grants https://www.far.ai/programs/grantmaking to technical researchers.
ABOUT FAR.RESEARCH
We explore promising research directions in AI safety and scale up only those showing a high potential for impact. Once the core research problems are solved, we work to scale them to a minimum viable prototype, demonstrating their validity to AI companies and governments to drive adoption.
Our recent and ongoing research includes:
Adversarial Robustness: working to rigorously solve security problems through building a science of security and robustness for AI, from demonstrating superhuman systems can be vulnerable https://far.ai/post/2023-07-superhuman-go-ais/, to scaling laws for robustness https://www.far.ai/news/does-robustness-improve-with-scale and jailbreaking constitutional classifiers https://arxiv.org/abs/2506.24068.
Mechanistic Interpretability: finding https://arxiv.org/abs/2502.12892 issues https://arxiv.org/abs/2508.16560 with https://arxiv.org/abs/2505.11756 Sparse Autoencoders, probing deception using AmongUs https://arxiv.org/abs/2504.04072, understanding learned planning https://far.ai/post/2024-07-learned-planners/ in SokoBan, and interpretable data attribution.
Red-teaming: conducting pre- and post-release adversarial evaluations of frontier models (e.g. Claude 4 Opus https://x.com/ARGleave/status/1926138376509440433, ChatGPT Agent https://cdn.openai.com/pdf/839e66fc-602c-48bf-81d3-b21eacc3459d/chatgpt_agent_system_card.pdf, GPT-5 https://cdn.openai.com/gpt-5-system-card.pdf); developing novel attacks https://www.far.ai/news/defense-in-depth to support this work.
Evals: developing evaluations for new threat models, e.g. persuasion https://arxiv.org/abs/2506.02873 and tampering risks https://arxiv.org/abs/2507.11630.
Mitigating AI deception: studying when lie detectors induce honesty or evasion https://www.far.ai/news/avoiding-ai-deception, and developing approaches to deception and sandbagging.
We are particularly looking to add Research Leads in the following pod shapes:
- Applied Interpretability — using interpretability to tackle concrete safety problems (better probes, backdoor detection, deception monitoring), aiming for fast feedback loops, often in collaboration with our other pods. A new pod, greenfield.
- Scalable Oversight / Alignment — methods that keep oversight robust as models become more capable than their supervisors: recursive reward modeling, debate, weak-to-strong generalization, process-based supervision.
- Adversarial Robustness —extending our independent-testing work into deployed-system protection: better safety guardrails, pre-training safety interventions (initially CBRN misuse, especially for open-weight models), backdoor detection and mitigation, realistic cybersecurity evaluations, and loss-of-control deception evaluations.
- Auditing / Evals — safety and alignment auditing: evaluation awareness (construct validity, safety-relevance, hyper-realistic evals), CoT monitorability and faithfulness training, black-box monitoring as a complement to our existing white-box work.
- Persuasion / Epistemic Risks — science of epistemic risks and intervention points, persuasion's role in loss of control risks, evaluations and independent testing, connections to broader harmful manipulation, solutions and epistemic uplift. Building on our existing work and shaping your own agenda in the area.
- Bring Your Own…
Skills asked for
- go
- cybersecurity
Similar jobs
- Research Lead - Pre-training SafetyFar Ai · Berkeley Office
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.