Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Anthropic in the United Kingdom is seeking a software engineer for the Safeguards team to build safety and oversight mechanisms for AI systems. You will monitor models, prevent misuse, and surface abuse patterns to researchers and analysts for action.
Responsibilities include building scalable detection systems, real-time defenses, and internal tooling, with a focus on safety, transparency, and policy enforcement across the stack. Premier compensation and relocation support are offered.
Proficiency in Python and TypescriptStrong communication skills and ability to explain complex technical concepts to non-technical stakeholdersAbility to work across the stackBachelor’s degree in Computer Science, Software Engineering or comparable experienceWe encourage you to apply even if you do not believe you meet every single qualificationHave experience with integrity, spam, fraud, or abuse detection and mitigation8+ years of experience in a software engineering positionHave experience building trust and safety detection mechanisms and intervention for AI/ML systemsHave experience with prompt engineering, jailbreak attacks, and other adversarial inputsHave worked closely with operational teams to build custom internal tooling