Job Description
We are seeking a Senior Staff Engineer, AI Safety to help design, develop, and evaluate safety mechanisms for large language models (LLMs) and other foundation models. This role combines machine learning, model evaluation, adversarial testing, and safety engineering to improve model robustness and reduce the risk of unsafe, non-compliant, or manipulated outputs. The position will provide technical leadership while partnering with research, engineering, and security teams to deploy scalable AI safety solutions.
Key Responsibilities
- Lead the development of safety controls to mitigate jailbreaks, prompt injection attacks, and unsafe model outputs.
- Design and implement evaluation frameworks to identify model weaknesses, regressions, edge cases, and emerging vulnerabilities.
- Develop and support adversarial testing and red-teaming programs to uncover safety risks in deployed and pre-release models.
- Build and improve safety enforcement mechanisms, including:
- Fine-tuning workflows
- Input and output filtering
- Pre-processing and post-processing controls
- Partner with researchers and red teams to convert emerging attack patterns into measurable evaluations and risk metrics.
- Monitor advancements in AI safety research, adversarial prompting techniques, and model exploitation methods, and incorporate findings into production defenses.
- Provide technical leadership and guidance for AI safety initiatives across multiple teams.
Equality & Inclusion
We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations and ordinances.
Skills and Requirements
- 5+ years of experience in machine learning, artificial intelligence, or AI platform engineering.
- 2+ years of experience in a technical leadership role.
- Experience integrating safety mechanisms into machine learning deployment environments, including inference services, moderation systems, or filtering layers.
- Strong understanding of transformer-based architectures and large language models.
- Experience with AI safety, model robustness, interpretability, or related domains.
- Experience evaluating model behavior in adversarial, edge-case, or failure-mode scenarios.
- Strong communication and collaboration skills, with the ability to align stakeholders across technical and non-technical teams.
- Bachelor's, Master\'s, or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field. Experience developing or deploying safety controls for large language models in production environments.
- Hands-on experience with adversarial testing, red teaming, and model vulnerability assessment.
- Experience designing or implementing prompt-level defenses and content moderation systems.
- Familiarity with techniques used to prevent jailbreaks, prompt injection attacks, and policy violations.
- Experience supporting large-scale model evaluation frameworks and automated safety testing pipelines.
- Background in AI safety research, model alignment, robustness, or responsible AI initiatives.