Staff ML Inference Engineer - Safety Platform & Scale

Socket.dev

San Francisco (CA)

Hybrid

USD 320,000 - 485,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity donation matching
Generous vacation
Parental leave
Flexible hours
SF office access

Job summary

Anthropic is seeking engineers to design, build, and operate production ML infrastructure that powers Claude's safety systems. You will work at the intersection of ML, distributed systems, and AI safety to scale classifiers and ML defenses across platforms.

The role emphasizes reliability, collaboration with researchers, and delivering production-ready safety tooling. Strong candidates will have 5+ years in production ML infrastructure, experience with transformers, monitoring systems, and

Qualifications

  • Proficient in Python and ML frameworks like PyTorch, TensorFlow, or JAX.
  • Understand distributed systems principles and have built systems that handle high-throughput, low-latency workloads.
  • Have built automated or self-service deployment pipelines and eval infrastructure allowing researchers to roll out classifiers and models independently.
  • Have implemented A/B testing frameworks and experimentation infrastructure for ML systems.
  • Are results-oriented, with a bias towards reliability and impact in safety-critical systems.
  • Enjoy collaborating with researchers and translating cutting-edge research into production systems.
  • Care deeply about AI safety and the societal impacts of your work.

Responsibilities

  • Design and build scalable ML infrastructure to support real-time safety deployments across our classifier and model ecosystem
  • Build monitoring and observability tools to track classifier performance, data quality, and system health for safety-critical applications
  • Collaborate with research teams to productionize safety research, translating experimental safety techniques into robust, scalable systems
  • Optimize inference latency and throughput for real-time safety evaluations while maintaining high reliability standards
  • Implement automated testing, deployment, and rollback systems for ML models in production safety applications
  • Partner with Safeguards, Security, and Alignment teams to understand requirements and deliver infrastructure that meets safety and production needs
  • Contribute to the development of internal tools and frameworks that accelerate safety research and deployment

Skills

Python
ML frameworks
Distributed systems
Deployment pipelines
A/B testing
Safety research collaboration

Education

Bachelor's degree

Job description

Anthropic is seeking engineers to design, build, and operate production ML infrastructure that powers Claude's safety systems. You will work at the intersection of ML, distributed systems, and AI safety to scale classifiers and ML defenses across platforms.

The role emphasizes reliability, collaboration with researchers, and delivering production-ready safety tooling. Strong candidates will have 5+ years in production ML infrastructure, experience with transformers, monitoring systems, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff+ ML Safety Infrastructure Engineer
Staff+ ML Safety Infrastructure Engineer

Anthropic • San Francisco (CA)

On-site
USD 180,000 - 280,000
Health insurance
Dental insurance
Vision insurance
+8
Staff ML Inference Systems Engineer, Real-Time Safety
Staff ML Inference Systems Engineer, Real-Time Safety

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 272,000 - 368,000
Staff Plus Software Engineer (ML Inference Path)
Staff Plus Software Engineer (ML Inference Path)

Anthropic • San Francisco (CA)

On-site
USD 180,000 - 280,000
Health insurance
Dental insurance
Vision insurance
+8
Staff+ Software Engineer, ML Inference Path
Staff+ Software Engineer, ML Inference Path

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 272,000 - 368,000
Scale ML Infrastructure Engineer for Safe AI
Scale ML Infrastructure Engineer for Safe AI

Anthropic • New York (NY)

Hybrid
USD 350,000 - 500,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1
Staff ML Systems Engineer - Safeguards Path
Staff ML Systems Engineer - Safeguards Path

Socket.dev • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
ML Systems Engineer: RL Training & AI Safety
ML Systems Engineer: RL Training & AI Safety

Anthropic • Seattle (WA)

Hybrid
USD 520,000 - 850,000
Equity donation matching
Generous vacation
Parental leave
+2
Engineering Manager, Safety ML — Production & Equity
Engineering Manager, Safety ML — Production & Equity

Discord • San Francisco (CA)

On-site
USD 272,000 - 306,000
Engineering Manager, Safety ML — Lead Production ML Systems
Engineering Manager, Safety ML — Lead Production ML Systems

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 272,000 - 306,000
Equity
Benefits
Staff+ Site Reliability Engineer, Safeguards ML Infra
Staff+ Site Reliability Engineer, Safeguards ML Infra

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000