Machine Learning Engineer

Jobzhr

San Francisco, Northern (CA, KY)

On-site

USD 130,000 - 200,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Performance-based annual bonus
Conferences/continuing education
Fully remote (U.S.-based)
Health benefits
Generous PTO & holidays

Job summary

10a Labs is seeking a Machine Learning Engineer to design, build, and evaluate advanced ML systems across AI safety and model evaluation applications. You will tackle problems involving reinforcement learning, language models, multimodal systems, and classifiers, turning ambiguous questions into rigorous experiments and scalable solutions.

Collaborate with engineers, analysts, red teamers, and subject‑matter experts to advance frontier AI, with focus on safety, robustness, and measurable impact

Qualifications

  • 3–5+ years in machine learning, research engineering, or a related technical field.
  • Strong Python skills and experience with ML frameworks such as PyTorch or JAX.
  • Hands‑on experience training, fine‑tuning, or evaluating modern ML models.
  • Strong understanding of experimental design, model evaluation, and quantitative analysis.
  • Familiarity with agentic AI fundamentals, including common harnesses and related risks.
  • Experience in one or more of reinforcement learning, NLP/LLMs, computer vision, or multimodal ML.
  • Strong software engineering fundamentals and the ability to work independently on ambiguous technical problems.

Responsibilities

  • Design and run ML experiments to evaluate capabilities, behavior, robustness, and limitations of advanced AI systems.
  • Develop and evaluate models across reinforcement learning, NLP/LLMs, computer vision, and multimodal ML.
  • Build evaluation pipelines, benchmarks, datasets, and metrics for frontier AI systems.
  • Train, fine‑tune, and evaluate models for safety, security, and other high‑impact applications.
  • Develop reliable tooling and infrastructure to run ML experiments and evaluations at scale.
  • Analyze results, identify model failure modes, and translate findings into new experiments and technical approaches.

Skills

Machine Learning experience
Python programming
PyTorch or JAX
Experimental design
Model evaluation
Agentic AI fundamentals
Independent problem solving

Job description

About 10a Labs

10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations, and intelligence collection enable engineering, safety, and security teams to stay ahead of evolving threats and deploy AI systems safely.

About the Role

We are seeking a Machine Learning Engineer to design, build, and evaluate advanced machine learning systems across AI safety and model evaluation applications.

This role combines strong ML engineering with an experimental mindset. You will work on problems involving reinforcement learning, model evaluations, language models, multimodal systems, and classifiers, taking ambiguous technical questions and turning them into rigorous experiments and scalable systems.

You will collaborate closely with engineers, analysts, red teamers, and subject‑matter experts supporting leading AI organizations.

What You’ll Do
  • Design and run ML experiments to evaluate the capabilities, behavior, robustness, and limitations of advanced AI systems.
  • Develop and evaluate models across reinforcement learning, NLP/LLMs, computer vision, and multimodal ML.
  • Build evaluation pipelines, benchmarks, datasets, and metrics for frontier AI systems.
  • Train, fine‑tune, and evaluate models for safety, security, and other high‑impact applications.
  • Develop reliable tooling and infrastructure to run ML experiments and evaluations at scale.
  • Analyze results, identify model failure modes, and translate findings into new experiments and technical approaches.
What We’re Looking For
  • 3–5+ years of experience in machine learning, research engineering, or a related technical field.
  • Strong Python skills and experience with ML frameworks such as PyTorch or JAX.
  • Hands‑on experience training, fine‑tuning, or evaluating modern ML models.
  • Strong understanding of experimental design, model evaluation, and quantitative analysis.
  • Familiarity with agentic AI fundamentals, including common harnesses, Model Context Protocol, agent benchmarks, and security risks to AI agents.
  • Experience in one or more of the following: reinforcement learning, NLP/LLMs, computer vision, or multimodal ML.
  • Strong software engineering fundamentals and the ability to work independently on ambiguous technical problems.
Nice to Have
  • Experience with RLHF/RLAIF, reward modeling, policy optimization, or other model post‑training techniques.
  • Experience evaluating frontier language or multimodal models.
  • Experience with adversarial evaluations, robustness testing, or AI safety.
  • Experience with distributed training, cloud ML infrastructure, or large‑scale ML systems.

We don't expect candidates to have experience across every area above. We value deep ML expertise, strong experimental instincts, and the ability to quickly learn new techniques.

Compensation & Benefits
  • Salary Range: $130K–$200K, depending on experience and location
  • Bonus: Performance-based annual bonus
  • Professional Development: Support for conferences, continuing education, or leadership training
  • Work Environment: Fully remote, U.S.-based
  • Health Benefits: Comprehensive health, dental, and vision coverage
  • Time Off: Generous PTO and paid holiday schedule
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote ML Engineer — AI Safety & Evaluation
Remote ML Engineer — AI Safety & Evaluation

Jobzhr • San Francisco (CA), Northern (KY)

Hybrid
USD 130,000 - 200,000
Performance-based annual bonus
Conferences/continuing education
Fully remote (U.S.-based)
+2
Remote ML Safety & Evaluation Engineer
Remote ML Safety & Evaluation Engineer

10a Labs • San Francisco (CA)

On-site
USD 130,000 - 200,000
Performance-based annual bonus
Professional development support
Comprehensive health benefits
+1
Software Engineer, Infrastructure & Platform
Software Engineer, Infrastructure & Platform

10a Labs • United States

Remote
USD 110,000 - 160,000
Fully remote, U.S.-based
Performance-based annual bonus
Professional development support (con-
+1
Senior ML Engineer - Safety & AI Evaluation (Remote)
Senior ML Engineer - Safety & AI Evaluation (Remote)

10a Labs • New York (NY)

Remote
USD 130,000 - 200,000
Comprehensive health, dental, and vision coverage
Performance-based annual bonus
Support for conferences and continuing education
Senior ML Engineer - Safety & AI Evaluation (Remote)
Senior ML Engineer - Safety & AI Evaluation (Remote)

10a Labs • Chicago (IL)

On-site
USD 130,000 - 200,000
Performance-based annual bonus
Comprehensive health, dental, and vision coverage
Support for conferences and continuing education
+1
Machine Learning Lead Engineer
Machine Learning Lead Engineer

Cox • Atlanta (GA)

On-site
USD 134,900 - 224,900
Flexible vacation policy
Paid wellness hours
Paid holidays
Remote ML Engineer - AI Safety & Evaluation Expert
Remote ML Engineer - AI Safety & Evaluation Expert

10a Labs • Seattle (WA)

On-site
USD 130,000 - 200,000
Comprehensive health, dental, and vision coverage
Performance-based annual bonus
Support for professional development and conferences
+1
Machine Learning Engineer
Machine Learning Engineer

Gray Swan AI • Pittsburgh

On-site
USD 140,000 - 225,000
401k with up to 4% matching
28 days annual leave (vacation +bahold
Health, dental, and vision coverage
+3
Machine Learning Engineer
Machine Learning Engineer

Gray Swan • United States

On-site
USD 160,000 - 257,000
401k match
Paid time off
Health insurance
+3
Principal AI/ML Engineer - AI Safety & Evaluation
Principal AI/ML Engineer - AI Safety & Evaluation

A10 Networks, Inc. • San Jose (CA)

On-site
USD 225,000 - 245,000