ML Systems Engineer - AI Trainer

Obsidian

San Francisco (CA)

On-site

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian is seeking MLOps Engineers to join their innovative GenAI team. This full-time position involves building foundational AI models, training, and evaluating AI systems while providing feedback and developing evaluation frameworks.

The ideal candidate has 2+ years of experience in ML infrastructure and hands-on skills with PyTorch, along with the ability to engage consistently for 40 hours a week.

Qualifications

  • 2+ years of dedicated professional experience in ML infrastructure, MLOps, or ML systems engineering at a recognized organization.
  • Hands-on production experience with PyTorch at scale.
  • Experience writing or optimizing custom GPU kernels using Triton or Pallas.

Responsibilities

  • Guide research and engineering teams to improve AI model performance.
  • Design domain-relevant tasks and write solutions to MLOps problems.
  • Evaluate MLOps tasks and provide technical feedback.
  • Develop guidelines for assessing training pipeline design.
  • Collaborate with subject matter experts for training data consistency.

Skills

Experience in ML infrastructure
Hands-on production experience with PyTorch
Knowledge of GPU kernels (Triton/Pallas)
Strong written communication skills

Job description

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models.

Overview

Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented MLOps Engineers with deep, hands-on expertise in PyTorch and kernel-level programming (Triton/Pallas). This role involves AI model training and evaluation work, including writing and assessing MLOps tasks and solutions to generate high-quality training data for frontier AI systems.

This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. This is a 40-hour full-time engagement, with no conflicts/no other engagements.

Key Responsibilities
  • Guide research and engineering teams to close knowledge gaps and improve AI model performance in MLOps, training infrastructure, and ML framework-level topics.
  • Design challenging, domain-relevant tasks, and write accurate and well-structured solutions to MLOps and ML systems problems.
  • Evaluate MLOps tasks and solutions and provide clear, written technical feedback.
  • Develop guidelines and detailed rubrics/evaluation frameworks to assess training pipeline design, distributed systems reasoning, and kernel-level optimization across tasks.
  • Collaborate with other subject matter experts to ensure consistency and accuracy in training data.
Core Qualifications
  • 2+ years of dedicated professional experience in ML infrastructure, MLOps, or ML systems engineering at a recognized, top-tier organization.
  • Hands-on production experience with PyTorch at scale.
  • Experience writing or optimizing custom GPU kernels using Triton or Pallas.
  • Demonstrable career progression.
  • Ability to engage reliably for at least 40 hours/week during weekdays.
  • Strong written communication skills and the ability to explain complex technical decisions clearly.
Equal Employment Opportunity

Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PyTorch MLOps Engineer - AI Specialist
PyTorch MLOps Engineer - AI Specialist

Obsidian • New York (NY)

On-site
USD 120,000 - 140,000
MLOps Engineer - AI Trainer
MLOps Engineer - AI Trainer

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 160,000
GenAI PyTorch MLOps Engineer (Triton/Pallas)
GenAI PyTorch MLOps Engineer (Triton/Pallas)

aitrainer • United States

On-site
Machine Learning Engineer — Model Evaluation & Experimentation
Machine Learning Engineer — Model Evaluation & Experimentation

Dorado • United States

Remote
USD 120,000 - 180,000
GenAI MLOps Engineer - JAX/PyTorch & GPU Kernels
GenAI MLOps Engineer - JAX/PyTorch & GPU Kernels

Dorado • United States

Remote
USD 120,000 - 160,000
GenAI MLOps Engineer & AI Training Systems
GenAI MLOps Engineer & AI Training Systems

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
GenAI MLOps Engineer — Build Frontier AI Models
GenAI MLOps Engineer — Build Frontier AI Models

Obsidian • New York (NY)

Remote
USD 100,000 - 130,000
Machine Learning Engineer
Machine Learning Engineer

Errgo • Town of Boston (NY)

Hybrid
USD 120,000 - 160,000
Medical, dental, and vision insurance
401(k)
Equity
+2