GenAI MLOps GPU Engineer - High-Throughput Inference

Obsidian

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Cincinnatus LLC is seeking MLOps Engineers to join a cutting-edge GenAI team building foundational AI models. This hands-on role focuses on large language model infrastructure, GPU kernels, performance profiling, and high-throughput inference serving, with opportunities to contribute to training data for frontier AI systems.

This is a 40-hour per week, full-time engagement as the employer-of-record, with placement at a leading AI Lab.

Qualifications

  • 2+ years hands-on experience in ML systems, ML infrastructure, or model serving.
  • Hands-on in GPU kernels, profiling tools, distributed workloads, or high-throughput inference.
  • Production experience with JAX and/or PyTorch is required.

Responsibilities

  • Design tasks across GPU kernels, profiling, debugging, and inference serving; provide detailed solutions.
  • Guide teams to improve model performance on training infrastructure and framework topics.
  • Evaluate MLOps solutions and provide clear, written feedback.
  • Develop evaluation rubrics for kernel optimization, profiler outputs, and serving latency.
  • Collaborate with experts to ensure training data quality and consistency.

Skills

ML systems
GPU kernel programming
Performance profiling
Distributed workloads
PyTorch/JAX
Technical communication

Tools

Kineto
Torch Profiler
Nsight
XLA
CUDA

Job description

Cincinnatus LLC is seeking MLOps Engineers to join a cutting-edge GenAI team building foundational AI models. This hands-on role focuses on large language model infrastructure, GPU kernels, performance profiling, and high-throughput inference serving, with opportunities to contribute to training data for frontier AI systems.

This is a 40-hour per week, full-time engagement as the employer-of-record, with placement at a leading AI Lab.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU MLOps Engineer for GenAI Inference & Kernel Ops
GPU MLOps Engineer for GenAI Inference & Kernel Ops

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
GenAI MLOps Engineer for LLM Systems
GenAI MLOps Engineer for LLM Systems

Dorado • United States

Remote
USD 140,000 - 220,000
GenAI ML Systems Engineer & AI Trainer
GenAI ML Systems Engineer & AI Trainer

Obsidian • Chicago (IL)

On-site
USD 120,000 - 180,000
GenAI ML Systems Engineer & Trainer (MLOps)
GenAI ML Systems Engineer & Trainer (MLOps)

Mercor • Chicago (IL)

On-site
USD 120,000 - 180,000
MLOps Engineer - GPU Specialist
MLOps Engineer - GPU Specialist

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000
MLOps Engineer - GPU Specialist
MLOps Engineer - GPU Specialist

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
GenAI MLOps Engineer — Build Frontier AI Models
GenAI MLOps Engineer — Build Frontier AI Models

Obsidian • New York (NY)

Remote
USD 100,000 - 130,000
ML Systems Engineer - AI Trainer
ML Systems Engineer - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 120,000 - 180,000
ML Systems Engineer - AI Trainer
ML Systems Engineer - AI Trainer

Mercor • Chicago (IL)

On-site
USD 120,000 - 180,000
GenAI ML Systems Engineer — MLOps & GPU Kernel Expert
GenAI ML Systems Engineer — MLOps & GPU Kernel Expert

Obsidian • New York (NY)

Remote
USD 90,000 - 130,000