MLOps Engineer (JAX, PyTorch, Pallas/Triton)

Weekday AI

United States

Remote

USD 96,000 - 152,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Weekday AI in the United States seeks an experienced MLOps Engineer to advance frontier AI systems. You will design and evaluate large-scale ML infrastructure, collaborate with researchers, and help build robust training pipelines using JAX, PyTorch, and custom GPU kernels.

This is a fully remote, 40-hour-per-week engagement. You will work with cross‑functional teams, review tasks, and contribute to performance optimization and scalable ML systems across the development lifecycle.

Qualifications

  • Minimum 2 years of professional experience in MLOps, ML infrastructure, or ML systems engineering.
  • Hands-on production experience with JAX and/or PyTorch in large-scale ML environments.
  • Practical experience developing or optimizing custom GPU kernels using Pallas (JAX) or Triton.
  • Strong understanding of distributed training systems, model optimization, and scalable ML infrastructure.
  • Availability to work 40 hours per week during standard weekday business hours.

Responsibilities

  • Partner with research and engineering teams to strengthen AI model capabilities in MLOps, ML infrastructure, and large-scale training systems.
  • Design challenging, real-world MLOps and machine learning systems tasks that reflect production engineering scenarios.
  • Develop accurate, well-documented solutions to complex ML infrastructure and training pipeline problems.
  • Review and evaluate technical tasks and AI-generated solutions, providing clear and actionable written feedback.
  • Create detailed evaluation rubrics and scoring frameworks for topics including: distributed training architectures, ML pipeline design, infrastructure optimization, kernel-level programming, and performance tuning.

Skills

MLOps
JAX
PyTorch
GPU kernels
Pallas
Triton
Distributed training
ML infrastructure

Tools

Pallas
Triton

Job description

This role is for one of our clients

Compensation: $70-$110 per hour

Join a cutting-edge AI research initiative at the forefront of Generative AI and contribute to the development of next-generation Large Language Models. We are seeking experienced MLOps Engineers with deep expertise in modern machine learning frameworks, large-scale training infrastructure, and kernel-level optimization.

In this role, you'll leverage your knowledge of JAX, PyTorch, and custom GPU kernel programming (Pallas/Triton) to create, evaluate, and refine high-quality technical tasks that help train frontier AI systems. You'll collaborate with AI researchers and engineering teams to improve model reasoning across MLOps, distributed training, and ML infrastructure topics.

This is a full-time, 40-hour-per-week remote engagement requiring full weekday availability.

Key Responsibilities
  • Partner with research and engineering teams to strengthen AI model capabilities in MLOps, ML infrastructure, and large-scale training systems.
  • Design challenging, real-world MLOps and machine learning systems tasks that reflect production engineering scenarios.
  • Develop accurate, well-documented solutions to complex ML infrastructure and training pipeline problems.
  • Review and evaluate technical tasks and AI-generated solutions, providing clear and actionable written feedback.
  • Create detailed evaluation rubrics and scoring frameworks for topics including:
    • Distributed training architectures
    • ML pipeline design
    • Infrastructure optimization
    • Kernel-level programming
    • Performance tuning
  • Collaborate with fellow subject matter experts to maintain consistency, quality, and technical accuracy across training datasets.
  • Contribute domain expertise to improve the reasoning capabilities of advanced AI systems.
Required Qualifications
  • Minimum 2 years of professional experience in MLOps, Machine Learning Infrastructure, or ML Systems Engineering within a recognized technology organization.
  • Hands-on production experience with JAX and/or PyTorch in large-scale machine learning environments.
  • Practical experience developing or optimizing custom GPU kernels using Pallas (JAX) or Triton.
  • Strong understanding of distributed training systems, model optimization, and scalable ML infrastructure.
  • Demonstrated career growth and increasing technical responsibility.
  • Availability to work 40 hours per week during standard weekday business hours.
  • Excellent written communication skills with the ability to clearly explain technical concepts and architectural decisions.
Preferred Skills
  • Experience designing and optimizing large-scale ML training pipelines.
  • Knowledge of distributed computing and GPU performance optimization.
  • Familiarity with evaluation methodologies for AI models and ML systems.
  • Experience collaborating with research teams on advanced machine learning projects.
  • Passion for advancing AI infrastructure and frontier model development.
Why Join
  • Help build and improve next-generation Large Language Models.
  • Work alongside leading AI researchers and experienced machine learning engineers.
  • Apply your expertise to high-impact projects involving large-scale ML systems and infrastructure.
  • Contribute directly to the development of cutting-edge AI technologies.
  • Enjoy a fully remote engagement with meaningful technical challenges.
Equal Opportunity

We welcome applications from qualified professionals regardless of legally protected characteristics and are committed to providing reasonable accommodations throughout the application and engagement process upon request.

Contract & Payment Terms
  • Engagement is offered on an independent contractor basis.
  • This is a fully remote opportunity that can be completed according to your own schedule.
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • The engagement does not require access to confidential or proprietary information belonging to any current employer, client, or institution.
  • Payments are processed weekly through Stripe or Wise based on approved work completed.
  • Please note: Applicants requiring H-1B sponsorship or participating in the STEM OPT program are not eligible for this opportunity.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer - Fully Remote | Upto $110/hr
MLOps Engineer - Fully Remote | Upto $110/hr

Obsidian • New York (NY)

Remote
USD 100,000 - 130,000
MLOps Engineer, LLM Systems · Mercor Mercor · 90-120/hr · remote in US, UK, CA · 2w ago 90-120/hr remote in US, UK, CA 2w ago
MLOps Engineer, LLM Systems · Mercor Mercor · 90-120/hr · remote in US, UK, CA · 2w ago 90-120/hr remote in US, UK, CA 2w ago

Benture • California (MO), Northern (KY)

Hybrid
USD 124,000 - 165,000
ML Systems Engineer - Fully Remote | Upto $110/hr
ML Systems Engineer - Fully Remote | Upto $110/hr

Remote Jobs • United States

Remote
USD 96,000 - 152,000
MLOps Engineer: JAX/PyTorch & GPU Kernel Expert — Remote
MLOps Engineer: JAX/PyTorch & GPU Kernel Expert — Remote

Weekday AI (YC W21) • United States

On-site
USD 96,000 - 152,000
Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour
Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour

24 Mag • New York (NY)

Remote
USD 124,000 - 165,000
MLOps Engineer - AI Trainer
MLOps Engineer - AI Trainer

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 160,000
MLOps Engineer - GPU Specialist
MLOps Engineer - GPU Specialist

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000
MLOps Engineer - GPU Specialist
MLOps Engineer - GPU Specialist

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
Remote MLOps Engineer—JAX/PyTorch, Triton Kernels
Remote MLOps Engineer—JAX/PyTorch, Triton Kernels

Weekday AI • United States

Remote
USD 96,000 - 152,000
MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling)
MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling)

HumanitApp • Northern (KY)

Hybrid
USD 257,887,000 - 343,849,000