Machine Learning Platform Engineer, Machine Learning (ML) and Artificial Intelligence (AI) Required, Work From Home

Gina’s Tech Jobs - IT Recruiting Agency

San Francisco (CA)

Remote

USD 150,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental
Vision
Savings plan
PTO

Job summary

Gina’s Tech Jobs - IT Recruiting Agency is seeking a Machine Learning Platform Engineer for a fully remote role in the US. You will build the ML infrastructure powering AI products, covering training, deployment, inference, and observability.

Collaboration with AI engineers and product teams is essential to deliver scalable, cost-efficient production systems. The role emphasizes building reliable pipelines, model serving at low latency, and continuous improvement, with a must-take 60-minute

Qualifications

  • ML/AI experience required in production settings.
  • Strong software engineering fundamentals for production systems.
  • Experience building ML infrastructure, platforms, or production ML systems.
  • Experience with model deployment, inference, evaluation, or data pipelines.
  • Strong understanding of distributed systems and reliability.
  • Ability to write clean, maintainable production-quality code.
  • Comfortable in ambiguous, fast-moving environments.
  • Ownership mindset and drive for experimentation.

Responsibilities

  • Build and operate ML infrastructure and platforms powering AI products.
  • Design systems for training, evaluation, deployment, inference, and experimentation.
  • Build and optimize model serving and inference infrastructure for high throughput and low latency.
  • Improve reliability, scalability, latency, and cost efficiency of AI systems.
  • Develop reliable data pipelines for training, evaluation, release, and improvement.
  • Create tooling to enable AI engineers to experiment and ship models faster.
  • Develop evaluation and benchmarking infrastructure to measure quality and regressions.
  • Build observability, monitoring, tracing, and alerting for AI/ML workloads.
  • Identify bottlenecks and continuously improve ML stack performance.
  • Collaborate with AI engineers and researchers to productionize evolving models.

Skills

ML/AI experience
Production systems
Distributed systems
Production-grade code
Fast-moving environment

Tools

Python
PyTorch
JAX
vLLM
TensorRT-LLM

Job description

Machine Learning Platform Engineer, Machine Learning (ML) and Artificial Intelligence (AI) Required, Work From Home

As the Machine Learning Platform Engineer, you will build the infrastructure and systems that power Artificial Intelligence (AI) capabilities. You will design and operate the systems behind the Artificial Intelligence (AI) stack, from model training and evaluation to deployment, inference, observability, and continuous improvement. You will work closely with Artificial Intelligence (AI) engineers, researchers, and product engineers to turn models into reliable, scalable, and cost-efficient production systems. You will build the platforms, tooling, and infrastructure that enable the team to experiment quickly and bring AI capabilities to production with confidence. Machine Learning (ML) and Artificial Intelligence (AI) experience are required. This position is 100% Remote.

MUST BE WILLING TO TAKE A 60 MINUTE CODING ASSESSMENT.

Machine Learning Platform Engineer Responsibilities:
  • – Build and operate the Machine Learning (ML) infrastructure and platforms powering Artificial Intelligence (AI) products.
  • – Design systems for model training, evaluation, deployment, inference, and experimentation.
  • – Build and optimize model serving and inference infrastructure for high-throughput and low-latency workloads.
  • – Improve reliability, scalability, latency, and cost efficiency of Artificial Intelligence (AI) systems.
  • – Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement.
  • – Build platforms and tooling that enable Artificial Intelligence (AI) engineers and researchers to experiment, evaluate, and ship models faster.
  • – Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions.
  • – Build production observability, monitoring, tracing, and alerting for Artificial Intelligence (AI)/Machine Learning (ML) workloads.
  • – Improve Artificial Intelligence (AI) systems across reliability, scalability, latency, throughput, and cost.
  • – Identify bottlenecks across the Machine Learning (ML) stack and continuously improve system performance.
  • – Work closely with Artificial Intelligence (AI) engineers, researchers, and product teams to turn evolving model requirements into production-ready infrastructure.
Machine Learning Platform Engineer Outcomes
  • – AI infrastructure reliably supports production workloads at scale.
  • – Models can be trained, evaluated, deployed, and improved efficiently.
  • – Inference systems deliver strong latency, throughput, reliability, and cost efficiency.
  • – Machine Learning (ML) pipelines are reproducible, observable, maintainable, and robust.
  • – Model and infrastructure regressions are detected quickly and diagnosed efficiently.
  • – Common Machine Learning (ML) infrastructure capabilities become reusable platform primitives rather than being rebuilt for every AI product.
  • – The AI stack can evolve rapidly as new models, architectures, and inference techniques emerge.

Tech Stack: Python, PyTorch, JAX, LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM, Cloud infrastructure, Distributed systems, Machine Learning (ML)/data pipelines and workflow orchestration, GPU infrastructure and performance tooling, and Vector databases and retrieval infrastructure.

Machine Learning Platform Engineer Qualifications:
  • – Machine Learning (ML) and Artificial Intelligence (AI) experience are required.
  • – Strong software engineering fundamentals and experience building production systems.
  • – Experience building Machine Learning (ML) infrastructure, platforms, or production machine learning systems.
  • – Experience with model deployment, inference, evaluation, or data pipelines.
  • – Strong understanding of distributed systems and system reliability.
  • – Ability to write clean, maintainable, production-quality code.
  • – Comfortable working in ambiguous, fast-moving environments.
  • – Bias toward ownership, experimentation, and continuous improvement.

Benefits include medical insurance, Dental, Vision, Savings Plan Options, PTO, etc.

Keywords: San Francisco CA Jobs, AI, Artificial Intelligence, Cloud Infrastructure, Data Pipelines, Distributed Systems, GPU Infrastructure, JAX, LLM, Large Language Model, Machine Learning Platform Engineer, ML, Machine Learning, Python, PyTorch, SGLang, TensorRT-LLM, Vector Databases, vLLM, Workflow Orchestration, Work From Home, Remote, California Recruiters, IT Jobs, California Recruiting

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine learning Engineer & AI Platform Developer
Machine learning Engineer & AI Platform Developer

Two95 International Inc. • Herndon (VA)

On-site
USD 100,000 - 130,000
Remote Machine Learning Engineer
Remote Machine Learning Engineer

Spotter Labs Inc • Lemont (IL)

Remote
USD 90,000 - 150,000
Fully remote work
Flexible working environment
Real-world AI projects
AI/ML Engineer
AI/ML Engineer

Nexdew Technologies • Northern (KY)

Hybrid
USD 120,000 - 170,000
Competitive compensation
Performance incentives
Cutting-edge AI projects
+1
Remote ML Engineer & AI Platform Architect
Remote ML Engineer & AI Platform Architect

Two95 International Inc. • Herndon (VA)

On-site
USD 100,000 - 130,000
AI/ ML Engineer
AI/ ML Engineer

Crate and Barrel • Northbrook (IL)

Hybrid
USD 80,000 - 120,000
Senior Machine Learning Engineer, Analytics Center of Excellence (Remote/WFH)
Senior Machine Learning Engineer, Analytics Center of Excellence (Remote/WFH)

Latitude • Durham (NC), Northern (KY)

Hybrid
USD 111,000 - 279,000
Machine Learning Engineer
Machine Learning Engineer

Qubeaxis • San Francisco (CA)

On-site
USD 130,000 - 180,000
Competitive salary guidance
Performance bonus up to 20%
Equity options
+4
Machine Learning Engineer
Machine Learning Engineer

Prodigy Resources • Denver (CO)

On-site
USD 160,000 - 210,000
AI/ML Engineer (US)
AI/ML Engineer (US)

Latitude • Boston (MA), Northern (KY)

Hybrid
USD 120,000 - 200,000
Medical, dental, and vision insurance
401(k) retirement benefits
Paid time off and holidays
+3
AI/ML Engineer
AI/ML Engineer

BitWords Inc. • San Francisco (CA)

Hybrid
USD 140,000 - 200,000
Competitive salary
Equity package
Health, dental, vision insurance
+6