Production ML Engineer: LLMs, RAG & MLOps on GCP

DeepHow

United States

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

DeepHow is seeking an AI Engineer to own our AI pipeline end-to-end, taking over the production ML stack, harden it, and ship improvements fast. Day one, your focus is MLOps and production deployment making sure our models are fast, reliable, and cost-efficient at scale.

This is a build-and-ship role, not research. If you like turning prototypes into production systems that real users depend on, read on.

Qualifications

  • 3–7+ years shipping ML/AI in production.
  • Strong Python; PyTorch or TensorFlow.
  • Hands-on with LLMs: prompting, fine-tuning, RAG, evals.
  • MLOps: CI/CD for models, monitoring, cost optimization.
  • Experience deploying on GCP or AWS (GCP preferred).
  • Comfort with vector DBs, embeddings, and retrieval systems.

Responsibilities

  • Deployment and scaling of LLM, VLM, and speech models in production (GCP).
  • Latency, cost, and reliability optimization across the stack.
  • RAG pipelines, prompting, and evaluation frameworks.
  • Infrastructure and tooling to accelerate experimentation and shipping.

Skills

Production ML in production
Python
PyTorch or TensorFlow
LLMs prompting/fine-tuning/RAG
MLOps CI/CD monitoring
GCP or AWS deployment
Vector DBs embeddings retrieval

Education

Bachelor’s or Master’s in CS/Engineering or related field

Tools

GCP
AWS
VectorDBs
PyTorch
TensorFlow

Job description

DeepHow is seeking an AI Engineer to own our AI pipeline end-to-end, taking over the production ML stack, harden it, and ship improvements fast. Day one, your focus is MLOps and production deployment making sure our models are fast, reliable, and cost-efficient at scale.

This is a build-and-ship role, not research. If you like turning prototypes into production systems that real users depend on, read on.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer: GenAI & LLM Production
MLOps Engineer: GenAI & LLM Production

InfoVision Inc. • Irving (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

InfoVision Inc. • Irving (TX)

On-site
USD 100,000 - 130,000
Senior AI/ML Engineer - LLMs, RL & Production ML
Senior AI/ML Engineer - LLMs, RL & Production ML

Socket.dev • Mountain View (CA)

On-site
USD 174,000 - 252,000
15% bonus target
Equity
Benefits
Production AI/ML Engineer - LLM & MLOps for Global DC Ops
Production AI/ML Engineer - LLM & MLOps for Global DC Ops

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+2
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
Senior ML Engineer: Production-Scale MLOps
Senior ML Engineer: Production-Scale MLOps

CoSourcing Partners - Enterprise-AI and IT Services Company • United States

On-site
USD 100,000 - 140,000
MLOps Engineer: Build & Deploy ML Pipelines
MLOps Engineer: Build & Deploy ML Pipelines

GCS Recruitment • United States

Remote
USD 90,000 - 130,000
Remote Project-Based ML Ops Engineer for Production AI
Remote Project-Based ML Ops Engineer for Production AI

Slalom • Milwaukee (WI)

Hybrid
401(k) with match
Health, dental, & vision coverage
Adoption and fertility assistance
+2
Remote MLOps Engineer: Scale Production ML Pipelines
Remote MLOps Engineer: Scale Production ML Pipelines

AgileEngine, LLC. • West Palm Beach (FL)

On-site
USD 120,000 - 180,000
Growth opportunities
Competitive compensation
Remote work
+3