Senior MLOps Engineer

pointwild

United States

Remote

USD 140,000 - 220,000

Full time

12 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Opportunity to influence cybersecurity
Enterprise-scale production AI

Job summary

pointwild seeks a senior MLOps engineer to build production ML infrastructure on Google Cloud, turning notebooks into resilient, auto-scaling services. You will work with researchers, data engineers and backend teams to architect pipelines and tooling.

Applicants should have 5+ years deploying ML workloads in the cloud, strong GCP expertise, and hands-on experience with Docker/Kubernetes, Airflow, Vertex AI Pipelines and CI/CD.

Qualifications

  • Five+ years of hands-on production ML in cloud environments.
  • GCP expertise: Vertex AI, Cloud Storage, GKE, Cloud Run, IAM/VPC.
  • Containerisation with Docker and Kubernetes; specialised serving tools (Triton, vLLM, MLflow).
  • Workflow orchestration with Airflow or Vertex AI Pipelines; CI/CD with GitHub Actions or ArgoCD.
  • Terraform for cloud resource management.
  • Python and SQL for scripting, automation and data handling.
  • Logging/telemetry via Grafana, Prometheus, GCP Monitoring or ML observability tools.

Responsibilities

  • Architect and manage scalable GCP-based ML infra (Vertex AI, GKE, Cloud Storage, Cloud Run, GPU/TPU).
  • Own end-to-end deployment lifecycle; build high-throughput, low-latency inference services.
  • Create automated, reproducible pipelines for training, testing, evaluation and deployment (Airflow, Vertex Pipelines, GitHub Actions).
  • Monitor system health and ML metrics; implement automated retraining triggers.
  • Provide scalable training environments and standardised deployment templates for researchers.
  • Collaborate with data engineers to integrate feature stores, versioning and streaming/batch workflows.
  • Lead migration of prototypes/notebooks into resilient, secure microservices.

Skills

GCP Vertex AI
Docker
Kubernetes
CI/CD
Terraform
Python
SQL
Grafana
Prometheus
Airflow
Vertex AI Pipelines
MLflow
Triton
vLLM

Tools

GKE
Cloud Run
GitHub Actions
ArgoCD
Cloud Storage
IAM/VPC

Job description

Role overview

This senior MLOps role focuses on building the production backbone for AI systems on Google Cloud. The engineer will architect infrastructure, pipelines and tooling that take models from notebooks and proofs of concept into resilient, observable, auto-scaling services, working closely with AI researchers, data engineers and backend teams across a cybersecurity organisation.

Responsibilities
  • Architect and manage scalable GCP-based ML infrastructure using Vertex AI, Google Kubernetes Engine, Cloud Storage, Cloud Run, and GPU or TPU compute
  • Own the end-to-end deployment lifecycle for machine learning models, building high-throughput, low-latency inference services with containerisation and specialised serving frameworks
  • Build automated, reproducible pipelines for model training, testing, evaluation and deployment using tools such as Airflow, Vertex AI Pipelines and GitHub Actions
  • Implement monitoring for system health and ML-specific metrics including feature drift, prediction accuracy and data distribution shifts, enabling automated retraining triggers
  • Provide AI and research engineers with scalable training environments, optimised runtime infrastructure and standardised deployment templates
  • Collaborate with data engineers to integrate pipelines with feature stores, dataset versioning and stream or batch processing workflows
  • Lead the technical transition of prototypes and notebooks into resilient, secure, auto-scaling microservices
Requirements
  • At least five years of hands‑on experience designing, deploying and maintaining production ML workloads in cloud environments
  • Deep practical experience with GCP, including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM and VPC configuration
  • Expertise with containerisation using Docker and Kubernetes or GKE, plus specialised serving tools such as Triton, vLLM or MLflow
  • Proven track record with workflow orchestrators like Airflow or Vertex AI Pipelines and modern CI/CD tools such as GitHub Actions or ArgoCD
  • Solid experience managing cloud resources with Terraform
  • Proficiency in Python and SQL for scripting, automation, API development and data manipulation
  • Hands‑on experience with logging, telemetry and drift detection using Grafana, Prometheus, GCP Cloud Monitoring or specialised ML observability frameworks
Nice to have
  • Experience running large-scale LLM or deep learning inference and training workloads
  • GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certification
  • Familiarity with feature stores such as Feast or Vertex AI Feature Store
Benefits and work setup
  • Opportunity to work across cloud, data and AI disciplines within a cybersecurity-focused product organisation
  • Exposure to production AI workloads at enterprise scale with autonomy to shape MLOps practices
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior MLOps Engineer: Scalable AI Infra on GCP
Senior MLOps Engineer: Scalable AI Infra on GCP

pointwild • United States

Remote
USD 140,000 - 220,000
Opportunity to influence cybersecurity
Enterprise-scale production AI
MLOps Engineer (Remote)
MLOps Engineer (Remote)

Cedent • United States

Remote
USD 55,104 - 82,656
Health Insurance
Dental Insurance
Vision Insurance
AI / ML Ops Engineer
AI / ML Ops Engineer

Zoho • United States

Remote
USD 140,000 - 210,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
MLOps Engineer
MLOps Engineer

Evlo AI • Seattle (WA)

On-site
USD 130,000 - 190,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
Senior MLOps Engineer
Senior MLOps Engineer

AppRecode, Inc. • Town of Middletown (NY)

On-site
USD 120,000 - 160,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Confidential
Architect - Platform Engineering - USA
Architect - Platform Engineering - USA

Quantiphi, Inc. • Chicago (IL)

On-site
USD 130,000 - 180,000
Opportunity to work with Fortune 500 companies
Exposure to cutting-edge AI technologies
Dynamic team environment