Production AI Infra & MLOps Engineer

LaStellar Group

New York (NY)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LaStellar Group is seeking a Platform Engineer to own and operate our production AI and data platform infrastructure across Azure and GCP. You will manage cloud infra, CI/CD pipelines, container orchestration, and agent operations to keep systems reliable, scalable, and production-grade.

You will build observability tooling, support data platform integration, and implement auto-scaling for container workloads. The role demands 3–5 years in DevOps/MLOps with strong curiosity about AI systems.

Qualifications

  • 3–5 years in software engineering, DevOps, MLOps, or platform engineering with production ownership.
  • Hands‑on Docker and container orchestration experience in Azure and/or GCP.
  • Terraform across cloud providers — you've designed it, not just configured it.
  • CI/CD pipeline experience with Git-based release management.
  • Systems thinker — you troubleshoot end‑to‑end, not just at the surface.
  • Genuine curiosity about AI and agentic systems — excited to grow into deeper platform concepts.
  • Strong signals in policy-as-code, observability depth, and multi-cloud tooling are a plus.

Responsibilities

  • Operate and scale live agentic AI systems across Azure and GCP, ensuring availability, performance, and resilience.
  • Build and maintain observability tooling for agent execution—logging, tracing, alerting, and performance monitoring.
  • Support integration of agents with data platforms and MCP servers.
  • Implement auto-scaling strategies for containerized workloads across Azure Container Apps, Cloud Run, and GKE.
  • Contribute to evaluation frameworks and quality standards for AI agents in production.
  • Own Python-based services’ lifecycle—from containerization to deployment and runtime behavior.
  • Build shared tooling and internal packages to accelerate data science workflows.
  • Write and maintain Terraform across Azure and GCP for registries, identities, secrets, and networks.
  • Develop CI/CD pipelines and release workflows across data science and engineering repos, enforcing security and runbooks.

Skills

Docker
Container orchestration
CI/CD pipelines
Infrastructure as code
End-to-end troubleshooting
AI/agented systems curiosity

Tools

Azure
GCP
Azure Container Apps
GKE
GCP Cloud Run
Vertex AI
Managed Identities
VNets
Key Vault
Secret Manager
OpenTelemetry
LangChain
MLflow
Prefect
dbt
Snowflake

Job description

LaStellar Group is seeking a Platform Engineer to own and operate our production AI and data platform infrastructure across Azure and GCP. You will manage cloud infra, CI/CD pipelines, container orchestration, and agent operations to keep systems reliable, scalable, and production-grade.

You will build observability tooling, support data platform integration, and implement auto-scaling for container workloads. The role demands 3–5 years in DevOps/MLOps with strong curiosity about AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure / MLOps Engineer — NYC
AI Infrastructure / MLOps Engineer — NYC

LaStellar Group • New York (NY)

On-site
USD 140,000 - 190,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Agentic AI Engineer — Production ML & Automation
Agentic AI Engineer — Production ML & Automation

Stellar Consulting Solutions, LLC • United States

On-site
USD 150,000 - 210,000
Senior Platform Cloud Engineer AI & MLOps Leader
Senior Platform Cloud Engineer AI & MLOps Leader

ContractStaffingRecruiters.com • Stamford (CT)

Hybrid
USD 120,000 - 150,000
Senior AI Platform Engineer — Production-Grade MLOps
Senior AI Platform Engineer — Production-Grade MLOps

TetraScience, Inc. • Cambridge (MA)

Hybrid
USD 200,000 - 270,000
Employer-paid benefits
Unlimited PTO
401K
+3
GenAI Platform Engineer - MLOps on AWS & Azure
GenAI Platform Engineer - MLOps on AWS & Azure

Veriipro • Charlotte (NC)

On-site
USD 120,000 - 180,000
Remote MLOps Engineer for Production AI (Contract)
Remote MLOps Engineer for Production AI (Contract)

Slalom • City of Albany (NY)

Hybrid
401(k) with match
Health, dental, vision coverage
Fertility and adoption assistance
+2
Senior AI Platform Engineer: Scalable LLM Infra on GCP
Senior AI Platform Engineer: Scalable LLM Infra on GCP

Plarium • Spain (TX)

On-site
USD 140,000 - 200,000
Senior AI Infra Engineer: LLMOps, GPU, Kubernetes
Senior AI Infra Engineer: LLMOps, GPU, Kubernetes

eloquentai • San Francisco (CA)

On-site
USD 130,000 - 160,000