Senior MLOps Engineer

Jobot

Atlanta (GA)

On-site

USD 150,000 - 175,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote work 100%
Competitive salary + bonus + equity
Medical, dental, vision insurance
Life/ Disability insurance
11 paid holidays
401k with match

Job summary

Jobot is a US-based fully remote role focused on operating and scaling a government-grade AI platform on Azure. You will own cost discipline, infrastructure, and reliability across AKS, networking, observability, and routing.

You will collaborate with an AI Operations Engineer to ensure secure, auditable deployment and production readiness. The role emphasizes cost accountability, production security, and a FedRAMP-conscious posture, with strong emphasis on automation and governance across cloud

Qualifications

  • 5+ years in DevOps, platform engineering, or SRE in SaaS environments.
  • Deep Azure experience with AKS, networking, Entra, and monitoring.
  • Infrastructure as code is your default (Terraform/Bicep or similar); you automate before you document.
  • MLOps experience: deploying/operating ML systems in production.
  • Demonstrated cost optimization work in cloud environments.

Responsibilities

  • Deploy and operate the Azure platform: AKS, networking, identity, storage, environments from development through production.
  • Own infrastructure as code end to end; ensure reproducibility and visibility.
  • Operate the AI infrastructure layer: observability, telemetry, model gateway, per-workload routing.
  • Own cloud and AI cost: metering, budgets, unit economics, remediation.
  • Harden production access with least privilege, secrets management, audit trails.
  • Collaborate on deploy-and-release path: Octopus Deploy, environment promotion, rollout, rollback.
  • Build platform reliability: monitoring, alerting, incident response, capacity planning.
  • Provide platform primitives for microservices: service infrastructure, scaling, and boundaries.

Skills

DevOps
Platform Eng
SRE
Azure
IaC
Scripting
MLOps
Cost optimization
Compliance
Production access discipline

Tools

Terraform
Bicep
Langfuse
Octopus Deploy

Job description

Job details
Own where changes land and what it costs to run

This Jobot Job is hosted by: Charles Simmons

Salary: $150,000 - $175,000 per year

A bit about us:

Small, mid staged AI native SaaS startup helps state and local governments modernize paper-based processes into intelligent, AI-driven digital workflows. As we evolve into an AI-first platform, our development velocity, model iteration frequency, and cross-team complexity increase. A reliable, cost-disciplined platform is essential to scale safely and predictably.

Why join us?
  • 100% work from home (US based only)
  • Own the platform that build, validation, and release loops run on and deploy to: infrastructure, environments, Kubernetes, Networking, observability, and the AI serving and routing layer.
  • Competitive base, bonus, and equity options
  • Medical, dental, and vision insurance plans, with significant employer contributions for employees AND dependents (contributions based on base-level plan; buyup plans available at additional costs)
  • Company-sponsored life, short-term, and long-term disability insurance
  • 11 Paid holidays
  • Flexible time off
  • 401k plan with 4% employer match
  • Monthly stipend for home office expenses
  • Monthly wellness stipend
Job Details

Everything runs on Azure, and the platform is getting more interesting: an AI product suite heading toward general availability, self-hosted AI observability and telemetry inside a FedRAMP-conscious boundary, autonomous agents participating in delivery, and a microservices decomposition in flight. You will deploy, operate, and scale that platform, and you will own its cost discipline.

This is a production seat with production access, and we treat that as an engineering responsibility, not a badge: least privilege, audit trails, and environment integrity are part of the job, because our customers are governments.

You will work alongside our AI Operations Engineer, who owns the agentic delivery system (the loops that build, validate, and release code). You own the platform those loops run on and deploy to: infrastructure, environments, Kubernetes, networking, observability, and the AI serving and routing layer. The boundary is simple: they own how changes move; you own where changes land and what it costs to run.

Responsibilities
  • Deploy and operate our Azure platform: AKS, networking, identity, storage, and environments from development through production
  • Own infrastructure as code end to end: environments are reproducible, drift is detected, and nothing reaches an environment without platform visibility
  • Operate the AI infrastructure layer: self-hosted observability and evaluation tooling (Langfuse), product telemetry, model gateway and per-workload routing, and compliant GovCloud inference paths
  • Own cloud and AI cost: metering, budgets, unit economics, MACC drawdown strategy, and active remediation; cost is an engineering metric here, not a finance afterthought
  • Harden production access and controls: least privilege, secrets management, audit evidence, and a FedRAMP-conscious security posture
  • Partner with AI Operations on the deploy-and-release path: Octopus Deploy, environment promotion, progressive rollout, and rollback
  • Build platform reliability: monitoring, alerting, incident response, and capacity planning
  • Give the microservices decomposition the platform primitives it needs: service infrastructure, scaling patterns, and clean environment boundaries
Qualifications
  • 5+ years in DevOps, platform engineering, or site reliability engineering in SaaS environments
  • Deep Azure experience: AKS, networking, identity (Entra), and monitoring; you have run production Kubernetes
  • Infrastructure as code as your default (Terraform, Bicep, or similar), plus strong scripting; you automate before you document
  • MLOps experience: deploying and operating LLM or ML systems in production, including model gateways, inference infrastructure, or AI observability stacks
  • Demonstrated cost work: you can point to cloud spend you found, explained, and reduced
  • Experience in compliance-heavy environments (FedRAMP, StateRAMP, SOC 2, or similar) is a strong plus
  • Comfortable holding production access, with the discipline that implies
Key Competencies
  • Treats environment integrity as sacred: no invisible changes, no snowflake servers, no heroics that cannot be audited
  • Cost literacy: reads a cloud bill the way an engineer reads a stack trace
  • Automates first: your instinct is a pipeline or a policy, not a runbook step
  • thinks in the open: surfaces risk early and documents what you build
  • Calm in production incidents; rigorous in the postmortem
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure / MLOps Engineer — NYC
AI Infrastructure / MLOps Engineer — NYC

LaStellar Group • New York (NY)

On-site
USD 140,000 - 180,000
MLOps / AIOps / LLMOps / AgentOps Engineer
MLOps / AIOps / LLMOps / AgentOps Engineer

FinOps Weekly • Northern (KY)

Hybrid
USD 120,000 - 160,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
MLOPS AI ENGINEER II
MLOPS AI ENGINEER II

VDart Inc • Coppell (TX)

Hybrid
USD 96,000 - 152,000
Senior MLOps Engineer
Senior MLOps Engineer

AppRecode, Inc. • Town of Middletown (NY)

On-site
USD 120,000 - 160,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

Codinix Consulting Services • California (MO)

On-site
USD 120,000 - 150,000
MLOps Engineer
MLOps Engineer

Atomic Machines • Emeryville (CA)

On-site
USD 200,000 - 250,000
Senior MLOps Engineer
Senior MLOps Engineer

Harnham • New York (NY)

On-site
USD 140,000 - 190,000
Base salary + bonus
Comprehensive benefits package