AI Platform Operations Manager

STACK Infrastructure

Denver (CO)

On-site

USD 128,000 - 146,000

Full time

6 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare
Dental Care
Vision Insurance
Life Insurance
Paid Time Off
Paid Leave Programs

Job summary

STACK INFRASTRUCTURE is hiring a DevOps Engineer, AI Platform in Denver to automate, deploy, and run the Azure-based AI platform pipelines. You will own infrastructure as code, CI/CD, containerization, observability, and release automation enabling AI teams to ship services reliably.

The role sits at the intersection of cloud engineering and MLOps, requiring strong collaboration with AI engineers and enterprise teams, with on‑site work in Denver, CO.

Qualifications

  • Bachelor’s degree in Computer Science, IT, Engineering, or related field.
  • 5+ years in DevOps/platform engineering with a proven delivery track record.
  • Strong Infrastructure as Code experience with Terraform, Bicep, or ARM.

Responsibilities

  • Build, maintain, and version IaC modules for Azure environments (Terraform/Bicep/ARM).
  • Automate provisioning of AI platform components (Azure AI Foundry, OpenAI Service, Databricks).
  • Maintain environment parity across Dev, Test, Production with config mgmt and drift detection.
  • Automate patching, cert rotation, key/secret rotation, backup validation, DR testing.
  • Design, build, and operate CI/CD pipelines in Azure DevOps or GitHub Actions for code, infra, containers, AI deployments.
  • Support model and prompt release workflows including versioning and rollout gates.
  • Operate AKS and Azure Container Apps, manage scaling, networking, and security boundaries.
  • Instrument platforms with logging/metrics/tracing; develop runbooks and automation for issues.

Skills

Terraform/IaC
Azure AKS/Container Apps
CI/CD pipelines
Containerization
Scripting (Python/PowerShell/Bash)
Observability tooling
Git workflows
Production troubleshooting
DevSecOps basics

Education

Bachelor's degree or equivalent

Tools

Terraform
Bicep
ARM
Azure DevOps
GitHub Actions
AKS
Azure Container Apps
Azure Monitor/Log Analytics

Job description

The Company

STACK INFRASTRUCTURE (STACK) provides digital infrastructure to scale the world’s most innovative companies. We are an award-winning industry leader in building, owning, and operating highly efficient, cost-effective wholesale, colocation, and cloud data centers. Each of our national facilities meets or exceeds the highest industry standards in all operational categories of availability, security, connectivity, and physical resilience.

The Company

STACK INFRASTRUCTURE (STACK) provides digital infrastructure to scale the world’s most innovative companies. We are an award-winning industry leader in building, owning, and operating highly efficient, cost-effective wholesale, colocation, and cloud data centers. Each of our national facilities meets or exceeds the highest industry standards in all operational categories of availability, security, connectivity, and physical resilience.

STACK offers the scale and geographic reach that rapidly growing hyperscale and enterprise companies need. The world runs on data. Data runs on STACK.

The Position

The DevOps Engineer, AI Platform is responsible for automating, deploying, and operating the infrastructure and delivery pipelines that support STACK’s enterprise AI platform on Azure. This is a hands‑on engineering role focused on build and run — not oversight. Reporting to Head of AI, Enterprise AI & Data Strategy org this individual owns the infrastructure-as-code, CI/CD, containerization, observability, and release automation that allow AI engineers and enterprise application teams to ship agentic AI solutions, RAG pipelines, and integration services reliably and repeatably. The role sits at the intersection of cloud infrastructure, platform engineering, and MLOps — turning platform architecture into automated, governed, observable, and cost‑efficient environments that teams across the organization build on.

Key Responsibilities
Infrastructure Automation & Infrastructure as Code
  • Build, maintain, and version infrastructure-as-code modules for Azure environments using Terraform, Bicep, or ARM, including compute, networking, storage, identity, and AI platform resources.
  • Automate provisioning of AI platform components — Azure AI Foundry, Azure OpenAI Service, Azure AI Search, Cosmos DB, ADLS Gen2, and Databricks — as reusable, parameterized deployment patterns.
  • Maintain environment parity across development, test, and production, including configuration management, drift detection, and remediation.
  • Implement and enforce tagging, naming, and resource organization standards that support governance, chargeback, and lifecycle management.
  • Automate routine platform operations — patching, certificate rotation, key and secret rotation, backup validation, and disaster recovery testing.
CI/CD & Release Engineering
  • Design, build, and operate CI/CD pipelines in Azure DevOps or GitHub Actions for application code, infrastructure code, container images, and AI/agent deployments.
  • Implement automated build, test, security scanning, artifact management, and promotion gates across environments.
  • Establish branching strategies, code review standards, and release management practices in partnership with AI engineering and enterprise application teams.
  • Build deployment automation for agentic AI services, MCP (Model Context Protocol) servers, and integration workloads running on Azure Container Apps and Azure Kubernetes Service (AKS).
  • Support model and prompt release workflows — versioning, staged rollout, evaluation gates, and rollback procedures for LLM-based applications.
Container Platform & AI Workload Operations
  • Operate and tune AKS and Azure Container Apps, including cluster upgrades, node pool sizing, autoscaling, ingress, networking, and workload isolation.
  • Build and maintain container images, base image standards, and registry governance in Azure Container Registry.
  • Manage compute scheduling and scaling for AI workloads, including GPU-backed and inference-heavy workloads where required.
  • Implement resiliency patterns — health probes, retries, throttling, quota management, and failover — for AI endpoints and integration services.
Observability, Reliability & Incident Response
  • Instrument platform and AI services with logging, metrics, tracing, and alerting using Azure Monitor, Log Analytics, Application Insights, and equivalent open-source tooling.
  • Build dashboards and service-level indicators covering platform availability, latency, throughput, error rates, token consumption, and model endpoint performance.
  • Participate in on‑call rotation, lead incident triage and resolution for platform issues, and drive root cause analysis and corrective actions.
  • Develop and maintain runbooks, operational documentation, and automated remediation for recurring issues.
Security, Governance & Cost Optimization
  • Implement DevSecOps practices — secrets management in Azure Key Vault, managed identity usage, least‑privilege access, dependency and container vulnerability scanning, and policy-as-code.
  • Partner with Information Security to ensure pipelines and environments meet enterprise security, data residency, and compliance requirements.
  • Support Azure FinOps practices through cost visibility, rightsizing, reserved capacity, and automated controls on non-production and idle resources.
  • Maintain audit trails and change records for infrastructure and release activity.
Delivery & Cross-Functional Collaboration
  • Work directly with AI engineers, data engineers, and enterprise application teams to remove deployment friction and improve time‑to‑production for AI solutions.
  • Translate platform architecture and standards into automated, self‑service capabilities that teams can consume without deep infrastructure knowledge.
  • Contribute to platform engineering standards, reference implementations, and internal documentation.
  • Provide technical escalation support for build, deployment, and environment issues.
The Details
  • Location: Denver, CO
  • Travel:
  • Benefits: Healthcare, Dental Care, Vision Insurance, Life Insurance, Paid Time Off, and Paid Leave Programs
  • Must be eligible to work in the United States
  • Must pass comprehensive background and drug screening
Must-have Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field, or equivalent practical experience.
  • 5+ years of hands‑on experience in DevOps, site reliability engineering, or platform engineering roles with a strong delivery track record.
  • Strong proficiency in Infrastructure as Code — Terraform, Bicep, or ARM — including module design, state management, and reusable patterns.
  • Proven experience building and operating CI/CD pipelines in Azure DevOps or GitHub Actions.
  • Hands‑on experience with containerization and orchestration — Docker, Azure Kubernetes Service (AKS), and Azure Container Apps or equivalent.
  • Solid working knowledge of Azure core services — compute, networking (VNet, NSG, Private Endpoints), storage, identity (Entra ID), and Key Vault.
  • Strong scripting and automation skills in Python, PowerShell, or Bash.
  • Experience with monitoring and observability tooling — Azure Monitor, Log Analytics, Application Insights, Prometheus, or Grafana.
  • Working knowledge of Git-based workflows, code review practices, and artifact/registry management.
  • Demonstrated ability to troubleshoot production issues across infrastructure, network, and application layers.
Preferred Qualifications
  • Microsoft Certified: DevOps Engineer Expert (AZ-400), Azure Administrator (AZ-104), or Certified Kubernetes Administrator (CKA).
  • Experience deploying and operating AI/ML workloads — model endpoints, RAG pipelines, vector databases, or agentic services in production.
  • Familiarity with MLOps tooling and practices — Azure Machine Learning, MLflow, Databricks, or equivalent model lifecycle platforms.
  • Experience deploying MCP (Model Context Protocol) servers or similar integration services connecting AI agents to enterprise systems.
  • Exposure to agentic AI frameworks such as Semantic Kernel, LangGraph, or AutoGen from a deployment and operations perspective.
  • Experience with GPU compute provisioning, quota management, and inference cost optimization.
  • Knowledge of FinOps frameworks and Azure cost optimization practices.
  • Experience integrating with enterprise systems such as Microsoft 365, Freshworks ITSM, Workday, NetSuite, or Procore.
  • Experience in data center, hyperscale, or infrastructure-intensive industry environments.
Compensation Range

$128,260.00 - $146,017.59

This Might Be Right For You If
  • You are a strong communicator, you are persuasive and clear, blending analytics with experience in decision‑making.
  • You do not get flustered easily. You can juggle multiple priorities while balancing urgent requests with shifting timelines and deliverables.
  • You are a team builder. You take the time to understand and develop the strengths of your resources while formulating long‑term plans for the growth and success of the team.
  • You are naturally curious and driven toward continual improvement. While you celebrate your successes, you take time to review and analyze campaigns for future learning.
WHY STACK?
  • We offer a competitive compensation package with strong benefits, including medical, dental, and vision insurance, a 401K program, flexible spending accounts – even a cell phone subsidy.
  • We foster a culture of appreciation, including peer‑to‑peer recognition and rewards programs.
  • Fun is part of our DNA, with events, game nights, happy hours, and barbecues.
  • We’re growing – this is a great time to join and make an impact!

STACK is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity and expression, age, national origin, mental or physical disability, genetic information, veteran status, or any other status protected by federal, state, or local law

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

StackAI • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Remote work available
Office near Salesforce Park
Senior DevOps Engineer
Senior DevOps Engineer

StackAI • New York (NY)

Hybrid
USD 130,000 - 210,000
Hybrid work model
Remote days
Office in SF
AI Infrastructure Analyst
AI Infrastructure Analyst

Heaven Hill Brands • Louisville (KY)

On-site
USD 86,000 - 120,000
Paid Vacation
11 Paid Holidays
Health, Dental & Vision eligibility
+4
Software Engineer II - Platform
Software Engineer II - Platform

Stacklok • New York (NY)

On-site
USD 158,000 - 193,000
Equity
Comprehensive benefits package
Hybrid/remote work environment
Sr. Software Engineer - Platform
Sr. Software Engineer - Platform

Stacklok • Atlanta (GA)

On-site
USD 189,000 - 231,000
Equity
Medical Insurance
Dental & Vision
+4
AI Platform DevOps Engineer — Infra Automation & CI/CD
AI Platform DevOps Engineer — Infra Automation & CI/CD

STACK Infrastructure • Denver (CO)

On-site
USD 128,000 - 146,000
Healthcare
Dental Care
Vision Insurance
+3
AI Platform Engineer
AI Platform Engineer

Worky • Conshohocken

Hybrid
USD 140,000 - 150,000
Medical, Dental, Vision
401(k) plan
Paid Time Off
+1
Staff Platform Engineer
Staff Platform Engineer

United States Digital Space LLC • San Francisco (CA)

On-site
USD 198,000 - 331,000
Equity options
Medical, dental, and vision benefits
Unlimited PTO
Sr Advanced AI Platform Engineer
Sr Advanced AI Platform Engineer

Honeywell • Atlanta (GA)

Hybrid
USD 120,000 - 150,000
Performance-driven salary
Employer-subsidized Medical, Dental, Vision, and Life Insurance
401(k) match
+3
Find the work that fits.
Find the work that fits.

United States Digital Space LLC • United States

Hybrid
USD 110,000 - 150,000
Competitive Salary
Mediclaim: 10 Lacs coverage for you &/
Hybrid Work: Flexible PTO & remote
+3