AI Platform Operations Manager

STACK Infrastructure US

United States

On-site

USD 128,000 - 146,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Life insurance
Paid time off
Paid leave programs

Job summary

STACK Infrastructure’s DevOps Engineer, AI Platform, drives automation and reliable delivery pipelines for enterprise AI on Azure. You’ll own infrastructure‑as‑code, CI/CD, containerization, observability, and release automation enabling AI teams to ship RAG pipelines and integration services efficiently.

You will work at the intersection of cloud infrastructure, platform engineering, and MLOps, turning architecture into automated, governable, and cost‑efficient environments used across the

Qualifications

  • Bachelor's degree in Computer Science, IT, Engineering, or related field, or equivalent.
  • 5+ years of hands-on DevOps, SRE, or platform engineering with delivery track record.
  • Strong Infrastructure as Code experience with Terraform, Bicep, or ARM.
  • Proven CI/CD experience in Azure DevOps or GitHub Actions; containerization knowledge.

Responsibilities

  • Automate infrastructure via IaC; manage Azure resources and AI platform components.
  • Design, build, and operate CI/CD pipelines for code, infra, and AI deployments.
  • Maintain environment parity across dev/test/prod; enforce governance and tagging standards.
  • Operate AKS and Azure Container Apps; manage container images and registries.
  • Support release workflows for agentic AI services and model endpoints.

Skills

DevOps
CI/CD
IaC
Azure
Python
Docker
AKS
Observability

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
Bicep
ARM
Azure DevOps
GitHub Actions
AKS
Azure Container Apps
Azure Monitor
Prometheus
Grafana

Job description

STACK INFRASTRUCTURE

provides digital infrastructure to scale the world’s most innovative companies. We are an award-winning industry leader in building, owning, and operating highly efficient, cost-effective wholesale, colocation, and cloud data centers. Each of our national facilities meets or exceeds the highest industry standards in all operational categories of availability, security, connectivity, and physical resilience. STACK offers the scale and geographic reach that rapidly growing hyperscale and enterprise companies need. The world runs on data. Data runs on STACK.

The Position

DevOps Engineer, AI Platform is responsible for automating, deploying, and operating the infrastructure and delivery pipelines that support STACK’s enterprise AI platform on Azure. This is a hands‑on engineering role focused on build and run — not oversight. Reporting to Head of AI, Enterprise AI & Data Strategy org this individual owns the infrastructure‑as‑code, CI/CD, containerization, observability, and release automation that allow AI engineers and enterprise application teams to ship agentic AI solutions, RAG pipelines, and integration services reliably and repeatably. The role sits at the intersection of cloud infrastructure, platform engineering, and MLOps — turning platform architecture into automated, governed, observable, and cost‑efficient environments that teams across the organization build on.

Key Responsibilities
  • Infrastructure Automation & Infrastructure as Code
    • Build, maintain, and version infrastructure-as-code modules for Azure environments using Terraform, Bicep, or ARM, including compute, networking, storage, identity, and AI platform resources.
    • Automate provisioning of AI platform components — Azure AI Foundry, Azure OpenAI Service, Azure AI Search, Cosmos DB, ADLS Gen2, and Databricks — as reusable, parameterized deployment patterns.
    • Maintain environment parity across development, test, and production, including configuration management, drift detection, and remediation.
    • Implement and enforce tagging, naming, and resource organization standards that support governance, chargeback, and lifecycle management.
    • Automate routine platform operations — patching, certificate rotation, key and secret rotation, backup validation, and disaster recovery testing.
  • CI/CD & Release Engineering
    • Design, build, and operate CI/CD pipelines in Azure DevOps or GitHub Actions for application code, infrastructure code, container images, and AI/agent deployments.
    • Implement automated build, test, security scanning, artifact management, and promotion gates across environments.
    • Establish branching strategies, code review standards, and release management practices in partnership with AI engineering and enterprise application teams.
    • Build deployment automation for agentic AI services, MCP (Model Context Protocol) servers, and integration workloads running on Azure Container Apps and Azure Kubernetes Service (AKS).
    • Support model and prompt release workflows — versioning, staged rollout, evaluation gates, and rollback procedures for LLM-based applications.
  • Container Platform & AI Workload Operations
    • Operate and tune AKS and Azure Container Apps, including cluster upgrades, node pool sizing, autoscaling, ingress, networking, and workload isolation.
    • Build and maintain container images, base image standards, and registry governance in Azure Container Registry.
    • Manage compute scheduling and scaling for AI workloads, including GPU-backed and inference-heavy workloads where required.
    • Implement resiliency patterns — health probes, retries, throttling, quota management, and failover — for AI endpoints and integration services.
  • Observability, Reliability & Incident Response
    • Instrument platform and AI services with logging, metrics, tracing, and alerting using Azure Monitor, Log Analytics, Application Insights, and equivalent open-source tooling.
    • Build dashboards and service-level indicators covering platform availability, latency, throughput, error rates, token consumption, and model endpoint performance.
    • Participate in on‑call rotation, lead incident triage and resolution for platform issues, and drive root cause analysis and corrective actions.
    • Develop and maintain runbooks, operational documentation, and automated remediation for recurring issues.
  • Security, Governance & Cost Optimization
    • Implement DevSecOps practices — secrets management in Azure Key Vault, managed identity usage, least‑privilege access, dependency and container vulnerability scanning, and policy-as-code.
    • Partner with Information Security to ensure pipelines and environments meet enterprise security, data residency, and compliance requirements.
    • Support Azure FinOps practices through cost visibility, rightsizing, reserved capacity, and automated controls on non‑production and idle resources.
    • Maintain audit trails and change records for infrastructure and release activity.
  • Delivery & Cross-Functional Collaboration
    • Work directly with AI engineers, data engineers, and enterprise application teams to remove deployment friction and improve time‑to‑production for AI solutions.
    • Translate platform architecture and standards into automated, self‑service capabilities that teams can consume without deep infrastructure knowledge.
    • Contribute to platform engineering standards, reference implementations, and internal documentation.
    • Provide technical escalation support for build, deployment, and environment issues.
The Details

Location: Denver, CO Travel: <10% Benefits: Healthcare, Dental Care, Vision Insurance, Life Insurance, Paid Time Off, and Paid Leave Programs Must be eligible to work in the United States Must pass comprehensive background and drug screening

MUST HAVE QUALIFICATIONS
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field, or equivalent practical experience.
  • 5+ years of hands‑on experience in DevOps, site reliability engineering, or platform engineering roles with a strong delivery track record.
  • Strong proficiency in Infrastructure as Code — Terraform, Bicep, or ARM — including module design, state management, and reusable patterns.
  • Proven experience building and operating CI/CD pipelines in Azure DevOps or GitHub Actions.
  • Hands‑on experience with containerization and orchestration — Docker, Azure Kubernetes Service (AKS), and Azure Container Apps or equivalent.
  • Solid working knowledge of Azure core services — compute, networking (VNet, NSG, Private Endpoints), storage, identity (Entra ID), and Key Vault.
  • Strong scripting and automation skills in Python, PowerShell, or Bash.
  • Experience with monitoring and observability tooling — Azure Monitor, Log Analytics, Application Insights, Prometheus, or Grafana.
  • Working knowledge of Git‑based workflows, code review practices, and artifact/registry management.
  • Demonstrated ability to troubleshoot production issues across infrastructure, network, and application layers.
PREFERRED QUALIFICATIONS
  • Microsoft Certified: DevOps Engineer Expert (AZ‑400), Azure Administrator (AZ‑104), or Certified Kubernetes Administrator (CKA).
  • Experience deploying and operating AI/ML workloads — model endpoints, RAG pipelines, vector databases, or agentic services in production.
  • Familiarity with MLOps tooling and practices — Azure Machine Learning, MLflow, Databricks, or equivalent model lifecycle platforms.
  • Experience deploying MCP (Model Context Protocol) servers or similar integration services connecting AI agents to enterprise systems.
  • Exposure to agentic AI frameworks such as Semantic Kernel, LangGraph, or AutoGen from a deployment and operations perspective.
  • Experience with GPU compute provisioning, quota management, and inference cost optimization.
  • Knowledge of FinOps frameworks and Azure cost optimization practices.
  • Experience integrating with enterprise systems such as Microsoft 365, Freshworks ITSM, Workday, NetSuite, or Procore.
  • Experience in data center, hyperscale, or infrastructure-intensive industry environments.

Compensation Range: $128,260.00 - $146,017.59

This Might Be Right for You If

You are a strong communicator, you are persuasive and clear, blending analytics with experience in decision‑making. You do not get flustered easily. You can juggle multiple priorities while balancing urgent requests with shifting timelines and deliverables. You are a team builder. You take the time to understand and develop the strengths of your resources while formulating long‑term plans for the growth and success of the team. You are naturally curious and driven toward continual improvement. While you celebrate your successes, you take time to review and analyze campaigns for future learning.

Why Stack?

We offer a competitive compensation package with strong benefits, including medical, dental, and vision insurance, a 401K program, flexible spending accounts – even a cell phone subsidy. We foster a culture of appreciation, including peer‑to‑peer recognition and rewards programs. Fun is part of our DNA, with events, game nights, happy hours, and barbecues. We’re growing – this is a great time to join and make an impact!

STACK is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity and expression, age, national origin, mental or physical disability, genetic information, veteran status, or any other status protected by federal, state, or local law.

Note to external agencies: We are not accepting any blind submissions or resumes/cvs from recruitment agencies. Any candidates sent to STACK Infrastructure, Inc. will not be accepted or considered as a submission without a signed agreement in place. Fees will not be paid in the event a candidate submitted by a recruiter without an agreement in place is hired; such resumes will be deemed the sole property of STACK Infrastructure, Inc.

With a culture rooted in inclusivity and growth, STACK invests in your future. Here, diverse perspectives are celebrated, and career development is central. Join us to be part of a team where your skills and ideas shape the future of digital infrastructure, building not just a job, but a lasting career with impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Operations Manager
AI Platform Operations Manager

STACK Infrastructure • Denver (CO)

On-site
USD 128,000 - 146,000
Healthcare
Dental Care
Vision Insurance
+3
AI Application Engineer
AI Application Engineer

STACK Infrastructure US • United States

Hybrid
USD 128,000 - 146,000
Healthcare
Dental Care
Vision Insurance
+3
AI Application Engineer
AI Application Engineer

STACK Infrastructure APAC • Denver (CO)

Hybrid
USD 128,000 - 146,000
Medical, dental, and vision insurance
401K program
Flexible spending accounts
+3
AI Application Engineer
AI Application Engineer

STACK Infrastructure • Denver (CO)

Hybrid
USD 128,000 - 146,000
Medical, dental, and vision insurance
401K program
Cell phone subsidy
Critical Operations Integration Lead
Critical Operations Integration Lead

STACK Infrastructure US • Plano (TX)

On-site
USD 140,000 - 190,000
Medical, dental, and vision insurance
401K with company match
Flexible spending accounts
+2
Senior DevOps Engineer
Senior DevOps Engineer

StackAI • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Remote work available
Office near Salesforce Park
AI Application Engineer
AI Application Engineer

STACK Infrastructure US • Denver (CO)

Hybrid
USD 128,000 - 146,000
Healthcare
Dental Insurance
Vision Insurance
+4
Senior Manager, Finance Systems & Operations
Senior Manager, Finance Systems & Operations

STACK Infrastructure US • United States

On-site
USD 141,000 - 163,000
Medical, dental, vision insurance
401K program
Cell phone subsidy
+1
Senior DevOps Engineer
Senior DevOps Engineer

StackAI • New York (NY)

Hybrid
USD 130,000 - 210,000
Hybrid work model
Remote days
Office in SF
Critical Operations Integration Lead
Critical Operations Integration Lead

STACK Infrastructure US • Sterling (VA)

On-site
USD 120,000 - 180,000
Medical, dental, and vision insurance
401K program
Annual bonus/merit or commission