Senior AI Operations Engineer (Senior Software Engineer)

Vizient, Inc.

Centennial (CO)

On-site

USD 102,000 - 179,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Comprehensive benefits plan
Incentive eligible

Job summary

Vizient, Inc. is seeking a senior backend engineer to design and build scalable cloud-based services powering AI/ML applications.

You will work with AI/ML engineers, data scientists, and cross-functional teams to productionize LLM workflows, RAG pipelines, and agentic AI solutions. You will leverage Pulumi, Terraform, Docker, and Kubernetes to implement infra-as-code, CI/CD pipelines, and secure, observable systems, while defining reusable patterns and engineering standards to accelerate

Qualifications

  • 5 or more years of relevant experience required.
  • Strong Python development experience with production backend services, APIs, distributed systems, and cloud-based applications.
  • Hands-on experience with infrastructure-as-code tools (Pulumi/Terraform) and cloud platforms (AWS/Azure/GCP).
  • Experience with Docker, CI/CD pipelines, container orchestration (Kubernetes), serverless architectures, and cloud deployment practices.
  • Knowledge of observability tools and production support practices.
  • Familiarity with ML/AI systems, model serving, LLM applications, inference pipelines, and RAG workflows.
  • Experience with Databricks, Azure AI Foundry, or similar AI/ML platform technologies.
  • Strong analytical, troubleshooting, problem-solving, and cross-functional collaboration skills.

Responsibilities

  • Design, develop, and maintain backend services, APIs, and platform components for AI/ML applications and distributed systems.
  • Build scalable cloud infrastructure using Pulumi and modern IaC practices; develop CI/CD pipelines and containerized cloud environments.
  • Support production deployment of ML models, LLM applications, RAG pipelines, and agentic AI systems with engineers to productionize model serving and data workflows.
  • Enhance observability through monitoring, logging, tracing, alerting, and incident management to improve reliability and performance.
  • Implement engineering standards for testing, code quality, security, maintainability, scalability, latency, and cost efficiency.
  • Define reusable platform patterns, developer tooling, and workflows to improve productivity and consistency.
  • Evaluate emerging AI engineering trends and tools to drive continuous improvement and innovation.
  • Partner with product, security, data, and platform teams to deliver production-ready AI solutions and contribute to platform strategy.
  • Troubleshoot complex production issues, perform root-cause analysis, and drive remediation to improve stability.
  • Mentor engineers through code reviews, knowledge sharing, and engineering best practices.

Skills

Python development
Distributed systems
Analytical thinking
Collaboration

Education

Relevant degree

Tools

Pulumi
Terraform
Docker
Kubernetes
CI/CD
Databricks
Azure AI Foundry

Job description

Summary:


In this role, you will design and build scalable backend systems, cloud infrastructure, and platform capabilities that power AI/ML products and applications. You will partner closely with AI/ML engineers, data scientists, and cross-functional teams to productionize LLM applications, RAG pipelines, and agentic AI workflows. You will leverage modern cloud technologies, infrastructure-as-code practices, and AI-assisted development tools to deliver reliable, secure, and maintainable solutions that accelerate innovation across the AI/ML team while helping establish engineering standards, reusable platform patterns, and operational best practices.

Responsibilities:
  • Design, develop, and maintain backend services, APIs, and platform components that support AI/ML applications and distributed systems.
  • Build scalable cloud infrastructure using Pulumi and modern infrastructure-as-code practices while developing CI/CD pipelines, deployment workflows, and containerized cloud environments.
  • Support production deployment of ML models, LLM applications, RAG pipelines, and agentic AI systems while collaborating with AI/ML engineers to productionize model serving, inference pipelines, and data workflows.
  • Enhance observability through monitoring, logging, tracing, alerting, and incident management practices to improve operational reliability and system performance.
  • Implement engineering standards for testing, code quality, security, maintainability, scalability, latency optimization, and cost efficiency.
  • Define reusable platform patterns, developer tooling, and engineering workflows that improve developer productivity and operational consistency across the AI/ML team.
  • Evaluate emerging AI engineering trends, AI-assisted development tools, and modern software practices to drive continuous improvement and innovation.
  • Partner with product, security, data, and platform teams to deliver production-ready AI solutions while contributing to architectural discussions and long-term platform strategy.
  • Troubleshoot complex production issues, perform root cause analysis, and drive remediation efforts to improve system stability and reliability.
  • Mentor engineers through technical collaboration, code reviews, knowledge sharing, and engineering best practices.
Qualifications:
  • Relevant degree preferred.
  • 5 or more years of relevant experience required.
  • Strong Python development experience with expertise building production backend services, APIs, distributed systems, and cloud-based applications required.
  • Hands-on experience with Pulumi, Terraform, or other infrastructure-as-code tools along with cloud platforms such as AWS, Azure, or GCP required.
  • Experience with Docker, CI/CD pipelines, infrastructure automation, container orchestration, Kubernetes, serverless architectures, and cloud deployment practices preferred.
  • Knowledge of observability tools, monitoring, logging, tracing, alerting frameworks, and production support practices.
  • Familiarity with ML/AI systems, model serving, LLM applications, inference pipelines, RAG workflows, data pipelines, or related AI platform technologies.
  • Experience with Databricks, Azure AI Foundry, or similar AI/ML platform technologies preferred.
  • Strong analytical, troubleshooting, problem-solving, verbal communication, and written communication skills with the ability to collaborate across technical and business teams.
  • Ability to operate effectively in fast-paced, evolving environments with a high level of ownership, accountability, and adaptability.
Estimated Hiring Range:

At Vizient, we consider skills, experience, and organizational needs in our compensation approach. Geographic factors may adjust the range estimate and hires typically fall below the top range. Compensation decisions are tailored to individual circumstances. The current salary range for this role is $102,400.00 to $179,000.00.

This position is also incentive eligible.

Equal Opportunity Employer: Females/Minorities/Veterans/Individuals with Disabilities

Vizient has a comprehensive benefits plan! Please view our benefits here:

about-us/careers

The Company is committed to equal employment opportunity to all employees and applicants without regard to race, religion, color, gender identity, ethnicity, age, national origin, sexual orientation, disability status, veteran status or any other category protected by applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Operations Engineer (Senior Software Engineer)
Senior AI Operations Engineer (Senior Software Engineer)

Vizient, Inc. • Irving (TX)

On-site
USD 102,000 - 179,000
Comprehensive benefits
Incentive pay
Senior AI Operations Engineer (Senior Software Engineer)
Senior AI Operations Engineer (Senior Software Engineer)

Vizient, Inc. • Chicago (IL)

On-site
USD 102,000 - 179,000
Incentive eligible
Senior AI Operations Engineer (Senior Software Engineer)
Senior AI Operations Engineer (Senior Software Engineer)

Vizient • Irving (TX)

On-site
USD 102,000 - 179,000
Comprehensive benefits plan
Senior Software Engineer
Senior Software Engineer

Vizient, Inc. • Chicago (IL)

On-site
USD 102,400 - 179,000
Senior Software Engineer
Senior Software Engineer

Vizient, Inc. • Centennial (CO)

On-site
USD 102,400 - 179,000
Senior Software Engineer
Senior Software Engineer

Vizient • Centennial (CO)

On-site
USD 102,000 - 179,000
Incentive eligible
Senior Software Engineer
Senior Software Engineer

Vizient • Chicago (IL)

On-site
USD 102,000 - 179,000
Senior Software Engineer
Senior Software Engineer

Vizient • Edina (MN)

On-site
USD 102,000 - 179,000
Comprehensive benefits plan
Incentive eligibility
Senior Software Engineer
Senior Software Engineer

RXinsider LTD. • Edina (MN)

On-site
USD 102,000 - 179,000
Sr AI/ML Engineer | Vizient
Sr AI/ML Engineer | Vizient

Vizient • Chicago (IL)

On-site
USD 102,000 - 179,000
Comprehensive benefits