Infra DevOps and Backend Engineer

GMI Cloud

Mountain View (CA)

On-site

USD 160,000 - 210,000

Full time

4 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

GMI Cloud is seeking an Infrastructure Backend Engineer to design, build, and maintain scalable AI infrastructure in Mountain View, CA. The role emphasizes cloud computing, distributed systems, and DevOps practices to enable efficient AI infrastructure operations.

You will contribute to large-scale training and inference, automate resource provisioning, and implement telemetry with Prometheus, Grafana, and Mimir while staying current with GPU technology.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • Proficiency in at least one programming language (Golang, Python, Bash) with strong coding practices.
  • Experience with infrastructure orchestration platforms, especially OpenStack and Kubernetes.
  • Experience with automation and CI/CD using Ansible, Jenkins, or GitLab CI.
  • Experience with telemetry/observability tools such as Prometheus, Grafana, and Mimir.
  • Knowledge of networking, GPU clusters, security, and performance tuning.

Responsibilities

  • Design, implement, and maintain AI/ML infrastructure optimized for large-scale training and inference.
  • Develop automation pipelines for GPU/CPU resource provisioning and workload scheduling using DevOps best practices.
  • Develop observability and telemetry solutions to pro-actively monitor hardware performance, utilization, and health to ensure cluster reliability and efficiency.
  • Optimize infrastructure for high-throughput data transfer and low-latency communication.
  • Manage infrastructure security, access controls, and compliance standards for on-prem GPU cluster environments.
  • Collaborate with relevant engineering teams to configure and troubleshoot GPU clusters and hardware resources.
  • Document infrastructure architecture, deployment procedures, automation workflow, and operational best practices.
  • Stay current with the latest GPU technology developments, infrastructure engineering and integrate new hardware/software solutions as appropriate.

Skills

Golang
Python
Bash

Education

Bachelor’s degree in Computer Science or related field

Tools

OpenStack
Kubernetes
Ansible
Jenkins
GitLab CI
Prometheus
Grafana
Mimir
HashiCorp Vault
Vastdata
Weka
DDN
Ceph

Job description

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents. Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter. From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud. One cloud for compute, inference, and agents.

Role Overview

We are seeking a talented and highly skilled Infrastructure Backend Engineering Development Engineer to design, build, and maintain the scalable infrastructure that supports GMI AI/ML initiatives. The ideal candidate will have a strong background in cloud computing, distributed systems, and DevOps practices to enable efficient AI infrastructure operations.

Responsibilities
  • Design, implement, and maintain AI/MLinfrastructure optimized for large-scale training and inference.
  • Develop automation pipelines for GPU/CPU resource provisioning and workload scheduling using DevOps best practices and methodologies.
  • Develop observability and telemetry solutions to pro-actively monitor hardware performance, utilization, and health to ensure cluster reliability and efficiency.
  • Optimize infrastructure for high-throughput data transfer and low-latency communication.
  • Manage infrastructure security, access controls, and compliance standards for on-prem GPU cluster environments.
  • Collaborate with relevant engineering teams to configure and troubleshoot GPU clusters and hardware resources.
  • Document infrastructure architecture, deployment procedures, automation workflow, and operational best practices.
  • Stay current with the latest GPU technology developments, infrastructure engineering and integrate new hardware/software solutions as appropriate.
Qualifications
  • Bachelor’s degree in Computer Science or related field.
  • Proficiency in at least one programming language (Golang, Python, Bash) with strong coding practices and system design skills.
  • Extensive experience with infrastructure orchestration platforms, especially OpenStack and Kubernetes.
  • Strong proficiency in automation, configuration management, and CI/CD pipelines using Ansible, Jenkins, GitLab CI, or similar.
  • Proven experience implementing telemetry and observability solutions (e.g. Prometheus, Grafana, Mimir and related technologies).
  • Knowledge of networking, security, and performance tuning in GPU clusters.
  • Hands-on experience in deploying GPU clusters and managing GPU workloads.
  • Experience with secret management using HashiCorp Vault.
  • Experience with distributed and high performance storage systems (e.g. Vastdata, Weka, DDN, Ceph)
  • Strong system thinking and abstraction skills, capable of designing complex distributed systems from an end-to-end perspective.
  • Strong understanding of DevOps principles, automation, and cloud-native architectures.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
Senior Technical Account Manager
Senior Technical Account Manager

GMI Cloud • Mountain View (CA)

On-site
USD 140,000 - 180,000
AI Infra DevOps & Backend Engineer
AI Infra DevOps & Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
Technical Program Manager, Data Center Infrastructure Delivery
Technical Program Manager, Data Center Infrastructure Delivery

GMI Cloud • United States

On-site
USD 170,000 - 230,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Product Manager - GPUaaS and OE Telemetry
Product Manager - GPUaaS and OE Telemetry

GMI Cloud • Mountain View (CA)

On-site
USD 120,000 - 160,000