AI Infra DevOps & Backend Engineer

GMI Cloud

Mountain View (CA)

On-site

USD 160,000 - 210,000

Full time

6 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

GMI Cloud is seeking an Infrastructure Backend Engineer to design, build, and maintain scalable AI infrastructure in Mountain View, CA. The role emphasizes cloud computing, distributed systems, and DevOps practices to enable efficient AI infrastructure operations.

You will contribute to large-scale training and inference, automate resource provisioning, and implement telemetry with Prometheus, Grafana, and Mimir while staying current with GPU technology.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • Proficiency in at least one programming language (Golang, Python, Bash) with strong coding practices.
  • Experience with infrastructure orchestration platforms, especially OpenStack and Kubernetes.
  • Experience with automation and CI/CD using Ansible, Jenkins, or GitLab CI.
  • Experience with telemetry/observability tools such as Prometheus, Grafana, and Mimir.
  • Knowledge of networking, GPU clusters, security, and performance tuning.

Responsibilities

  • Design, implement, and maintain AI/ML infrastructure optimized for large-scale training and inference.
  • Develop automation pipelines for GPU/CPU resource provisioning and workload scheduling using DevOps best practices.
  • Develop observability and telemetry solutions to pro-actively monitor hardware performance, utilization, and health to ensure cluster reliability and efficiency.
  • Optimize infrastructure for high-throughput data transfer and low-latency communication.
  • Manage infrastructure security, access controls, and compliance standards for on-prem GPU cluster environments.
  • Collaborate with relevant engineering teams to configure and troubleshoot GPU clusters and hardware resources.
  • Document infrastructure architecture, deployment procedures, automation workflow, and operational best practices.
  • Stay current with the latest GPU technology developments, infrastructure engineering and integrate new hardware/software solutions as appropriate.

Skills

Golang
Python
Bash

Education

Bachelor’s degree in Computer Science or related field

Tools

OpenStack
Kubernetes
Ansible
Jenkins
GitLab CI
Prometheus
Grafana
Mimir
HashiCorp Vault
Vastdata
Weka
DDN
Ceph

Job description

GMI Cloud is seeking an Infrastructure Backend Engineer to design, build, and maintain scalable AI infrastructure in Mountain View, CA. The role emphasizes cloud computing, distributed systems, and DevOps practices to enable efficient AI infrastructure operations.

You will contribute to large-scale training and inference, automate resource provisioning, and implement telemetry with Prometheus, Grafana, and Mimir while staying current with GPU technology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infra DevOps and Backend Engineer
Infra DevOps and Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
AI Infrastructure Engineer — Cloud, GPU & MLOps
AI Infrastructure Engineer — Cloud, GPU & MLOps

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
AI Systems Engineer: Cloud, GPUs & Automation
AI Systems Engineer: Cloud, GPUs & Automation

MCI • United States

On-site
USD 90,000 - 130,000
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
Senior AI Infra Engineer — Telemetry & ML Ops Equity
Senior AI Infra Engineer — Telemetry & ML Ops Equity

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000
Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
AI Infrastructure Intern: GPU & Cloud Platform (Remote)
AI Infrastructure Intern: GPU & Cloud Platform (Remote)

Meshy AI • San Francisco (CA)

On-site
USD 55,000 - 69,000
Equity
Flexible time off
Remote work options
+1
Senior AI Infrastructure Engineer - Remote & Scaled GPU
Senior AI Infrastructure Engineer - Remote & Scaled GPU

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 160,000