AI Infrastructure System Engineer

Lever, Inc.

India

On-site

INR 1,800,000 - 3,000,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote-friendly (Bangalore-based)

Job summary

Lever, Inc. in Bangalore (India) seeks an AI Infrastructure System Engineer to build and operate large-scale infrastructure for AI training and inference workloads. You will engineer systems that manage thousands of GPUs with a focus on automation, reliability, performance, and availability.

You will collaborate with infrastructure, hardware, networking, platform, and AI teams to solve complex challenges, building software-driven, autonomous infrastructure at scale.

Qualifications

  • 3+ years building distributed systems or backend infrastructure.
  • Proficient in Python, Go, or Rust.
  • Experience with Linux and modern infra tools like Kubernetes, Terraform, Ansible.
  • Ability to design software-driven infrastructure and automation.
  • Strong problem-solving and collaboration across teams.

Responsibilities

  • Design and build fleet automation to manage GPU clusters with minimal human intervention.
  • Develop AI infrastructure agents to automate deployment workflows and remediation.
  • Build fleet intelligence platforms monitoring health, firmware, networking, and perf.
  • Create predictive capabilities to anticipate infrastructure failures.
  • Improve deployment velocity and reliability via automation.
  • Collaborate with hardware, networking, platform, and AI teams.

Skills

Python
Go
Rust
Distributed systems
Linux

Tools

Kubernetes
Terraform
Ansible

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Infrastructure System Engineerbased in India.

This role offers the opportunity to build and operate large-scale infrastructure supporting advanced AI training and inference workloads.
You’ll engineer systems that manage thousands of GPUs with a strong focus on automation, reliability, performance, and availability.
The position goes beyond traditional infrastructure operations, emphasizing software-driven solutions and autonomous systems.
You’ll develop platforms that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal manual intervention.
Your work will span hardware, networking, storage, distributed systems, and AI workloads.
You’ll collaborate closely with infrastructure, hardware, networking, platform, and AI engineering teams to solve complex systems challenges.
It’s an ideal environment for an automation-focused engineer who enjoys building intelligent infrastructure at significant scale.

Accountabilities:
  • Design and build fleet automation systems capable of provisioning, validating, deploying, upgrading, repairing, and retiring GPU clusters with minimal human intervention.
  • Develop AI infrastructure agents that automate deployment workflows, investigate root causes, triage incidents, and support autonomous remediation.
  • Build fleet intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance.
  • Develop predictive capabilities that identify potential infrastructure failures before they affect customers or workloads.
  • Build software and automation systems that maximize GPU availability, utilization, performance, and reliability across large accelerator fleets.
  • Create automated validation frameworks for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage systems, and distributed AI workloads.
  • Develop internal infrastructure platforms and developer tools that enable infrastructure to be managed programmatically rather than through manual operations.
  • Continuously improve deployment velocity, system reliability, and operational efficiency through automation and software engineering.
  • Collaborate with hardware, networking, platform, and AI teams to identify infrastructure challenges and develop scalable solutions.
  • Apply strong systems thinking to problems spanning hardware and software components across large-scale AI infrastructure.
Requirements:
  • 3+ years of experience building distributed systems, infrastructure platforms, or large-scale backend software.
  • Strong software engineering skills in Python, Go, or Rust.
  • Proven experience developing platforms, automation systems, developer infrastructure, or similar software-driven infrastructure solutions.
  • Experience working with Linux and modern infrastructure technologies such as Kubernetes, Terraform, Ansible, or comparable tools.
  • Strong understanding of distributed systems and the ability to reason across hardware and software layers.
  • Passion for solving complex infrastructure problems through software and automation.
  • Strong automation-first mindset, with an instinct to build systems that eliminate repetitive manual tasks.
  • Ability to work effectively on complex technical challenges involving reliability, performance, scalability, and operational efficiency.
  • Strong collaboration skills and the ability to partner effectively with engineering teams across infrastructure, hardware, networking, and AI.
  • Experience with GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch, or related technologies is a plus.
  • Experience with InfiniBand or RoCE networking is advantageous.
  • Familiarity with bare-metal provisioning and infrastructure lifecycle management is desirable.
  • Experience supporting large-scale AI training or inference clusters is a plus.
  • Knowledge of hardware health monitoring and predictive failure detection is beneficial.
  • Experience with distributed storage systems is advantageous.
  • Familiarity with AI agents or autonomous infrastructure operations is a plus.
Benefits:
  • Opportunity to work on large-scale AI infrastructure supporting advanced training and inference workloads.
  • Exposure to complex systems spanning GPUs, networking, storage, distributed computing, and AI workloads.
  • Opportunity to build highly automated infrastructure systems and developer platforms.
  • Collaboration with multidisciplinary engineering teams working across hardware, networking, platform, and AI.
  • Environment focused on software-driven infrastructure, automation, scalability, reliability, and performance.
  • Opportunity to contribute to systems operating at significant GPU scale.
  • Remote-friendly job listing based in Bangalore, India.
How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Software Engineer - NV Cloud Functions
Senior Systems Software Engineer - NV Cloud Functions

Jobgether • India

On-site
INR 4,000,000 - 7,000,000
Competitive salary
Open-source collaboration
Global engineering exposure
+1
AI infrastructure System Engineer Bangalore
AI infrastructure System Engineer Bangalore

Together AI • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Technical Support Engineer (GPU Cluster), India
Technical Support Engineer (GPU Cluster), India

Togetherai • Pune District, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Health insurance
Remote work
Equity
Technical Support Engineer (GPU Cluster), India
Technical Support Engineer (GPU Cluster), India

Together AI • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Startup equity
Health insurance
Remote work flexibility
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • India

Remote
INR 3,500,000 - 7,000,000
Remote work
International team
Cutting-edge AI projects
Senior Infrastructure Engineer
Senior Infrastructure Engineer

SourcingXPress • Hyderabad

On-site
INR 2,000,000 - 3,500,000
High ownership
Rapid learning opportunities
Career advancement potential
Technical Support Engineer (GPU Cluster), India
Technical Support Engineer (GPU Cluster), India

Together • Pune District, Bengaluru

On-site
INR 900,000 - 1,500,000
Health insurance
Startup equity
Remote work flexibility
Junior/Senior or Staff Software Engineer, Inference / Compute Infrastructure Engineering
Junior/Senior or Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together • India

Remote
INR 4,000,000 - 6,000,000
Customer Support Engineer (GPU Cluster), India
Customer Support Engineer (GPU Cluster), India

Together AI • India

On-site
INR 1,800,000 - 3,200,000
Competitive compensation
Startup equity
Health insurance
+1
Technical Support Engineer (Inference) - India Weekends
Technical Support Engineer (Inference) - India Weekends

Togetherai • Pune District, Bengaluru

On-site
INR 2,500,000 - 3,500,000
Health insurance
Startup equity
Remote work flexibility