Infrastructure Engineer

Emergys

Pune District

On-site

INR 900,000 - 1,500,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Emergys is seeking a senior systems engineer to deploy, manage, and optimize AI/ML workloads across GPU clusters, HPC, and cloud environments. You will build and maintain scalable platform infrastructure with Kubernetes, containers, and orchestration platforms, and administer Linux servers with focus on security, performance, and reliability.

The role requires hands-on expertise in cloud-native infrastructure, GPU-based workloads, and enterprise networking, with strong scripting skills in Bash

Qualifications

  • 8+ years of experience in Linux systems administration, cloud-native infrastructure, HPC environments, or platform engineering.
  • 4+ years of experience supporting AI/ML workloads or large-scale distributed compute environments in production.
  • Comfortable leveraging AI-assisted tools for collaborative development and productivity.

Responsibilities

  • Deploy, manage, and optimize AI/ML inference workloads across GPU clusters, HPC, and cloud environments.

Skills

Linux administration
Kubernetes administration
Cloud-native infrastructure
SRE/Platform engineering
Bash scripting
Python scripting
CI/CD workflows
Networking fundamentals

Tools

Kubernetes
Bash
Python
CI/CD tooling

Job description

  • Deploy, manage, and optimize AI/ML and LLM inference workloads across GPU clusters, HPC infrastructure, and cloud environments.
  • Build and maintain scalable AI platform infrastructure using Kubernetes, containers, and enterprise orchestration platforms.
  • Administer and optimize Linux servers including system configuration, patching, security hardening, performance tuning, and troubleshooting.
  • Manage physical infrastructure including servers, storage, networking, and bare metal environments within enterprise data centers.
  • Implement and maintain CI/CD and automation workflows for platform and infrastructure deployments.
  • Optimize infrastructure performance, GPU utilization, resource allocation, and distributed workloads to meet operational requirements.
  • Benchmark and evaluate AI workloads for scalability, latency, throughput, and resource efficiency.
  • Collaborate with infrastructure, SRE, and platform engineering teams to provision compute resources and maintain enterprise-scale AI environments.
  • Implement monitoring, logging, observability, and alerting solutions for platform reliability and operational visibility.
  • Apply security patches, upgrades, compliance controls, and operational best practices for Linux and Kubernetes environments.
  • Troubleshoot issues across hardware, networking, operating systems, Kubernetes clusters, and AI/ML workloads.
  • Support enterprise operations through efficient incident, change, and ticket management processes.
  • Automate infrastructure operations using scripting and infrastructure automation tools.
Required Qualifications:
  • 8+ years of experience in Linux systems administration, cloud-native infrastructure, HPC environments, or platform engineering.
  • At least 4 years of experience supporting AI/ML workloads or large-scale distributed compute environments in production.
  • Comfortable leveraging AI-assisted tools for collaborative development, code generation, refactoring, and productivity enhancement.
  • Strong hands-on expertise with Linux administration (RHEL, Ubuntu, or similar).
  • Experience with Kubernetes administration, container orchestration, and cloud native infrastructure platforms.
  • Strong understanding of GPU infrastructure, distributed computing, and HPC systems.
  • Hands-on experience with bare-metal infrastructure, servers, storage systems, and enterprise networking.
  • Strong understanding of networking fundamentals including TCP/IP, DNS, load balancing, and firewalls.
  • Experience with scripting and infrastructure automation with Bash, Python etc.
  • Experience with CI/CD, DevOps, or infrastructure deployment workflows.
  • Experience with monitoring, observability, and logging platforms.
  • Strong troubleshooting and performance optimization skills across Linux, Kubernetes, networking, and infrastructure stacks.
  • Excellent problem-solving, communication, and collaboration skills.
  • Ability to work effectively in fast-paced, mission-critical production environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)

Annova Solutions • Indore District

On-site
INR 2,800,000 - 4,200,000
AI Infrastructure and Platform Architect
AI Infrastructure and Platform Architect

Ignatiuz Inc. • Indore District

On-site
INR 3,500,000 - 6,000,000
Senior Cloud Engineer
Senior Cloud Engineer

Neysa • Mumbai

On-site
INR 400,000 - 650,000
DevOps Engineer
DevOps Engineer

Hermes Corporate • India

On-site
INR 900,000 - 1,400,000
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)
Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)

Annova Solutions Corp. • Indore District

On-site
INR 2,400,000 - 4,200,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA Gruppe • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Head – AI Infrastructure/Data Center Operations
Head – AI Infrastructure/Data Center Operations

LinkCxO (The CxO's Marketplace) • Chennai District

On-site
INR 4,000,000 - 6,500,000
Senior HPC Cluster Engineer - AI, ML
Senior HPC Cluster Engineer - AI, ML

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,400,000
Senior DevOps Cloud Platform Engineer AWS Kubernetes AI Infrastructure
Senior DevOps Cloud Platform Engineer AWS Kubernetes AI Infrastructure

Annova Solutions India • Indore District

On-site
INR 1,800,000 - 3,200,000