Senior Infrastructure Engineer

Jobgether

India

Hybrid

INR 8,247,000 - 12,371,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary
Annual discretionary bonus
Remote or hybrid options
Employee wellbeing benefits
25 days annual holiday
Autonomy and ownership

Job summary

Jobgether is seeking a Senior Infrastructure Engineer to own and evolve a high-performance GPU-focused infrastructure. You will design, deploy, and operate OpenStack and Kubernetes at scale, with strong emphasis on reliability, security, and automation.

In a fast-paced, globally distributed environment, you will collaborate with Platform, DevOps, AI, and Product teams to deliver scalable infrastructure and proactive incident response strategies.

Qualifications

  • Extensive hands-on Linux systems administration
  • Experience building, cabling, and commissioning servers and racks
  • Willingness to travel to data-centre sites in Quebec
  • Production experience with OpenStack and Kubernetes at scale

Responsibilities

  • Design, deploy, operate, and improve OpenStack and Kubernetes for GPU workloads
  • Take ownership of infrastructure performance, scalability, and reliability
  • Automate provisioning, deployment, and recurring workflows
  • Improve monitoring, logging, alerting, and observability
  • Lead incident response and post-incident improvements
  • Maintain security controls across infrastructure and container environments
  • Build and commission physical server and GPU infrastructure in data-centres
  • Collaborate with Platform, DevOps, AI, Product, and Support teams

Skills

Linux admin
OpenStack
Kubernetes
Networking
GPU infra
Automation
GitOps
CI/CD
Security
SRE

Education

Bachelor's in CS

Tools

NVIDIA GPUs
CI/CD tools

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Infrastructure Engineer based in Canada.

As a Senior Infrastructure Engineer, you will take ownership of business-critical infrastructure supporting demanding AI and GPU workloads.
You will design, deploy, and operate large-scale OpenStack and Kubernetes environments with a strong focus on performance, reliability, and scalability.
The role combines hands-on systems engineering with automation, infrastructure-as-code, observability, security, and incident response.
You will work directly with physical servers, racks, networking, storage, and GPU infrastructure across data-centre environments.
Your work will have a visible impact on platform stability, customer experience, and the ability to scale globally.
You will collaborate closely with platform, DevOps, AI, product, and support teams in a fast-moving, technically ambitious environment.
This is an ownership-driven opportunity for an engineer who enjoys solving complex infrastructure problems and improving systems end-to-end.

Accountabilities:
  • Design, deploy, operate, and continuously improve OpenStack and Kubernetes environments supporting high-performance GPU workloads.
  • Take end-to-end ownership of infrastructure performance, scalability, resilience, and operational reliability.
  • Build and maintain infrastructure using infrastructure-as-code, GitOps, CI/CD, and Git-based operational practices.
  • Automate provisioning, deployment, configuration, and recurring infrastructure workflows to improve efficiency and consistency.
  • Optimize GPU workload scheduling using Kubernetes and NVIDIA technologies to maximize platform performance and resource utilization.
  • Implement and improve monitoring, logging, alerting, and observability systems to identify and resolve infrastructure issues proactively.
  • Lead incident response activities, troubleshoot complex production problems, and drive post-incident improvements that strengthen overall reliability.
  • Maintain robust security controls across infrastructure and container environments, including RBAC, network policies, access controls, and tenant isolation.
  • Build, rack, cable, configure, and commission physical server and GPU infrastructure within data-centre environments.
  • Work with networking and storage systems to ensure reliable, scalable infrastructure for compute-intensive workloads.
  • Collaborate with Platform, DevOps, AI, Product, and Support teams to align infrastructure capabilities with technical and customer requirements.
  • Contribute to the continuous evolution of infrastructure architecture and operational practices as global environments scale.
  • Identify opportunities to improve automation, performance, reliability, and operational efficiency across the infrastructure platform.
Requirements
  • Extensive hands-on experience administering Linux systems, with strong technical depth across operating systems, troubleshooting, and infrastructure operations.
  • Proven experience physically building, assembling, cabling, configuring, and commissioning servers and racks.
  • Direct hands-on experience working in data centres, including physically installing and racking hardware on-site.
  • Willingness and ability to travel to data-centre sites in Quebec as required.
  • Strong understanding of networking and storage technologies and their role within large-scale infrastructure environments.
  • Experience installing, racking, and configuring GPU hardware is highly advantageous, particularly with NVIDIA platforms.
  • Production experience operating OpenStack and/or Kubernetes at scale is strongly preferred.
  • Experience with infrastructure automation, infrastructure-as-code, CI/CD pipelines, GitOps, and Git-based workflows is an advantage.
  • Exposure to high-performance computing, large-scale compute, or other demanding infrastructure environments is desirable.
  • Experience with Kubernetes GPU scheduling and NVIDIA tooling is a plus.
  • Strong troubleshooting and analytical abilities, with a methodical approach to diagnosing complex infrastructure issues.
  • Ability to take ownership of systems from design through deployment, operation, troubleshooting, and continuous improvement.
  • Comfortable working in a fast-paced environment where priorities can evolve and engineers are trusted to make decisions independently.
  • Strong collaboration and communication skills, with the ability to work effectively across technical and non-technical teams.
  • Contributions to open-source infrastructure or technology projects are a plus.
Benefits
  • Competitive salary.
  • Annual discretionary bonus scheme.
  • Employee wellbeing benefits.
  • 25 days of annual holiday plus public holidays.
  • Flexible working arrangements, with remote or hybrid options depending on role and location.
  • Significant autonomy and ownership, with the freedom to take initiative and experiment.
  • Opportunity to work on challenging, high-performance AI and GPU infrastructure.
  • Visible impact on infrastructure reliability, scalability, and customer experience.
  • Clear career progression and professional growth opportunities.
  • Collaborative international environment built around trust, transparency, and ownership.
  • Opportunity to contribute to the evolution of a rapidly scaling AI infrastructure platform and its engineering culture.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior Staff SRE – Compute Platform
Senior Staff SRE – Compute Platform

NVIDIA Corporation • India

On-site
INR 4,500,000 - 9,000,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA AI • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Senior DevOps Engineer
Senior DevOps Engineer

NVIDIA Gruppe • Pune District

On-site
INR 4,000,000 - 7,000,000
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
Senior Solution Architect, Cloud Infrastructure (Maharashtra)

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Corporation • Pune District

On-site
INR 4,000,000 - 6,500,000
Senior CUDA Driver and DevOps Engineer
Senior CUDA Driver and DevOps Engineer

NVIDIA • India

On-site
INR 1,500,000 - 2,000,000
Competitive salaries
Comprehensive benefits package
Principal Engineer – Cluster Deployment
Principal Engineer – Cluster Deployment

Nava • Bengaluru

On-site
INR 4,200,000 - 6,200,000
Senior DevOps Engineer
Senior DevOps Engineer

NVIDIA • Pune District

On-site
INR 1,500,000 - 2,100,000