Senior Site Reliability and HPC Infrastructure Engineer

Synopsys India Pvt Ltd

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Synopsys India Pvt Ltd is seeking a Senior Site Reliability professional to design, build, and optimize large-scale HPC compute farm platforms powering chip design tools. You will tune schedulers, lead capacity planning, and drive automation to improve reliability and performance for thousands of concurrent engineering workloads.

You will collaborate with R&D, Cloud, Infrastructure, and Security teams on cloud-integrated HPC initiatives and mentor engineers across global sites.

Qualifications

  • 8+ years of Linux/UNIX systems administration experience.
  • 5+ years HPC or compute farm administration, including workload scheduling and capacity planning.
  • Strong expertise in IBM Spectrum LSF, Slurm, or equivalent schedulers and policy configuration.
  • Advanced knowledge of LDAP, NFS, DNS, and enterprise storage networking for distributed compute.
  • Proven Python and Shell scripting for automation, monitoring, and orchestration.
  • Solid experience with Grafana, Prometheus, Elastic, or Splunk for proactive incident detection.
  • Experience with cloud-integrated HPC on Azure or AWS, Kubernetes, Docker, Ansible, or Terraform.

Responsibilities

  • Design, build, and optimize large-scale HPC compute farm platforms for thousands of workloads across global sites.
  • Tune IBM Spectrum LSF, Slurm, or equivalents to maximize utilization and minimize queue times.
  • Lead capacity planning and translate engineering demand into infra requirements and timelines.
  • Drive automation using Python and Shell scripting to reduce manual toil and accelerate incident response.
  • Lead complex troubleshooting and root cause analysis across storage, networking, LDAP, NFS, and scheduler layers.
  • Collaborate on strategic initiatives including cloud-integrated HPC and AI/ML workload enablement.
  • Mentor junior engineers and provide technical leadership across global teams; participate in 24x5 support and major infra projects.

Skills

Linux/UNIX admin
HPC admin
Scheduling
Python scripting
Shell scripting
Monitoring
Cloud HPC
Kubernetes

Tools

IBM Spectrum LSF
Slurm
LDAP
NFS
DNS
Grafana
Prometheus
Elastic
Splunk
Azure
AWS
Kubernetes
Docker
Ansible
Terraform

Job description

General Information

Site Reliability, Staff / HPC Infrastructure Engineer

Job Title Site Reliability, Staff / HPC Infrastructure Engineer

Job ID 18460

Country India

City Bengaluru

Date Posted 24-Aug-2026

Job Category Information Technology

Job Subcategory Site Reliability

Hire Type Employee

Remote Eligible No

Descriptions & Requirements

Job Description and Requirements

We Are Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to rapidly innovate AI-powered products. We deliver industry-leading silicon design, IP, simulation and analysis solutions, and design services. We partner closely with our customers across a wide range of industries to maximize their R&D capability and productivity, powering innovation today that ignites the ingenuity of tomorrow.

You Are

You have spent years keeping large-scale compute environments running, not just available but actually performant, under the kind of load that makes most infrastructure buckle. You understand that when an engineer's simulation job sits in queue for six hours instead of six minutes, that is not a scheduler problem, it is a planning problem, a tuning problem, or a resource allocation problem, and you are the person who figures out which one and fixes it before it becomes a pattern. You think in systems, not tickets. When LSF or Slurm throws an error, you do not just restart the service, you trace it back to the workload spike, the misconfigured policy, or the storage bottleneck that triggered it. You have built enough automation to know that the script is the easy part, the hard part is designing it so it does not create three new problems when the environment shifts next quarter. You are comfortable in a room with R&D engineers who need 10,000 cores by Tuesday and cloud architects who want to move half the farm to Azure, and you can translate between those conversations without losing the thread of what actually has to work. At Synopsys, you will work on HPC infrastructure that powers the chip design tools the world depends on, and the decisions you make will directly affect how fast our engineers can build the next generation of silicon.

What You'll Be Doing
  • Design, build, and optimize large-scale HPC compute farm platforms that support thousands of engineering workloads across global sites
  • Administer and tune IBM Spectrum LSF, Slurm, or equivalent workload schedulers to maximize resource utilization and minimize job queue times
  • Lead capacity planning and workload optimization efforts, translating engineering demand into infrastructure requirements and deployment timelines
  • Drive automation using Python and Shell scripting to eliminate manual toil, improve reliability, and accelerate incident response
  • Lead complex troubleshooting and root cause analysis for platform issues, working across storage, networking, LDAP, NFS, and scheduler layers
  • Collaborate with R&D, Cloud, Infrastructure, and Security teams on strategic initiatives including cloud-integrated HPC and AI/ML workload enablement
  • Mentor junior engineers and provide technical leadership across global teams, setting standards for operational excellence and engineering rigor
  • Participate in 24x5 support operations and lead major infrastructure projects from design through deployment
The Impact You Will Have
  • Enable faster product development cycles by ensuring high availability and performance of the compute infrastructure that powers Synopsys EDA tools
  • Maximize infrastructure utilization and license efficiency, directly reducing costs and improving engineering productivity across the company
  • Reduce incident response time and platform downtime through automation, monitoring, and proactive capacity management
  • Accelerate cloud adoption and modernization efforts, helping Synopsys scale compute resources dynamically to meet global engineering demand
  • Improve workload throughput and job completion times, giving engineers more iterations per day and faster feedback loops
  • Build operational resilience into the platform, ensuring that infrastructure scales reliably as the business grows
  • Mentor and elevate the technical capabilities of the global infrastructure engineering team, raising the bar for how we operate and support critical systems
What You'll Need
  • 8+ years of Linux/UNIX systems administration experience with deep expertise in performance tuning, troubleshooting, and large-scale operations
  • 5+ years of hands-on HPC or compute farm administration, including workload scheduling, resource management, and capacity planning
  • Strong expertise in IBM Spectrum LSF, Slurm, or equivalent schedulers, including policy configuration, job prioritization, and performance optimization
  • Advanced knowledge of LDAP, NFS, DNS, enterprise storage systems, and networking in the context of distributed compute environments
  • Proven experience with Python and Shell scripting for automation, monitoring, and infrastructure orchestration
  • Solid understanding of monitoring and observability tools such as Grafana, Prometheus, Elastic, or Splunk for proactive incident detection and analysis
  • Experience with EDA environments is a strong plus, as is familiarity with cloud-integrated HPC platforms on Azure or AWS, Kubernetes, Docker, Ansible, or Terraform
Who You Are
  • You take ownership of problems end to end, from the first alert to the post-incident review, and you do not consider something fixed until you understand why it broke
  • You are self-driven and proactive, the kind of engineer who sees a pattern in the logs and writes the script to prevent it before anyone files a ticket
  • You can explain a complex infrastructure tradeoff to a VP in two sentences without losing the technical nuance, and you can turn that same conversation into a design doc for your team
  • You are collaborative and team-oriented, comfortable working across time zones and disciplines to align on solutions that work for everyone
  • You have strong analytical and problem-solving instincts, you do not guess, you instrument, measure, and validate before you change production
  • You are focused on service excellence and customer impact, you measure success by whether engineering teams can do their jobs faster and more reliably because of the infrastructure you built
The Team

You'll Be Part Of The Compute Farm Infrastructure Engineering team operates and enhances Synopsys' global HPC and compute farm platforms. You will work with a distributed team responsible for providing reliable, scalable engineering compute services that support mission-critical EDA workloads. The team drives automation, modernization, and cloud adoption while maintaining high availability and operational excellence across all global sites.

Rewards and Benefits

We offer a comprehensive range of health, wellness, and financial benefits to cater to your needs. Our total rewards include both monetary and non-monetary offerings. At Synopsys, we want talented people of every background to feel valued and supported to do their best work. Synopsys considers all applicants for employment without regard to race, color, religion, national origin, gender, sexual orientation, age, military veteran status, or disability.

Experience Level Senior Level

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Research and Development Engineer - Embedded Software and FPGA
Senior Research and Development Engineer - Embedded Software and FPGA

Synopsys India Pvt Ltd • Bengaluru

On-site
INR 1,200,000 - 1,800,000
HPC Engineer
HPC Engineer

Sandisk • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Senior Applications Engineer
Senior Applications Engineer

Synopsys India Pvt Ltd • Bengaluru

On-site
INR 1,400,000 - 2,100,000
Cognizant Hiring For Sr. HPC Engineer
Cognizant Hiring For Sr. HPC Engineer

Cognizant • Pune District, Chennai District, Bengaluru

On-site
INR 3,000,000 - 6,000,000
R&D Engineer, Sr Engineer (C/C++, Data structures, Algorithm, FPGA)
R&D Engineer, Sr Engineer (C/C++, Data structures, Algorithm, FPGA)

Synopsys, Inc. • India

On-site
INR 1,500,000 - 2,100,000
Health insurance
Lead Technology Specialist
Lead Technology Specialist

Airbus • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Site Reliability Engineer - Cloud Infrastructure
Site Reliability Engineer - Cloud Infrastructure

Synechron Technologies • Bengaluru

On-site
INR 2,400,000 - 4,000,000
Specialist Software Engineer
Specialist Software Engineer

Amgen SA • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Staff Data Engineer (HPC cluster software such as Slurm, NC, LSF or Grid Engine) experience wit[...]
Staff Data Engineer (HPC cluster software such as Slurm, NC, LSF or Grid Engine) experience wit[...]

SanDisk • Bengaluru

On-site
INR 3,000,000 - 5,000,000
Staff Application Engineer - Physical Implementation
Staff Application Engineer - Physical Implementation

Synopsys Inc • Dadri

On-site
INR 4,000,000 - 7,000,000