Site Reliability, Staff / HPC Infrastructure Engineer

Synopsys, Inc.

Hinoba-an

On-site

PHP 1,674,000 - 2,567,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Synopsys, Inc. is seeking an experienced HPC Infrastructure Engineer to design, build, and optimize large-scale compute farms that power flagship EDA workloads.

You will administer schedulers, lead capacity planning, and drive automation across global teams. You will collaborate with R&D, Cloud, Infrastructure, and Security to enable cloud-integrated HPC and AI/ML workloads while mentoring engineers and delivering high-availability platforms.

Qualifications

  • 8+ years of Linux/UNIX systems administration with performance tuning and large-scale operations.
  • 5+ years in HPC/compute farm administration including workload scheduling, resource management, and capacity planning.
  • Deep expertise with IBM Spectrum LSF, Slurm, or equivalent schedulers and policy configuration.
  • Strong knowledge of LDAP, NFS, DNS, enterprise storage, and networking in distributed compute environments.
  • Proven automation experience with Python and shell scripting for orchestration and monitoring.
  • Proficient with Grafana, Prometheus, Elastic, or Splunk for proactive incident detection and analysis.
  • Familiarity with cloud-integrated HPC on Azure/AWS, Kubernetes, Docker, Ansible, or Terraform.

Responsibilities

  • Design, build, and optimize large-scale HPC compute farm platforms supporting thousands of workloads across global sites.
  • Administer and tune IBM Spectrum LSF, Slurm, or equivalents to maximize resource utilization and minimize queue times.
  • Lead capacity planning and translate engineering demand into infrastructure requirements and timelines.
  • Drive automation using Python and Shell scripting to reduce toil, improve reliability, and accelerate incident response.
  • Lead complex troubleshooting and root cause analysis across storage, networking, LDAP, NFS, and scheduler layers.
  • Collaborate with R&D, Cloud, Infrastructure, and Security on cloud-integrated HPC and AI/ML workload enablement.
  • Mentor junior engineers and provide technical leadership across global teams, guiding operational standards.

Skills

Linux/UNIX
HPC/Compute
Python
Shell scripting
LDAP
NFS
DNS
Monitoring
Cloud platforms
Kubernetes
Docker
Ansible
Terraform
Azure
AWS
LSF
Slurm

Tools

IBM Spectrum LSF
Slurm
Azure
AWS
Kubernetes
Docker
Ansible
Terraform

Job description

We Are

Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to rapidly innovate AI-powered products. We deliver industry-leading silicon design, IP, simulation and analysis solutions, and design services. We partner closely with our customers across a wide range of industries to maximize their R&D capability and productivity, powering innovation today that ignites the ingenuity of tomorrow.

You Are

You have spent years keeping large-scale compute environments running, not just available but actually performant, under the kind of load that makes most infrastructure buckle. You understand that when an engineer's simulation job sits in queue for six hours instead of six minutes, that is not a scheduler problem, it is a planning problem, a tuning problem, or a resource allocation problem, and you are the person who figures out which one and fixes it before it becomes a pattern.

You think in systems, not tickets. When LSF or Slurm throws an error, you do not just restart the service, you trace it back to the workload spike, the misconfigured policy, or the storage bottleneck that triggered it. You have built enough automation to know that the script is the easy part, the hard part is designing it so it does not create three new problems when the environment shifts next quarter.

You are comfortable in a room with R&D engineers who need 10,000 cores by Tuesday and cloud architects who want to move half the farm to Azure, and you can translate between those conversations without losing the thread of what actually has to work. At Synopsys, you will work on HPC infrastructure that powers the chip design tools the world depends on, and the decisions you make will directly affect how fast our engineers can build the next generation of silicon.

What You'll Be Doing
  • Design, build, and optimize large-scale HPC compute farm platforms that support thousands of engineering workloads across global sites
  • Administer and tune IBM Spectrum LSF, Slurm, or equivalent workload schedulers to maximize resource utilization and minimize job queue times
  • Lead capacity planning and workload optimization efforts, translating engineering demand into infrastructure requirements and deployment timelines
  • Drive automation using Python and Shell scripting to eliminate manual toil, improve reliability, and accelerate incident response
  • Lead complex troubleshooting and root cause analysis for platform issues, working across storage, networking, LDAP, NFS, and scheduler layers
  • Collaborate with R&D, Cloud, Infrastructure, and Security teams on strategic initiatives including cloud-integrated HPC and AI/ML workload enablement
  • Mentor junior engineers and provide technical leadership across global teams, setting standards for operational excellence and engineering rigor
  • Participate in 24x5 support operations and lead major infrastructure projects from design through deployment
The Impact You Will Have
  • Enable faster product development cycles by ensuring high availability and performance of the compute infrastructure that powers Synopsys EDA tools
  • Maximize infrastructure utilization and license efficiency, directly reducing costs and improving engineering productivity across the company
  • Reduce incident response time and platform downtime through automation, monitoring, and proactive capacity management
  • Accelerate cloud adoption and modernization efforts, helping Synopsys scale compute resources dynamically to meet global engineering demand
  • Improve workload throughput and job completion times, giving engineers more iterations per day and faster feedback loops
  • Build operational resilience into the platform, ensuring that infrastructure scales reliably as the business grows
  • Mentor and elevate the technical capabilities of the global infrastructure engineering team, raising the bar for how we operate and support critical systems
What You'll Need
  • 8+ years of Linux/UNIX systems administration experience with deep expertise in performance tuning, troubleshooting, and large-scale operations
  • 5+ years of hands-on HPC or compute farm administration, including workload scheduling, resource management, and capacity planning
  • Strong expertise in IBM Spectrum LSF, Slurm, or equivalent schedulers, including policy configuration, job prioritization, and performance optimization
  • Advanced knowledge of LDAP, NFS, DNS, enterprise storage systems, and networking in the context of distributed compute environments
  • Proven experience with Python and Shell scripting for automation, monitoring, and infrastructure orchestration
  • Solid understanding of monitoring and observability tools such as Grafana, Prometheus, Elastic, or Splunk for proactive incident detection and analysis
  • Experience with EDA environments is a strong plus, as is familiarity with cloud-integrated HPC platforms on Azure or AWS, Kubernetes, Docker, Ansible, or Terraform
Who You Are
  • You take ownership of problems end to end, from the first alert to the post-incident review, and you do not consider something fixed until you understand why it broke
  • You are self-driven and proactive, the kind of engineer who sees a pattern in the logs and writes the script to prevent it before anyone files a ticket
  • You can explain a complex infrastructure tradeoff to a VP in two sentences without losing the technical nuance, and you can turn that same conversation into a design doc for your team
  • You are collaborative and team-oriented, comfortable working across time zones and disciplines to align on solutions that work for everyone
  • You have strong analytical and problem-solving instincts, you do not guess, you instrument, measure, and validate before you change production
  • You are focused on service excellence and customer impact, you measure success by whether engineering teams can do their jobs faster and more reliably because of the infrastructure you built
The Team You'll Be Part Of

The Compute Farm Infrastructure Engineering team operates and enhances Synopsys' global HPC and compute farm platforms. You will work with a distributed team responsible for providing reliable, scalable engineering compute services that support mission-critical EDA workloads. The team drives automation, modernization, and cloud adoption while maintaining high availability and operational excellence across all global sites.

Rewards and Benefits

We offer a comprehensive range of health, wellness, and financial benefits to cater to your needs. Our total rewards include both monetary and non-monetary offerings.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Infrastructure SRE — Large-Scale Compute
Senior HPC Infrastructure SRE — Large-Scale Compute

Synopsys, Inc. • Hinoba-an

On-site
PHP 1,674,000 - 2,567,000
Lead Systems & Network Engineer | Onsite
Lead Systems & Network Engineer | Onsite

TerraBarn Inc • Cebu City

On-site
PHP 900,000 - 1,500,000
Application Engineering, Engineer
Application Engineering, Engineer

Synopsys, Inc. • Hinoba-an

On-site
PHP 600,000 - 1,000,000
Site Reliability Engineer
Site Reliability Engineer

Philtech Inc. • Taguig

On-site
PHP 670,000 - 1,339,000
Health insurance
Retirement plans
Career growth opportunities
Solutions Engineer - Sr Staff, DFT/ Design for Test
Solutions Engineer - Sr Staff, DFT/ Design for Test

Synopsys, Inc. • Hinoba-an

On-site
PHP 900,000 - 1,400,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 2,000,000 - 4,500,000
Hybrid work model
Linux System Engineer (Hybrid)
Linux System Engineer (Hybrid)

Broadridge Financial Solutions • Manila

Hybrid
PHP 800,000 - 1,200,000
Senior Devops Engineer
Senior Devops Engineer

V2 Solutions • Hinoba-an

Hybrid
PHP 1,200,000 - 2,400,000
Platform Engineer, Golang & IaC (Remote, European Union)
Platform Engineer, Golang & IaC (Remote, European Union)

Talanto • España

Hybrid
PHP 5,646,000 - 8,156,000
Equity package
Remote-first
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8 • Manila, Hinoba-an

On-site
PHP 1,000,000 - 2,000,000