Senior Specialist IT

Infineon Technologies AG

Bengaluru

On-site

INR 1,500,000 - 2,300,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Infineon Technologies AG is seeking an HPC System Reliability Engineer to manage and optimize large-scale HPC infrastructure. You will ensure reliability, monitor performance bottlenecks, and automate deployment using Terraform and Ansible.

The role requires strong Linux, scripting, and containerization skills, plus experience with HPC schedulers like LSF. You will collaborate across development, operations, and vendor teams, implementing self-healing, proactive monitoring, and runbooks to drive

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field
  • Proven experience managing and optimizing HPC clusters, compute nodes, storage and interconnects
  • Hands-on experience with workload managers and job schedulers, especially LSF
  • Strong scripting/programming skills in Python, Bash, or Go
  • Solid Linux administration (RHEL/CentOS/Ubuntu) and networking knowledge
  • Experience with containerization (Docker, Kubernetes)
  • Proficiency with monitoring/observability tools (Prometheus, Grafana, Nagios)
  • Experience with log analysis tools (ELKStack)
  • Strong troubleshooting and communication skills
  • DevOps practices including CI/CD and security best practices for HPC

Responsibilities

  • Ensure reliability, availability, and performance of HPC systems
  • Monitor and resolve performance bottlenecks across clusters, storage, and interconnects
  • Develop proactive monitoring, alerting, and automation solutions
  • Automate infrastructure deployment using IaC tools like Terraform and Ansible
  • Implement self-healing and automated recovery mechanisms
  • Manage incidents, perform root cause analysis, and drive preventive actions
  • Create and maintain runbooks and playbooks
  • Optimize HPC workloads, LSF scheduling, and resource utilization
  • Benchmark and improve hardware and software performance
  • Collaborate with development, operations, and vendor teams
  • Provide training and documentation to stakeholders
  • Stay updated on HPC tech including GPUs, accelerators, InfiniBand
  • Drive continuous improvement and scalability initiatives across HPC environment

Skills

LSF
Python
Bash
Go
Linux
Docker
Kubernetes
Prometheus
Grafana
Nagios
ELK Stack
CI/CD
RHCE

Education

Bachelor’s or Master’s degree in Computer Science/Engineering

Tools

Terraform
Ansible
Docker
Kubernetes
Terraform
Ansible

Job description

We are looking for a skilled HPC System Reliability Engineer with expertise in managing and optimizing large-scale HPC infrastructure. If you have a passion for Linux, automation, performance tuning, and DevOps practices, this could be the perfect opportunity for you.

Your role Key responsibilities in your new role
  • Ensure reliability, availability, and performance of HPC systems.
  • Monitor and resolve performance bottlenecks across clusters, storage, and interconnects.
  • Develop proactive monitoring, alerting, and automation solutions.
  • Automate infrastructure deployment and management using IaC tools like Terraform and Ansible.
  • Implement self-healing and automated recovery mechanisms.
  • Manage incidents, perform root cause analysis, and drive preventive actions.
  • Create and maintain operational runbooks and playbooks.
  • Optimize HPC workloads, LSF scheduling, and resource utilization.
  • Benchmark and improve hardware and software performance.
  • Collaborate with development, operations, and vendor teams for efficient HPC operations.
  • Provide technical guidance, training, and documentation to stakeholders.
  • Stay updated on HPC technologies, including GPUs, accelerators, and InfiniBand.
  • Drive continuous improvement and scalability initiatives across the HPC environment.
Your profile Qualifications and skills to help you succeed
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field (or equivalent experience).
  • Proven experience managing and optimizing HPC clusters, compute nodes, storage, and interconnects.
  • Hands‑on expertise with workload managers and job schedulers, particularly LSF.
  • Strong scripting/programming skills in Python, Bash, or Go.
  • Solid Linux administration experience (RHEL, CentOS, Ubuntu) and networking knowledge.
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Proficiency with monitoring and observability tools like Prometheus, Grafana, and Nagios.
  • Experience with log management and analysis tools such as the ELKStack .
  • Strong troubleshooting, analytical, and problem-solving capabilities .
  • Excellent communication and collaboration skills with cross-functional teams .
  • Ability to prioritize and manage multiple tasks in a fast‑paced environment.
  • Good understanding of DevOps practices, including CI/CD pipelines.
  • Knowledge of security best practices for HPC environments.
  • Relevant certifications such as RHCE (RedHatCertified Engineer) are an added advantage.

Provide your feedback on BizChat #WeAreIn for driving decarbonization and digitalization.

As a global leader in semiconductor solutions in powersystems and IoT, Infineon enables game-changing solutions for green and efficient energy, clean and safe mobility, as well as smart and secure IoT.

Together

We drive innovation and customer success, while caring for our people and empowering them to reach ambitious goals.

Be a part of making life easier, safer and greener.

Are you in? Weare on a journey to create the best Infineon for everyone.

We strive to create the best Infineon for everyoneby embracing diversity and inclusion.

Wewelcome applicants for who they are and offer a workplace built on trust, openness, respect and equal opportunity.

Our hiring decisions are based on skills and experience.

Even if youdon’t meet every requirement, we encourage you to apply.

Please let your recruiter know if you need any accommodations during the interview process.

Learn more about our variouscontact channels and aboutDiversity&Inclusionat Infineon.

Anand Rawal

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Specialist IT
Senior Specialist IT

Infineon Technologies • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Specialist IT
Specialist IT

Infineon Technologies AG • Bengaluru

On-site
INR 600,000 - 1,200,000
Specialist IT
Specialist IT

Infineon Technologies • Bengaluru

On-site
INR 600,000 - 1,200,000
Junior Software Developer (Contract)
Junior Software Developer (Contract)

Infineon Technologies AG • Bengaluru

On-site
INR 800,000 - 1,200,000
Third-party payroll contract
Benefits via partner company
Devops Architect
Devops Architect

Infineon Technologies AG • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Principal Engineer Software
Principal Engineer Software

Infineon Technologies AG • Ahmedabad District

On-site
INR 2,500,000 - 4,500,000
Staff Engineer Maintenance Test
Staff Engineer Maintenance Test

Infineon Technologies AG • Malacca

On-site
INR 1,800,000 - 3,200,000
Senior Staff Engineer Digital Twin Equipment & Process
Senior Staff Engineer Digital Twin Equipment & Process

Infineon Technologies AG • Malacca

On-site
INR 2,000,000 - 2,800,000
Senior Engineer
Senior Engineer

Infineon Technologies AG • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Senior Specialist SPM
Senior Specialist SPM

Infineon Technologies AG • Ahmedabad District

On-site
INR 900,000 - 1,300,000