Senior Specialist IT

Infineon Technologies

Bengaluru

On-site

INR 2,500,000 - 4,500,000

Full time

9 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Infineon Technologies in Bengaluru seeks an HPC System Reliability Engineer to manage large-scale HPC infrastructure, optimize performance, and drive DevOps practices using Linux, automation, and modern tooling.

The role emphasizes monitoring, proactive automation, and collaboration with cross‑functional teams to ensure reliable, scalable HPC operations and continuous improvement across GPUs, InfiniBand, and accelerators.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field (or equivalent experience).
  • Proven experience managing and optimizing HPC clusters, compute nodes, storage, and interconnects.
  • Hands‑on expertise with workload managers and job schedulers, particularly LSF.
  • Strong scripting/programming skills in Python, Bash, or Go.
  • Solid Linux administration experience (RHEL, CentOS, Ubuntu) and networking knowledge.
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Proficiency with monitoring and observability tools like Prometheus, Grafana, and Nagios.
  • Experience with log management and analysis tools such as the ELK Stack.
  • Strong troubleshooting, analytical, and problem‑solving capabilities.
  • Excellent communication and collaboration skills with cross‑functional teams.
  • Ability to prioritize and manage multiple tasks in a fast‑paced environment.
  • Good understanding of DevOps practices, including CI/CD pipelines.
  • Knowledge of security best practices for HPC environments.
  • Relevant certifications such as RHCE (Red Hat Certified Engineer) are an added advantage.

Responsibilities

  • Ensure reliability, availability, and performance of HPC systems.
  • Monitor and resolve performance bottlenecks across clusters, storage, and interconnects.
  • Develop proactive monitoring, alerting, and automation solutions.
  • Automate infrastructure deployment and management using IaC tools like Terraform and Ansible.
  • Implement self‑healing and automated recovery mechanisms.
  • Manage incidents, perform root cause analysis, and drive preventive actions.
  • Create and maintain operational runbooks and playbooks.
  • Optimize HPC workloads, LSF scheduling, and resource utilization.
  • Benchmark and improve hardware and software performance.
  • Collaborate with development, operations, and vendor teams for efficient HPC operations.
  • Provide technical guidance, training, and documentation to stakeholders.
  • Stay updated on HPC technologies, including GPUs, accelerators, and InfiniBand.
  • Drive continuous improvement and scalability initiatives across the HPC environment.

Skills

Linux scripting
Python
Bash scripting
Go programming
DevOps practices
CI/CD pipelines
Networking knowledge
Troubleshooting
Communication

Education

Bachelor's or Master's in CS/Engineering

Tools

Terraform
Ansible
Docker
Kubernetes
Prometheus
Grafana
Nagios
ELK Stack
LSF

Job description

We are looking for a skilled HPC System Reliability Engineer with expertise in managing and optimizing large-scale HPC infrastructure. If you have a passion for Linux, automation, performance tuning, and DevOps practices, this could be the perfect opportunity for you.

Your role
Key responsibilities in your new role
  • Ensure reliability, availability, and performance of HPC systems.
  • Monitor and resolve performance bottlenecks across clusters, storage, and interconnects.
  • Develop proactive monitoring, alerting, and automation solutions.
  • Automate infrastructure deployment and management using IaC tools like Terraform and Ansible.
  • Implement self‑healing and automated recovery mechanisms.
  • Manage incidents, perform root cause analysis, and drive preventive actions.
  • Create and maintain operational runbooks and playbooks.
  • Optimize HPC workloads, LSF scheduling, and resource utilization.
  • Benchmark and improve hardware and software performance.
  • Collaborate with development, operations, and vendor teams for efficient HPC operations.
  • Provide technical guidance, training, and documentation to stakeholders.
  • Stay updated on HPC technologies, including GPUs, accelerators, and InfiniBand.
  • Drive continuous improvement and scalability initiatives across the HPC environment.
Your profile
Qualifications And Skills To Help You Succeed
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field (or equivalent experience).
  • Proven experience managing and optimizing HPC clusters, compute nodes, storage, and interconnects.
  • Hands‑on expertise with workload managers and job schedulers, particularly LSF.
  • Strong scripting/programming skills in Python, Bash, or Go.
  • Solid Linux administration experience (RHEL, CentOS, Ubuntu) and networking knowledge.
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Proficiency with monitoring and observability tools like Prometheus, Grafana, and Nagios.
  • Experience with log management and analysis tools such as the ELK Stack.
  • Strong troubleshooting, analytical, and problem‑solving capabilities.
  • Excellent communication and collaboration skills with cross‑functional teams.
  • Ability to prioritize and manage multiple tasks in a fast‑paced environment.
  • Good understanding of DevOps practices, including CI/CD pipelines.
  • Knowledge of security best practices for HPC environments.
  • Relevant certifications such as RHCE (Red Hat Certified Engineer) are an added advantage.
  • Provide your feedback on BizChat

#WeAreIn for driving decarbonization and digitalization.

As a global leader in semiconductor solutions in power systems and IoT, Infineon enables game-changing solutions for green and efficient energy, clean and safe mobility, as well as smart and secure IoT. Together, we drive innovation and customer success, while caring for our people and empowering them to reach ambitious goals. Be a part of making life easier, safer and greener.

Are you in?

We are on a journey to create the best Infineon for everyone.

We strive to create the best Infineon for everyone by embracing diversity and inclusion. We welcome applicants for who they are and offer a workplace built on trust, openness, respect and equal opportunity.

Our hiring decisions are based on skills and experience.

Anand Rawal

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Specialist IT
Senior Specialist IT

Infineon Technologies AG • Bengaluru

On-site
INR 1,500,000 - 2,300,000
Specialist IT
Specialist IT

Infineon Technologies AG • Bengaluru

On-site
INR 600,000 - 1,200,000
Specialist IT
Specialist IT

Infineon Technologies • Bengaluru

On-site
INR 600,000 - 1,200,000
Junior Software Developer (Contract)
Junior Software Developer (Contract)

Infineon Technologies AG • Bengaluru

On-site
INR 800,000 - 1,200,000
Third-party payroll contract
Benefits via partner company
Principal Engineer Software
Principal Engineer Software

Infineon Technologies AG • Ahmedabad District

On-site
INR 2,500,000 - 4,500,000
Staff Specialist Talent Attraction
Staff Specialist Talent Attraction

Infineon Technologies • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Devops Architect
Devops Architect

Infineon Technologies AG • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Senior Engineer
Senior Engineer

Infineon Technologies AG • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Staff Engineer Maintenance Test
Staff Engineer Maintenance Test

Infineon Technologies AG • Malacca

On-site
INR 1,800,000 - 3,200,000
Senior Specialist SPM
Senior Specialist SPM

Infineon Technologies AG • Ahmedabad District

On-site
INR 900,000 - 1,300,000