LTM — a Larsen & Toubro company — is an AI‑centric global technology services company and the Business Creativity partner to the world’s largest and most disruptive enterprises. We bring human insights and intelligent systems together to help clients create greater value at the intersection of technology and domain expertise. Our capabilities span integrated operations, transformation, and business AI — enabling new ways of working, new productivity paradigms, and new roads to value. Together with over 87,000 employees across 40 countries and our global network of partners, LTM owns business outcomes for our clients, helping them not just outperform the market, but to Outcreate it.
Role
HPC Site Reliability Engineer (HPC SRE) – Fixed Term Employment, 1 year, on‑site at least 1 day per week – Gothenburg, Sweden.
Key Accountabilities
- Develop, deliver, and operate research computing services and applications.
- Adopt a Site Reliability Engineering approach to manage HPC services, handling development, deployment, monitoring, and incident response end‑to‑end.
- Solve complex technical problems related to scientific computing applications, services, and their usage by end‑users.
- Provide advanced research software engineering expertise to assist users in debugging and optimising workflows and applications.
Essential Knowledge, Skills, and Experience
- Installation, optimisation, and configuration of scientific applications.
- Effective use of HPC job schedulers such as SLURM.
- Experience working in a Linux environment.
- Competency in multiple programming and scripting languages, including Python, R, Shell Scripts, C/C++, and Golang, with deep expertise in at least one.
- Strong understanding of factors influencing HPC application performance.
- Highly customer‑focused, with the ability to explain IT technical concepts to non‑IT experts.
Required Skills and Knowledge
- Scientific degree and/or experience in the computationally intensive analysis of scientific data.
- Prior experience in high‑performance computing (HPC) environments, especially at large scales (10,000+ cores).
- Experience with high‑performance parallel filesystems at petabyte scale (e.g., GPFS, Lustre).
- Hands‑on knowledge of a range of scientific and HPC applications (e.g., simulation software, bioinformatics tools, 3D data visualisation packages).
- Experience with software build frameworks such as Easybuild or Spack.
- Expertise in GPU AI/ML tools and frameworks (e.g., CUDA, TensorFlow, PyTorch).
- Strong understanding of parallel programming techniques (e.g., MPI, pthreads, OpenMP) and code profiling/optimisation.
- Experience with workflow engines (e.g., Apache Airflow, Nextflow, Cromwell, AWS Step Functions).
- Familiarity with container runtimes such as Docker, Singularity, or Enroot.
- Expertise in scientific domains relevant to early drug development, such as deep learning, medical imaging, molecular dynamics, or omics.
- Experience with frameworks for regression tests and benchmarks for HPC applications (e.g., Reframe HPC).
- Experience working in GxP‑validated environments.
Additional Areas of Experience (Desirable)
- Experience administering and optimising an HPC job scheduler (e.g., SLURM).
- Experience with configuration automation and infrastructure as code (e.g., Ansible, HashiCorp Terraform, AWS CloudFormation, Amazon Cloud Development Kit).
- Experience deploying infrastructure and code to public cloud, especially AWS.
- Hands‑on experience working in a DevOps team and using agile methodologies.