Senior Consultant - System Management

LTM

Göteborgs kommun

On-site

SEK 60,000 - 80,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

LTM, a Larsen & Toubro company, is looking for a HPC Site Reliability Engineer in Gothenburg, Sweden. This fixed-term role (1 year) involves developing and managing research computing services with a focus on high-performance computing. Candidates must have a scientific degree and extensive experience in HPC environments, including job schedulers and programming languages. The position is on-site at least 1 day per week and requires strong customer service skills. Opportunities for optimization and advanced technical support are key aspects of this role.

Qualifications

  • Significant experience in high-performance computing environments, especially at large scales (10,000+ cores).
  • Expertise in GPU AI/ML tools and frameworks such as CUDA and TensorFlow.
  • Hands-on knowledge of scientific and HPC applications like bioinformatics tools and 3D visualization packages.

Responsibilities

  • Develop, deliver, and operate research computing services.
  • Adopt a Site Reliability Engineering approach to manage HPC services.
  • Provide advanced software engineering expertise to assist users.

Skills

Installation and optimization of scientific applications
HPC job schedulers such as SLURM
Linux environment experience
Python, R, Shell Scripts, C/C++, Golang
Understanding of HPC application performance
Customer-focused communication of technical concepts

Education

Scientific degree
Experience in high-performance computing

Tools

GPFS
Lustre
CUDA
TensorFlow
PyTorch
Apache Airflow
Docker

Job description

LTM — a Larsen & Toubro company — is an AI‑centric global technology services company and the Business Creativity partner to the world’s largest and most disruptive enterprises. We bring human insights and intelligent systems together to help clients create greater value at the intersection of technology and domain expertise. Our capabilities span integrated operations, transformation, and business AI — enabling new ways of working, new productivity paradigms, and new roads to value. Together with over 87,000 employees across 40 countries and our global network of partners, LTM owns business outcomes for our clients, helping them not just outperform the market, but to Outcreate it.

Role

HPC Site Reliability Engineer (HPC SRE) – Fixed Term Employment, 1 year, on‑site at least 1 day per week – Gothenburg, Sweden.

Key Accountabilities
  • Develop, deliver, and operate research computing services and applications.
  • Adopt a Site Reliability Engineering approach to manage HPC services, handling development, deployment, monitoring, and incident response end‑to‑end.
  • Solve complex technical problems related to scientific computing applications, services, and their usage by end‑users.
  • Provide advanced research software engineering expertise to assist users in debugging and optimising workflows and applications.
Essential Knowledge, Skills, and Experience
  • Installation, optimisation, and configuration of scientific applications.
  • Effective use of HPC job schedulers such as SLURM.
  • Experience working in a Linux environment.
  • Competency in multiple programming and scripting languages, including Python, R, Shell Scripts, C/C++, and Golang, with deep expertise in at least one.
  • Strong understanding of factors influencing HPC application performance.
  • Highly customer‑focused, with the ability to explain IT technical concepts to non‑IT experts.
Required Skills and Knowledge
  • Scientific degree and/or experience in the computationally intensive analysis of scientific data.
  • Prior experience in high‑performance computing (HPC) environments, especially at large scales (10,000+ cores).
  • Experience with high‑performance parallel filesystems at petabyte scale (e.g., GPFS, Lustre).
  • Hands‑on knowledge of a range of scientific and HPC applications (e.g., simulation software, bioinformatics tools, 3D data visualisation packages).
  • Experience with software build frameworks such as Easybuild or Spack.
  • Expertise in GPU AI/ML tools and frameworks (e.g., CUDA, TensorFlow, PyTorch).
  • Strong understanding of parallel programming techniques (e.g., MPI, pthreads, OpenMP) and code profiling/optimisation.
  • Experience with workflow engines (e.g., Apache Airflow, Nextflow, Cromwell, AWS Step Functions).
  • Familiarity with container runtimes such as Docker, Singularity, or Enroot.
  • Expertise in scientific domains relevant to early drug development, such as deep learning, medical imaging, molecular dynamics, or omics.
  • Experience with frameworks for regression tests and benchmarks for HPC applications (e.g., Reframe HPC).
  • Experience working in GxP‑validated environments.
Additional Areas of Experience (Desirable)
  • Experience administering and optimising an HPC job scheduler (e.g., SLURM).
  • Experience with configuration automation and infrastructure as code (e.g., Ansible, HashiCorp Terraform, AWS CloudFormation, Amazon Cloud Development Kit).
  • Experience deploying infrastructure and code to public cloud, especially AWS.
  • Hands‑on experience working in a DevOps team and using agile methodologies.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Site Reliability Engineer for Research Compute
HPC Site Reliability Engineer for Research Compute

LTM • Göteborgs kommun

On-site
SEK 60,000 - 80,000
Systems Engineer for HPC
Systems Engineer for HPC

Lunds Universitet • Lunds kommun

Hybrid
SEK 480,000 - 720,000
Flexible working hours
Generous vacation benefits
Favourable pension schemes
Technical Lead Platform Engineering – Scientific Computing Platform (SCP)
Technical Lead Platform Engineering – Scientific Computing Platform (SCP)

AstraZeneca AB • Göteborgs kommun

On-site
SEK 713,000 - 988,000
HPC Field Support Engineer
HPC Field Support Engineer

Atos SE • Linköpings kommun

On-site
SEK 40,000 - 55,000
Hands-on experience with cutting-edge HPC systems
Structured technical training program
Opportunities for professional growth
Onsite HPC & AI Systems Engineer (Linköping)
Onsite HPC & AI Systems Engineer (Linköping)

Atos • Linköpings kommun

On-site
Linux System Administrator
Linux System Administrator

emagine • Stockholms kommun

On-site
SEK 550,000 - 750,000
Network Infrastructure Engineer
Network Infrastructure Engineer

LTM • Stockholms kommun

On-site
SEK 500,000 - 700,000
Remote-friendly Systems Engineer for HPC & Research Infra
Remote-friendly Systems Engineer for HPC & Research Infra

Lunds Universitet • Lunds kommun

Hybrid
SEK 480,000 - 720,000
Flexible working hours
Generous vacation benefits
Favourable pension schemes
High Performance Computing System administrator
High Performance Computing System administrator

Swediumglobal • Västerås kommun

On-site
SEK 50,000 - 70,000
DevOps / SRE Specialist IRC303673
DevOps / SRE Specialist IRC303673

GlobalLogic • Göteborgs kommun

On-site
SEK 900,000 - 1,200,000