HPC and IT Infrastructure Engineer

Nostrum Biodiscovery

Barcelona

Híbrido

EUR 70.000 - 100.000

Jornada completa

Hace 9 días
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca la empresa.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Health care insurance
Flexible working hours
Annual bonus
Teleworking policy
Transport/Meal tickets

Descripción de la vacante

Nostrum Biodiscovery seeks an HPC and IT Infrastructure Engineer to manage the computing environment for science runs in Barcelona. You will run on-prem HPC and AWS ParallelCluster, secure endpoints, and maintain scientific software stacks. Expect ~60% HPC infra work and 40% IT services and security operations.

Target start date: immediately. Competitive salary with annual bonus, flexible hours, health insurance, and transport or meal tickets are offered.

Formación

  • Solid Linux systems administration experience, including networking, storage, and OS troubleshooting.
  • Hands-on experience administering HPC clusters with Slurm: partitions, QOS, accounting, and job troubleshooting.
  • Experience building or maintaining scientific software environments: module systems, containers, compilers, and MPI.
  • Experience with AWS, particularly compute, storage, networking services, and cost awareness in cloud environments.
  • Scripting and automation (Bash, Python) and infrastructure as code.
  • Experience managing endpoints and user accounts across mixed operating systems, including IAM.
  • Working knowledge of information security practice: access control, endpoint protection, vulnerability management, and incident response.

Responsabilidades

  • Administer the on-premises HPC cluster and AWS ParallelCluster environments, including scheduler, queues, accounting, and capacity.
  • Maintain the scientific software stack (modules, containers, compilers, MPI, CUDA) and troubleshoot performance and reliability with researchers.
  • Support clusters dedicated to client projects, including maintenance and documentation.
  • Operate the office network, servers, storage, and backups.
  • Administer internally hosted services and applications, including intranet and ticketing system.
  • Provide first- and second-line user support through internal ticketing system, onboarding/offboarding, credentials and access control.
  • Manage endpoint security, alert response, and vulnerability remediation, and maintain ISO/IEC 27001 controls.
  • Automate recurring operational work and keep configuration standards documented.

Conocimientos

Linux administration
HPC cluster management
Slurm expertise
AWS know-how
Python/Bash scripting

Herramientas

Slurm
AWS
CUDA toolchains
MPI

Descripción del empleo

We are seeking an HPC and IT Infrastructure Engineer to manage the computing environment our science runs on. In this role, you will be in charge of the clusters where our simulations execute, the network and devices our team works from, and the security controls that protect both.

You will work directly with computational chemists and software engineers whose results depend on the cluster being fast and available, and you will have a real say in how the environment is designed. Expect roughly 60% of your time on HPC and infrastructure engineering and 40% on IT services and security operations.

Target start date: immediately.

Job Responsibilities
  • Administer the on-premises HPC cluster and AWS ParallelCluster environments, including scheduler, queues, accounting, and capacity, with an eye on cloud spend.
  • Maintain the scientific software stack (modules, containers, compilers, MPI, CUDA) and troubleshoot performance and reliability issues alongside the researchers running the jobs.
  • Support clusters dedicated to client projects, including maintenance and documentation.
  • Operate the office network, servers, storage, and backups.
  • Administer internally hosted services and applications, including the intranet and the ticketing system.
  • Provide first- and second-line user support through our internal ticketing system, covering user onboarding and offboarding, credentials and access control, and the management of laptops, workstations, and company devices.
  • Manage endpoint security, alert response, and vulnerability remediation, and maintain the technical controls required by ISO/IEC 27001.
  • Automate recurring operational work and keep configuration standards documented.
Required Skills
  • Solid Linux systems administration experience, including networking, storage, and troubleshooting at the OS level.
  • Hands-on experience administering HPC clusters with Slurm: partitions, QOS, accounting, and job troubleshooting.
  • Experience building or maintaining scientific software environments: module systems, containers, compilers, and MPI.
  • Experience with AWS, particularly compute, storage, and networking services, and cost awareness in cloud environments.
  • Scripting and automation (Bash, Python) and configuration management or infrastructure as code.
  • Experience managing endpoints and user accounts across mixed operating systems, including identity and access management.
  • Working knowledge of information security practice: access control, endpoint protection, vulnerability management, and incident response.
  • Ability to document your work clearly and to communicate with non-specialist users.
We will value experience in
  • AWS Parallel Cluster or other cloud HPC deployments.
  • GPU infrastructure: CUDA toolchains, drivers, and scheduling GPU workloads.
  • Compiling and installing open-source or scientific software, resolving dependency issues, and delivering functional builds to end users.
  • ISO/IEC 27001 implementation or audit support and GDPR technical controls.
  • Security monitoring and SIEM tooling, and infrastructure monitoring.
  • Working in a client-facing or regulated environment where infrastructure is part of a service commitment.
Benefits of working at Nostrum
  • Competitive salary based on experience and Barcelona market benchmarks, plus an annual bonus tied to company and department/individual performance
  • Flexible working hours and teleworking policy.
  • Health care insurance.
  • Possibility to have food, transportation, or nursery tickets (tax benefits).
  • As part of the career plan and development plans, Nostrum will facilitate all necessary training and future certifications to bring this role to the next level.
  • Exposure to international projects and teams, with offices in Barcelona and Boston and a growing footprint in Asia
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

HPC & IT Infrastructure Engineer — Flexible Hours, Remote-Ready
HPC & IT Infrastructure Engineer — Flexible Hours, Remote-Ready

Nostrum Biodiscovery • Barcelona

Híbrido
EUR 70.000 - 100.000
Health care insurance
Flexible working hours
Annual bonus
+2
HPC Architect
HPC Architect

Quantori • España

Híbrido
EUR 90.000 - 130.000
Remote or office work
Healthcare benefits
Professional development
HPC Administrator
HPC Administrator

Atos SE • Madrid

Presencial
EUR 42.000 - 65.000
Senior Site Reliability Engineer (Remote)
Senior Site Reliability Engineer (Remote)

Pragmatike • Madrid

Presencial
EUR 90.000 - 130.000
100% remote
Flexible hours
Autonomy
+2
HPC scientific software engineer
HPC scientific software engineer

Barcelona Supercomputing Center • Barcelona

Presencial
EUR 40.000 - 80.000
Flexible working hours
Private health insurance
Restaurant tickets
+4
Senior Site Reliability Engineer / Kubernetes (Remote)
Senior Site Reliability Engineer / Kubernetes (Remote)

Pragmatike • Madrid

Presencial
EUR 90.000 - 140.000
100% remote
Flexible hours
Collaborative international team
+1
Senior Scientific Computing Engineer
Senior Scientific Computing Engineer

United States Digital Space LLC • Madrid

Presencial
EUR 60.000 - 90.000
Senior HPC Cluster Engineer
Senior HPC Cluster Engineer

Nebius Group • Madrid

Presencial
EUR 80.000 - 120.000
Competitive compensation
Career growth
Flexible work
+3
Hpc Engineer For Earth Sciences Applications
Hpc Engineer For Earth Sciences Applications

Barcelona Supercomputing Center • Barcelona

Presencial
EUR 55.000 - 75.000
Restaurant tickets
Private health insurance
relocation support
Senior Cloud Engineer
Senior Cloud Engineer

Migx • Barcelona

Híbrido
EUR 80.000 - 110.000
Hybrid work model
EST business hours
25 holiday days
+5