Senior IT Systems Administrator (Linux & HPC)

KBR, Inc

Leatherhead

Hybrid

GBP 70,000 - 100,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

KBR, Inc. seeks a Senior IT Systems Administrator (Linux & HPC) to own and operate Linux-based enterprise and HPC platforms, including SLURM scheduling, Nvidia Base Command Manager, and CycleCloud. You will work with engineering and scientific users to diagnose issues spanning compute, storage, and networking.

The role emphasizes hands-on ownership, patching, security hardening, and on-call support to maintain high availability across complex systems.

Qualifications

  • Substantial hands-on Linux administration in complex enterprise or research environments.
  • Experience supporting HPC clusters and diagnosing across compute, storage and network layers.
  • Strong SLURM administration and workload troubleshooting.

Responsibilities

  • Administer, configure, patch, harden and upgrade enterprise Linux server platforms with emphasis on RHEL or similar.
  • Operate and support HPC clusters including SLURM queues, partitions and job scheduling.
  • Manage NVIDIA Base Command Manager and Azure CycleCloud integration with SLURM.
  • Administer NetApp storage used by Linux/HPC platforms, including permissions and performance.
  • Ensure secure configuration, backups and disaster recovery planning.

Skills

Linux administration
HPC clusters
SLURM
Bash scripting
Networking basics
Documentation

Education

Bachelor's degree in computing, engineering or related field
Certs advantageous

Tools

NVIDIA Base Command Manager
Azure CycleCloud
Bright Cluster Manager
NetApp storage
Ansible
Git
Virtualisation

Job description

Senior IT Systems Administrator (Linux & HPC)KBR is a global provider of differentiated, professional services and technologies delivered across a wide government, defense and industrial base. Drawing from its rich 100-year history and culture of innovation and mission focus, KBR creates sustainable value by combining engineering, technical and scientific expertise with its full life cycle capabilities to help our clients meet their most pressing challenges today and into the future.We deliver science, technology and engineering solutions to governments and companies around the world. KBR employs approximately 37,000 people worldwide with customers in more than 80 countries and operations in over 29 countries. KBR is proud to work with its customers across the globe to provide technology, value-added services, and long-term operations and maintenance services to ensure consistent delivery with predictable results. At KBR, We Deliver.KBR is looking for a Senior IT Systems Administrator (Linux & HPC).**The Opportunity:**The Senior IT Systems Administrator (Linux & HPC) is responsible for the administration, maintenance, security and operational reliability of Linux-based enterprise and high-performance computing platforms. The role will provide hands-on technical ownership across Linux operating systems, HPC compute, workload scheduling and associated storage and network services.A major focus will be the operation of cloud and co-located HPC platforms. The position requires strong Linux engineering experience, practical knowledge of HPC environments, including NVIDIA Base Command Manager, Azure CycleCloud and SLURM. The ability to diagnose complex cross-platform issues, and the confidence to work independently while acting as a senior technical resource for colleagues is also essential. This role does not include formal supervisory or people-management accountability**Key Responsibilities****Linux Systems Administration*** Administer, configure, patch, harden and upgrade enterprise Linux server platforms, with particular emphasis on Red Hat Enterprise Linux or comparable distributions.* Manage core Linux services including identity and access integration, SSH, DNS client configuration, time synchronisation, software repositories, filesystems, logging and scheduled services.* Automate repeatable administration using shell scripting and configuration-management or orchestration tooling.* Monitor Linux performance, availability, capacity and security, and resolve complex operating-system and application integration issues.* Maintain build standards, technical documentation, operational procedures and recovery runbooks.**HPC Platform & Workload Scheduling*** Operate and support HPC clusters spanning management, login, compute and storage components.* Administer SLURM, including queues and partitions, scheduling policies, job submission, accounting, fair-share, reservations and troubleshooting failed or poorly performing workloads.* Support NVIDIA Base Command Manager and Azure CycleCloud for cluster provisioning, node lifecycle management, monitoring and integration with SLURM.* Work with engineering and scientific users to diagnose job, compiler, library, MPI, resource-allocation and performance issues.* Plan and execute maintenance activities while protecting service availability and active workloads.**Compute, Storage & Infrastructure Integration*** Administer physical and virtual server infrastructure and support hardware lifecycle, firmware and operating-system maintenance.* Support Cisco compute infrastructure and its integration with Linux and HPC management services.* Operate and support NetApp file and data services used by Linux and HPC platforms, including provisioning, permissions, capacity, performance and availability.* Apply a working understanding of high-speed networking, IP addressing, routing, DNS, NFS, network dependencies and storage connectivity to end-to-end troubleshooting.* Collaborate with network, security, application and storage specialists, vendors and service providers to resolve cross-domain incidents.**Security, Resilience & Operational Support*** Apply secure configuration, least privilege, vulnerability remediation and audit controls to Linux and HPC systems.* Ensure backup, restore and disaster-recovery requirements are defined, implemented and regularly tested for supported platforms.* Monitor service health and capacity, respond to incidents, identify root cause and contribute to problem management.* Plan and deliver changes through established change-management processes, including risk assessment, testing, implementation and back-out planning.* Participate in an appropriate operational support or on-call arrangement where required.* Provide technical guidance, peer review and knowledge transfer to infrastructure colleagues and service-desk teams.**Qualifications, Skills and Experience****Essential Technical Skills & Experience*** Substantial hands-on experience administering Linux in a complex enterprise or research-computing environment.* Practical experience supporting HPC clusters and diagnosing issues across compute, scheduler, storage and network layers.* Strong operational knowledge of SLURM administration and workload troubleshooting.* Competence in Bash or another relevant scripting language, with a track record of automating administrative tasks.* Experience with Linux performance analysis, capacity management, patching, security hardening and vulnerability remediation.* Experience operating shared file services and troubleshooting NFS, permissions, throughput, latency and capacity issues.* Working knowledge of enterprise networking fundamentals and distributed-system dependencies.* Experience delivering controlled technical change and maintaining accurate documentation in a structured IT service-management environment.**Desirable Technical Skills & Experience*** Experience with NVIDIA Base Command Manager, Bright Cluster Manager or a comparable HPC cluster-management platform.* Experience with Cisco compute platforms and associated server-management tooling.* Experience administering NetApp storage in Linux or HPC environments.* Knowledge of InfiniBand or other high-speed, low-latency fabrics.* Experience with MPI workloads, environment modules, compilers, engineering or scientific applications, and common HPC software stacks.* Experience with Ansible, Git and infrastructure-as-code practices.* Experience with virtualisation, containers or cloud-hosted Linux workloads.* Knowledge of backup, recovery, monitoring and observability platforms used in enterprise infrastructure.**Professional Capability*** Able to work independently with minimal supervision and take technical ownership through to resolution.* Demonstrated ability to analyse complex problems, prioritise operational risk and exercise sound technical judgement.* Clear written and verbal communication, including the ability to explain technical issues to specialists and non-specialists.* Strong operational discipline, attention to detail and commitment to secure, reliable and supportable services.* Collaborative approach and willingness to coach colleagues and share knowledge.**Education & Qualifications*** Bachelor's degree in computing, engineering or a related field with relevant experience, or an equivalent combination of professional training and substantial practical experience.* Relevant Linux, HPC, Cisco, NVIDIA or NetApp certifications are advantageous but not essential.**KBR Company Information**When you become part of the KBR team, your opportunities are endless. Through collaboration with our customers, we’re defining tomorrow’s challenges, then providing the solutions and services to overcome those challenges, always maintaining our commitment to total safety and reliability.At KBR, we partner with government and industry clients to provide purposeful and comprehensive solutions with an emphasis on efficiency and safety. With a full portfolio of services, proprietary technologies and expertise, our employees are ready to handle projects and missions throughout their entire lifecycle, from planning and design to sustainability and maintenance. Whether at the bottom of the ocean or in outer space, our clients trust us to deliver the impossible on a daily basis.Working at KBR means being rewarded for your contributions. In addition to competitive benefits and professional development, our people are empowered to use all their potential, creating meaningful change for themselves and our clients. We attract the best minds in the world because our expertise thrives on creativity, resourcefulness and collaboration. That is how we supply our clients with cutting-edge solutions and services.As the needs of the world change, we’re ready to respond and guide the way forward with strategic, sustainable, and technological advancements grounded in more than a century of practical application and execution.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior IT Systems Administrator (Linux & HPC)
Senior IT Systems Administrator (Linux & HPC)

KBR, Inc. • Leatherhead

Hybrid
GBP 65,000 - 95,000
Senior Network Administrator
Senior Network Administrator

KBR, Inc. • Leatherhead

On-site
GBP 60,000 - 85,000
Information Manager - Systems Integration
Information Manager - Systems Integration

KBR, Inc • Leatherhead

Hybrid
GBP 50,000 - 70,000
Competitive benefits
Professional development opportunities
Senior Manager Information Management - 3D CAD E3D
Senior Manager Information Management - 3D CAD E3D

KBR, Inc • Leatherhead

On-site
GBP 70,000 - 100,000
Senior Linux & HPC Systems Engineer
Senior Linux & HPC Systems Engineer

KBR, Inc • Leatherhead

Hybrid
GBP 70,000 - 100,000
Senior Offshore Pipeline Engineer
Senior Offshore Pipeline Engineer

KBR, Inc • Leatherhead

Hybrid
GBP 75,000 - 110,000
Senior Linux & HPC Systems Administrator (Hybrid)
Senior Linux & HPC Systems Administrator (Hybrid)

KBR, Inc. • Leatherhead

Hybrid
GBP 65,000 - 95,000
Site Manager
Site Manager

KBR, Inc. • Leeds

On-site
GBP 106,000 - 145,000
Principal Consultant - Ammonia & Derivatives
Principal Consultant - Ammonia & Derivatives

KBR, Inc • Leatherhead

On-site
GBP 80,000 - 100,000
Competitive benefits
Professional development opportunities
Data Warehouse Developer
Data Warehouse Developer

KBR, Inc • United Kingdom

Hybrid
GBP 55,000 - 75,000