Cyberinfrastructure Engineer

uchicago

Chicago (IL)

Presencial

USD 80.000 - 90.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Convierte este puesto en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

The University of Chicago's MANIAC Lab is seeking a Cyberinfrastructure Engineer to build and operate advanced cyberinfrastructure for science collaborations such as ATLAS at CERN. You will work in campus data centers, connecting researchers to national grids and AI platforms, with hands-on Linux administration and cloud-native operations.

You'll implement Kubernetes, Helm, containers, and GitOps tooling; manage Ceph storage, monitoring, and security patches; and participate in a distributed

Formación

  • Bachelor's degree or equivalent in a related field.
  • 2–5 years of relevant work experience in IT/ops.
  • Strong Linux server administration experience.
  • Experience with distributed storage and monitoring tooling.

Responsabilidades

  • Operate Linux compute clusters and storage systems for ATLAS production.
  • Manage Ceph/Rook storage and Kubernetes deployments.
  • Perform on-site hardware maintenance and vendor support.
  • Maintain monitoring with Prometheus and Grafana; apply patches.
  • Develop AI-enabled infrastructure with container and GitOps workflows.

Conocimientos

Linux system administration
Kubernetes
HTCondor
Ceph
GPU/server hardware
Python scripting
Git/GitOps
Networking

Educación

Bachelor's degree in Computer Science, Computer Engineering, Physics or related field

Herramientas

Puppet
Chef
Ansible
Flux
Kubespray
Docker
Kubernetes
Helm
OpenShift
Rook/Ceph integration

Descripción del empleo

Department

PSD Enrico Fermi Institute: Gardner Group

About the Department

The MANIAC Lab, located within the Enrico Fermi Institute of the Physical Sciences Division at the University of Chicago, builds and operates advanced cyberinfrastructure for scientific instruments investigating the fundamental mysteries of nature. The Lab operates data-intensive high-throughput computing facilities that are part of a global computing grid used to reconstruct and analyze particle collisions recorded by the ATLAS detector at the Large Hadron Collider (LHC) at CERN in Geneva, Switzerland (https://atlas.cern). The Lab also serves as a gateway for U.S. researchers to national cyberinfrastructure resources through the Open Science Grid Consortium (OSG) and the Institute for Research and Innovation in Software for High Energy Physics (IRIS-HEP).

The Lab operates the NSF Shared Tier 3 Analysis Facility (https://af.uchicago.edu). The facility supports ATLAS physicists analyzing the complete LHC Run 2 and Run 3 datasets as the collaboration prepares for the High-Luminosity LHC (HL-LHC), which begins with Run 4 around 2030. It provides computing infrastructure and services for the South Pole Telescope (SPT-3G) and operates a worldwide distributed data network for the XENON dark matter search experiment at the Gran Sasso National Laboratory in Italy.

The Lab is increasingly focused on AI-enabled research computing. Its GPU platforms support machine learning training and inference, including models served through NVIDIA Triton and large language models hosted on facility hardware. The Lab has also built an agentic AI platform based on the Model Context Protocol (MCP). It gives researchers' AI assistants secure, identity-brokered access to facility services such as HTCondor, Kubernetes, Rucio, and ServiceX. The same platform supports AI-assisted facility operations, where sandboxed agents analyze operational metrics and propose corrective actions for human review.

The Lab provides the Scalable Systems Laboratory (SSL), a software testing and integration platform for IRIS-HEP. IRIS-HEP develops software and computing solutions for the HL-LHC era. Through the SSL Deployment Factory, the Lab is packaging the Kubernetes infrastructure it already operates into tested, versioned bundles that other facilities can deploy. These bundles are bootstrapped with Kubespray and managed through GitOps with Flux, and cover services such as JupyterHub, BinderHub, ServiceX, and the MCP platform. The goal is to shorten a process that currently takes months of manual work at each site.

Job Summary

As the Cyberinfrastructure Engineer at the MANIAC Lab within the Enrico Fermi Institute, you will report directly to Research Assistant Professor Giordon Stark. As part of our team (https://maniaclab.uchicago.edu), you will help build and operate advanced cyberinfrastructure for science collaborations such as the ATLAS experiment at the CERN LHC. These collaborations rely on data-intensive, distributed computing technologies such as HTCondor, Kubernetes, and Ceph.

The Cyberinfrastructure Engineer will help operate the storage, compute, and GPU infrastructure in the University's campus data centers. This infrastructure connects researchers to the national-scale OSG fabric and to emerging AI/ML service platforms. The work combines hands-on Linux systems administration and on-site hardware support with cloud-native operations using Kubernetes, Helm, containers, and GitOps tooling. It also involves distributed storage (Ceph via Rook, ZFS), configuration management (Puppet), and monitoring and alerting (Prometheus, Grafana).

The team is collaborative and partly distributed. We meet weekly over Zoom and coordinate daily in Slack. Team members are expected to take ownership of their work and function independently. The position is jointly supported by the ATLAS Midwest Tier 2 Center (MWT2), focused on production operations, and by IRIS-HEP, focused on building the infrastructure that enables AI agents to drive physics analysis and facility operations for the HL-LHC era.

Responsibilities
  • Facility operations (MWT2 and the UChicago Analysis Facility, approximately 50%)
  • Provides systems administrative services for Linux compute clusters (CPU and GPU), storage systems, and related support servers. These systems support ATLAS production and analysis at the Midwest Tier 2 Center and the Analysis Facility.
  • Operates, upgrades, and scales Ceph distributed storage (including Rook-managed Ceph on Kubernetes). The goal is to meet growing HL-LHC capacity and throughput requirements.
  • Performs on-site hardware work in campus data centers, including installing, retrofitting, and replacing servers, storage, and GPUs. Manages vendor support cases and spare-parts inventory.
  • Maintains monitoring, dashboards, and alerting (e.g., Prometheus, Grafana) for storage, compute, network, and hardware health.
  • Applies operating system and service security patches and vulnerability mitigations in coordination with University security requirements.
  • Performs network diagnostics, throughput measurement, and analysis for both LAN and WAN. Supports the facility's participation in WLCG Data Challenges and HL-LHC readiness scale tests.
  • Participates in the team's operations support rotation and maintains documentation and operational runbooks.
  • Agentic research infrastructure (IRIS-HEP, approximately 50%)
  • Deploys and operates the infrastructure for agent-driven physics analysis. This includes the Lab's MCP gateway, the MCP servers it brokers access to, and the credential services behind them. Together, these let researchers' AI assistants securely use facility resources such as HTCondor, Kubernetes, Rucio, and ServiceX.
  • Supports IRIS-HEP integration challenges and demonstrations of end-to-end agentic analysis workflows. The work spans dataset discovery, batch processing, analysis, and inference run through AI agents.
  • Operates GPU platforms for machine learning and AI workloads, including locally hosted large language models and sandboxed agent environments.
  • Helps develop and operate AI-assisted operations agents that monitor HTCondor, Kubernetes, Ceph, and related services. These agents propose corrective actions under human review.
  • Deploys and operates services using container-based approaches (Docker, Kubernetes, Helm) and GitOps workflows (e.g., Flux). Contributes reusable deployment bundles to the SSL Deployment Factory so other facilities can adopt these services.
  • Learns new distributed computing, AI/ML, and infrastructure-as-a-service technologies.
  • Maintains complex system and network administration functions. Works with moderate guidance to administer simple systems and assists in the administration of larger systems.
  • Installs, configures, and maintains operating system workstations and servers. Performs software installations and upgrades to operating systems and layered software packages. Monitors and tunes the system to achieve optimum performance levels, acquiring higher-level skills in the process.
  • Performs other related work as needed.
Minimum Qualifications
Education

Minimum requirements include a college or university degree in related field.

Work Experience

Minimum requirements include knowledge and skills developed through 2-5 years of work experience in a related job discipline.

Certifications

---

Preferred Qualifications
Education
  • Bachelor's degree in Computer Science, Computer Engineering, Physics or related field.
Experience
  • Strong experience managing Linux operating systems.
  • Configuration management and build systems for large numbers of computers using tools such as Puppet, Chef and Ansible.
  • Experience operating distributed storage at scale, particularly Ceph (including Rook on Kubernetes).
  • Hands-on data center hardware experience, including server diagnostics, component replacement, and working with vendor support (e.g., Dell iDRAC/warranty processes).
  • Experience with monitoring and alerting tools such as Nagios, Prometheus, Grafana, and Alertmanager.
Technical Knowledge or Skills
  • Unix/Linux operating systems administration tools and shell scripts.
  • Distributed storage systems such as Ceph; local file systems such as ZFS.
  • Git version control, scripting (Bash, Python) and automation.
  • Knowledge and expertise in technologies such as TCP/IP and related protocols; networked file systems, including NFS.
  • Knowledge or experience with batch scheduling systems such as Slurm or HTCondor.
  • Knowledge of container technologies such as Docker, Kubernetes, Helm, OpenShift/OKD, OpenStack.
  • Familiarity with identity and access management (e.g., Keycloak, OAuth/OIDC).
Preferred Competencies
  • Strong oral and written communication skills.
  • Initiative and capacity for teamwork and creativity.
  • Ability to effectively communicate and collaborate with team members, supervisors, and researchers.
  • Ability to manage complex technical details and switch between projects.
  • Ability to work independently with minimal supervision, take ownership of issues through resolution, and keep the team informed.
  • Comfortable collaborating in a distributed team through weekly Zoom meetings and daily Slack communication.
  • Knowledge or experience with Spark, Dask or Ray.
  • Knowledge of GitOps and cluster lifecycle tooling such as Flux, Argo CD, or Kubespray.
  • Experience with GPU servers, including NVIDIA drivers, CUDA, and GPU scheduling in Kubernetes or HTCondor.
  • Familiarity with AI/ML infrastructure, such as model serving, LLM-based tooling, or MCP, is a plus.
Working Conditions
  • Able to work on-site in University data centers on a regular basis to physically install, move, and replace hardware of up to 50lbs/person.
Additional Documents
  • Resume/CV (required)
  • Cover letter (required)
  • Professional Reference Information (preferred)
Job Family

Information Technology

Role Impact

Individual Contributor

Scheduled Weekly Hours

37.5

Drug Test Required

No

Health Screen Required

No

Motor Vehicle Record Inquiry Required

No

Pay Rate Type

Salary

FLSA Status

Exempt

Pay Range

$80,000.00 - $90,000.00

The included pay rate or range represents the University's good faith estimate of the possible compensation offer for this role at the time of posting.

Benefits Eligible

Yes

The University of Chicago offers a wide range of benefits programs and resources for eligible employees, including health, retirement, and paid time off. Information about the benefit offerings can be found in the Benefits Guidebook.

Posting Statement

The University of Chicago is an equal opportunity employer and does not discriminate on the basis of race, color, religion, sex, sexual orientation, gender, gender identity, or expression, national or ethnic origin, shared ancestry, age, status as an individual with a disability, military or veteran status, genetic information, or other protected classes under the law. For additional information please see the University's Notice of Nondiscrimination.

Job seekers in need of a reasonable accommodation to complete the application process should call 773-702-5800 or submit a request via Applicant Inquiry Form.

All offers of employment are contingent upon a background check that includes a review of conviction history. A conviction does not automatically preclude University employment. Rather, the University considers conviction information on a case-case basis and assesses the nature of the offense, the circumstances surrounding it, the proximity in time of the conviction, and its relevance to the position.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Cyberinfrastructure Engineer
Cyberinfrastructure Engineer

The University of Chicago • Chicago (IL)

Presencial
USD 80.000 - 90.000
Cyberinfrastructure Engineer
Cyberinfrastructure Engineer

University of Chicago • EE. UU.

Presencial
USD 80.000 - 90.000
Cyberinfrastructure Engineer
Cyberinfrastructure Engineer

The-University-of-Chicago • Chicago (IL)

Presencial
USD 80.000 - 90.000
Cloud Computing Developer
Cloud Computing Developer

uchicago • Illinois

Presencial
USD 95.000 - 110.000
Staff Platform Engineer
Staff Platform Engineer

uchicago • Illinois

Presencial
USD 100.000 - 140.000
Computational Scientist- HPC/AI Generalist
Computational Scientist- HPC/AI Generalist

uchicago • Chicago (IL)

Híbrido
USD 85.000 - 103.000
Infrastructure Platform Engineer
Infrastructure Platform Engineer

University of Chicago • Chicago (IL)

Presencial
USD 106.000 - 117.000
Sr. HPC System Administrator
Sr. HPC System Administrator

uchicago • Chicago (IL)

Híbrido
USD 100.000 - 125.000
Systems Administrator 3
Systems Administrator 3

uchicago • Illinois

Presencial
USD 100.000 - 110.000
Systems Administrator 3
Systems Administrator 3

University of Chicago • Chicago (IL)

Presencial
USD 100.000 - 110.000