HPC Observability Engineer

EIT Professionals Corp

United States

Remote

USD 100,000 - 130,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

EIT Professionals Corp is seeking a skilled HPC Observability Engineer to work remotely. The ideal candidate will have extensive experience in setting up Grafana dashboards and maintaining InfluxDB services. You will be responsible for developing Python scripts for data ingestion and visualizing critical metrics in a high-performance computing environment. This role requires strong analytical skills and the ability to work independently, with a focus on high-quality code and documentation.

Qualifications

  • 5-8 years of experience with Grafana, InfluxDB, and Telegraf.
  • Knowledge of Docker and Google Cloud Platform is advantageous.
  • Experience in Python and Bash scripting is a plus.

Responsibilities

  • Develop and leverage Grafana dashboards and Telegraf configurations.
  • Visualize non-native resources in Grafana.
  • Collaborate with teams to troubleshoot and optimize pipelines.

Skills

Python
Communication
Analysis
Documentation

Education

B.Tech, MS, or PhD in Computer Science

Tools

Grafana
InfluxDB
Telegraf
Docker
Git
Google Cloud Platform

Job description

22 hours ago Be among the first 25 applicants

Get AI-powered advice on this job and more exclusive features.

Direct message the job poster from EIT Professionals Corp

Role: HPC Observability Engineer (Python, HPC)

Location: Remote

Contract

Description:

The client has Grafana and InfluxDB services running on K8S in-house on-premises. Telegraf is used to ingest data from a GPU HPC cluster into InfluxDB. This engineer will help collect and visualize data for the “Terra” platform. The HPC Observability Engineer should have experience in:

  • Setting up and maintaining Grafana dashboards for HPC environments
  • Creating drill-down dashboards for servers, including metrics like memory, network, and CPU utilization
  • Exploring and utilizing out-of-the-box metrics from InfluxDB
  • Writing Python scripts for data ingestion into InfluxDB with examples
  • Developing a proof of concept with a simple Python script to monitor load
  • Ingesting Infiniband packet data
  • Monitoring LSF jobs in various states
  • Visualizing server-specific and cluster-wide metrics in Grafana
  • Optional: Integrating third-party plugins like DDN’s Lustre, Mellanox fabric, etc.
Qualifications and Skills:
  • B.Tech, MS, or PhD in Computer Science or related field
  • 5-8 years of experience with Grafana, InfluxDB, and Telegraf
  • Experience in Python and Bash scripting is a plus
  • Knowledge of Docker and Google Cloud Platform is advantageous
  • HPC operations experience is beneficial
  • Strong communication skills and ability to work independently
  • Proficiency in requirements analysis and automated testing
  • Ability to write efficient, secure, and well-documented Python code
  • Experience with Git and pipeline development
  • Awareness of modern security and development practices
Responsibilities:
  • Develop and leverage Grafana dashboards and Telegraf configurations
  • Create dashboards for server and cluster metrics
  • Develop Python scripts for data ingestion and documentation
  • Visualize non-native resources in Grafana
  • Optional: Integrate third-party plugins
  • Maintain high-quality code and documentation
  • Collaborate with teams to troubleshoot and optimize pipelines
Desired Skills:
  • Python (good to have)
  • Bash scripting (good to have)
  • Docker (must)
  • HPC operations and LSF (good to have)
  • Experience with DDN Lustre, Mellanox fabric (good to have)
  • Google Cloud Platform (good to have)
  • Knowledge of Git (must)
Seniority level:
  • Mid-Senior level
Employment type:
  • Contract
Job function:
  • Engineering and Information Technology
Industries:
  • IT Services and IT Consulting

This job is active and accepting applications.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC (High-Performance Computing) Consultant @ Remote
HPC (High-Performance Computing) Consultant @ Remote

BURGEON IT SERVICES LLC • United States

Remote
USD 120,000 - 180,000
HPC Data Center Production Engineer (Automation) - Banking & Finance
HPC Data Center Production Engineer (Automation) - Banking & Finance

Hamilton Barnes • Illinois

On-site
USD 175,000 - 235,000
Premium health, dental, and vision
Performance-based bonus
Meals and high-end office amenities
+2
HPC Systems Engineer
HPC Systems Engineer

EITR Technologies LLC • Annapolis (MD)

On-site
USD 120,000 - 180,000
Senior HPC Engineer
Senior HPC Engineer

RCH Solutions • United States

Remote
USD 120,000 - 180,000
Competitive salary + bonus
Health and wellness benefits
401(k) plan with match
+2
Senior DevOps Engineer – HPC / EDA
Senior DevOps Engineer – HPC / EDA

Whiz Global LLC • Rancho Cordova (CA)

Hybrid
USD 140,000 - 170,000
HPC Performance and Validation Engineer
HPC Performance and Validation Engineer

Addison Group • Dallas (TX)

On-site
USD 180,000 - 260,000
100% paid medical, dental, vision
401(k)
25 days PTO
+3
Lead Platform Engineer/Architect - HPC, Kubernetes
Lead Platform Engineer/Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 150,000 - 230,000
HPC Platform Engineer
HPC Platform Engineer

Addison Group • Dallas (TX)

Hybrid
USD 180,000 - 260,000
Medical, dental, and vision insurance
401(k)
25 days PTO
+3
HPC Engineer
HPC Engineer

Tata Consultancy Services • Indianapolis (IN)

On-site
USD 75,000 - 80,000
Discretionary annual incentive
Comprehensive medical coverage (M/D/V,
Family support leaves
+6
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000