Senior Site Reliability Engineer for HPC & Control Systems

CT19

Massachusetts

On-site

USD 140,000 - 210,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

CT19 is seeking a Site Reliability Engineer (SRE) to integrate and maintain hardware and software systems that enable advanced computing control platforms. You will collaborate with software, hardware, and test engineers to install, upgrade, maintain, test, and troubleshoot complex control systems for development and production environments.

The role emphasizes reliability, stability, and operational functionality across development and production environments, including CI/CD, observability,

Qualifications

  • 10+ years of experience in Network SQA, Systems Engineering, SRE, or infra roles.
  • Strong Linux and Windows administration skills.
  • Proficient in Python, Bash, or Go scripting and DevOps tools.

Responsibilities

  • Implement, maintain, and test software and hardware within heterogeneous control systems.
  • Define and test operational procedures for advanced computing platforms.
  • Manage test infrastructure, including HIL setups and containerized services like Kubernetes.
  • Automate provisioning, configuration, and orchestration of compute systems.
  • Collaborate with software and test teams to deploy DevOps tools with hardware workflows.
  • Maintain dashboards and infrastructure for regression and system health monitoring.
  • Enforce access control, system configuration, and lab operations best practices.
  • Support incident response and root cause analysis for CI/CD failures.

Skills

Linux administration
Networking
Scripting (Python, Bash, Go)
CI/CD
Observability
Rack-mounted servers

Education

Bachelor’s degree in Computer Science or related field

Tools

Docker
Git
Kubernetes
Terraform
Ansible

Job description

CT19 is seeking a Site Reliability Engineer (SRE) to integrate and maintain hardware and software systems that enable advanced computing control platforms. You will collaborate with software, hardware, and test engineers to install, upgrade, maintain, test, and troubleshoot complex control systems for development and production environments.

The role emphasizes reliability, stability, and operational functionality across development and production environments, including CI/CD, observability,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

CT19 • Massachusetts

On-site
USD 140,000 - 210,000
Site Reliability Engineer: HPC Automation & Monitoring
Site Reliability Engineer: HPC Automation & Monitoring

LTD GLOBAL, LLC • Berkeley (CA)

Hybrid
USD 91,000 - 129,000
Senior Site Reliability Engineer - 24/7 HPC Ops
Senior Site Reliability Engineer - 24/7 HPC Ops

Bay Systems • Berkeley (CA)

On-site
USD 120,000 - 150,000
Senior SRE: Global HPC & Multi-Cloud Reliability
Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation • Durham (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior SRE & DevOps Engineer - Automation & Reliability
Senior SRE & DevOps Engineer - Automation & Reliability

Compunnel, Inc. • New Jersey

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer – 24/7 Uptime & Automation
Lead Site Reliability Engineer – 24/7 Uptime & Automation

HCL Technologies Limited • Santa Clara (CA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
HPC Site Reliability Engineer — Onsite 24/7
HPC Site Reliability Engineer — Onsite 24/7

Bay Systems Consulting Inc. • Berkeley (CA)

On-site
USD 83,000 - 166,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior SRE: Scalable AI Platform & HPC
Senior SRE: Scalable AI Platform & HPC

Mistral Ai • New York (NY)

On-site
USD 150,000 - 180,000