Site Reliability Engineer

CT19

Massachusetts

On-site

USD 140,000 - 210,000

Full time

22 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

CT19 is seeking a Site Reliability Engineer (SRE) to integrate and maintain hardware and software systems that enable advanced computing control platforms. You will collaborate with software, hardware, and test engineers to install, upgrade, maintain, test, and troubleshoot complex control systems for development and production environments.

The role emphasizes reliability, stability, and operational functionality across development and production environments, including CI/CD, observability,

Qualifications

  • 10+ years of experience in Network SQA, Systems Engineering, SRE, or infra roles.
  • Strong Linux and Windows administration skills.
  • Proficient in Python, Bash, or Go scripting and DevOps tools.

Responsibilities

  • Implement, maintain, and test software and hardware within heterogeneous control systems.
  • Define and test operational procedures for advanced computing platforms.
  • Manage test infrastructure, including HIL setups and containerized services like Kubernetes.
  • Automate provisioning, configuration, and orchestration of compute systems.
  • Collaborate with software and test teams to deploy DevOps tools with hardware workflows.
  • Maintain dashboards and infrastructure for regression and system health monitoring.
  • Enforce access control, system configuration, and lab operations best practices.
  • Support incident response and root cause analysis for CI/CD failures.

Skills

Linux administration
Networking
Scripting (Python, Bash, Go)
CI/CD
Observability
Rack-mounted servers

Education

Bachelor’s degree in Computer Science or related field

Tools

Docker
Git
Kubernetes
Terraform
Ansible

Job description

A Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable advanced computing control systems and software platforms. You will work closely with software engineers, scientists, hardware engineers, and test engineers to install, upgrade, maintain, test, and troubleshoot multiple hardware and software control systems for complex computing platforms. This role is a foundational systems engineering position responsible for ensuring the reliability, stability, and operational functionality of both development and production environments.

Key Responsibilities:
  • Implement, maintain, and test software and hardware within heterogeneous systems that control diverse computing devices.
  • Define, document, implement, and test operational procedures for advanced computing platforms.
  • Manage research, development, and testing infrastructure, including hardware-in-the-loop (HIL) setups, containerized services (e.g., Kubernetes), networking equipment (routers, switches), and test systems.
  • Propose and/or implement automated provisioning, configuration, and orchestration of local and remote compute systems used for control, testing, and simulation.
  • Collaborate with software and test engineering teams to ensure smooth integration and deployment of DevOps tools with custom hardware and specialized workflows.
  • Maintain artifact repositories, test result dashboards, and infrastructure for regression tracking and system health monitoring.
  • Establish and enforce best practices for access control, system configuration, and laboratory operations.
  • Support incident response, troubleshooting, and root cause analysis for CI/CD failures or system anomalies.
  • Implement monitoring and alerting automation by integrating logs and metrics from embedded systems, test environments, and orchestration layers.
  • Work with CI/CD pipelines for building, testing, and deploying software and firmware across control systems.
Required Qualifications:
  • Practical experience implementing LAN and WAN technologies on switches and routers (including VLAN configuration, DNS, DHCP, TCP/IP-based services).
  • Practical Linux (e.g., Ubuntu, Debian, Red Hat) and Windows administration experience, including network operations, application installation, and debugging.
  • Proficiency in scripting languages such as Python, Bash, or Go, and familiarity with tools like Docker, Git, and Kubernetes.
  • Experience with CI/CD tools (e.g., GitLab CI, Jenkins) and infrastructure-as-code platforms (e.g., Ansible, Terraform, or similar).
  • Familiarity with observability tools (e.g., Grafana, Prometheus, ELK stack) and logging systems for real-time monitoring.
  • Hands‑on experience with rack‑mounted servers.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 10+ years of experience in Network SQA, Systems Engineering, SRE, or infrastructure engineering roles.
Preferred Qualifications:
  • Experience supporting hybrid systems involving embedded devices, custom hardware, or real‑time control systems.
  • Knowledge of secure software deployment and product lifecycle processes.
  • Experience with hardware‑in‑the‑loop pipelines or distributed lab/test automation environments.
  • Exposure to scientific computing, high‑performance computing (HPC), or advanced computing software stacks.
  • Ability to thrive in fast‑paced, interdisciplinary R&D teams.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Veritas Search Group • Tustin (CA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • New Jersey

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE) – Evening Shift
Site Reliability Engineer (SRE) – Evening Shift

Peraton • Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer (Secret Clearance)
Site Reliability Engineer (Secret Clearance)

ROI Services LLC • Huntsville (AL)

On-site
USD 110,000 - 150,000