Site Reliability Engineer

BridgeSource Utilities Solutions

United States

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BridgeSource Utilities Solutions seeks three Site Reliability Engineer positions (SRE, Senior SRE, Principal SRE). Roles emphasize building, operating, and improving large-scale AI, cloud, and infrastructure platforms with focus on reliability and automation.

Ideal candidates have strong Linux, Kubernetes, cloud experience (AWS/Azure/GCP), IaC (Terraform, Ansible), and SRE/Platform engineering background. Senior/Principal candidates should show leadership and mentoring abilities.

Qualifications

  • Experience in Site Reliability Engineering, Platform/Systems/Software/Infrastructure Engineering.
  • Strong knowledge of Linux, Kubernetes & cloud platforms (AWS, Azure, GCP).
  • Proficiency in automation programming (Python, Go) for tooling.
  • Experience with Infrastructure as Code (Terraform, Ansible), CI/CD, monitoring/observability.

Responsibilities

  • Help build, operate, and continuously improve large-scale AI, cloud and infrastructure platforms.
  • Focus on reliability, observability, incident response and on-call processes.
  • Collaborate across engineering teams and mentor engineers in senior/principal roles.

Skills

SRE
Platform Engineering
Systems Engineering
Software Engineering
Infrastructure Engineering

Tools

Terraform
Ansible

Job description

Site Reliability Engineer / Senior Site Reliability Engineer / Principal Site Reliability Engineer (Remote with possibility of on-site needed as programs and responsibilities continue to grow)

We are hiring for three Site Reliability Engineering positions (Site Reliability Engineer, Senior Site Reliability Engineer, and Principal Site Reliability Engineer). The level will be determined based on your experience, technical expertise, and leadership background.

In these roles, you will help build, operate, and continuously improve large-scale AI, cloud and infrastructure platforms supporting mission-critical workloads. You'll focus on reliability, automation, observability, incident response, and operational excellence across distributed production environments.

Key Requirements:

  • Experience in Site Reliability Engineering, Platform Engineering, Systems Engineering, Software Engineering, or Infrastructure Engineering.
  • Strong knowledge of Linux, Kubernetes, cloud platforms (AWS, Azure, or GCP), networking, and distributed systems.
  • Proficiency in Python, Go, or a similar programming language for automation and tooling.
  • Experience with Infrastructure as Code (Terraform, Ansible, etc.), CI/CD, monitoring, and observability tools.
  • Hands-on experience supporting production environments, troubleshooting complex issues, and performing root cause analysis (RCA).
  • Understanding of SLIs, SLOs, incident management, and on-call best practices.
  • Experience with AI infrastructure, GPU environments, HPC, InfiniBand, or RDMA is a plus.
  • Strong communication skills with the ability to collaborate across engineering teams. Senior and Principal-level candidates should also demonstrate technical leadership, architecture/design experience, and the ability to mentor and guide other engineers.
  • Nice to have: experience building HPC clusters for 1,000+ GPUs to power AI/ML workloads with Hyperscale clients

Preference will be given for Candidates living in the Seattle, New York, San Francisco and Houston Metro areas.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Lead Site Reliability Engineer - Infrastructure & DevOps
Lead Site Reliability Engineer - Infrastructure & DevOps

SRI Tech Solutions Inc. • Orlando (FL)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

asobbi • California (MO)

On-site
USD 170,000 - 220,000
Fully remote (US timezone)
Lead Site Reliability Engineer
Lead Site Reliability Engineer

KellyMitchell Group • United States

On-site
Medical, Dental, & Vision Insurance
ESOP
401K
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • San Francisco (CA)

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Washington

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000