Remote Site Reliability Engineer: Scale & Resilience

Bright Vision Technologies

Nashua (NH)

On-site

USD 100,000 - 180,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking an experienced Site Reliability Engineer to ensure the availability, performance, and reliability of large-scale distributed systems in production. This fully remote role works across the US with strong emphasis on automation and software engineering principles.

You will define SLOs/SLIs, lead incident response, build robust monitoring and CI/CD pipelines, and collaborate with development teams to embed reliability early in design, while continually reducing

Qualifications

  • Ten or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
  • Bachelor’s degree in Computer Science, Engineering, or related technical discipline.
  • Strong programming skills in Python, Go, or Java with ability to build robust automation and tooling.
  • Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting.
  • Production experience operating Kubernetes and container-based workloads.
  • Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or equivalents.
  • Hands-on experience designing and operating CI/CD pipelines for infrastructure and applications.
  • Solid understanding of distributed systems, consistency, partitioning, and failure semantics.
  • Demonstrated experience leading incident response and post-incident reviews.
  • Excellent communication and documentation skills.

Responsibilities

  • Define, instrument, and refine SLOs, SLIs, and error budgets for critical services and drive prioritization.
  • Lead incident response and post-incident reviews to drive lasting improvements.
  • Design and implement robust monitoring, logging, and tracing strategies using modern tooling.
  • Develop on-call processes, runbooks, and escalation paths to reduce toil and protect engineers.
  • Automate operational toil with production-grade tooling in Python/Go/Bash and similar languages.
  • Architect and operate large-scale Kubernetes clusters and container workloads.
  • Design CI/CD pipelines for safe, frequent, observable releases with automated tests and feature flags.
  • Lead capacity planning and performance engineering with modeling and load testing.
  • Collaborate with development teams to embed reliability early in design.

Skills

SRE practices
Automation
Monitoring & observability
Programming (Python/Go/Java)
Incident management
Communication

Education

Bachelor’s degree in CS/Engineering

Tools

Kubernetes
Prometheus
Grafana
OpenTelemetry
ELK/EFK
Datadog
CI/CD tooling
Linux

Job description

Bright Vision Technologies is seeking an experienced Site Reliability Engineer to ensure the availability, performance, and reliability of large-scale distributed systems in production. This fully remote role works across the US with strong emphasis on automation and software engineering principles.

You will define SLOs/SLIs, lead incident response, build robust monitoring and CI/CD pipelines, and collaborate with development teams to embed reliability early in design, while continually reducing

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Remote
Senior Site Reliability Engineer - Remote

Bright-Vision-Technologies • United States

Remote
USD 100,000 - 150,000
Remote SRE — Scale, Resilience & Observability
Remote SRE — Scale, Resilience & Observability

Bright Vision Technologies • United States

On-site
USD 100,000 - 150,000
Competitive base salary
Health benefits
Long-term stability
Senior Reliability Engineer - 100% Remote
Senior Reliability Engineer - 100% Remote

Bright Vision Technologies • Redwood City (CA), San Mateo (CA)

On-site
USD 75,000 - 95,000
Remote SRE Technical Lead — Automate & Scale Reliability
Remote SRE Technical Lead — Automate & Scale Reliability

Bright Vision Technologies • New Albany (IN), City of Albany (NY)

On-site
USD 100,000 - 150,000
Remote Site Reliability Engineer — Automation & Cloud
Remote Site Reliability Engineer — Automation & Cloud

Jobot • Akron (OH)

Remote
USD 100,000 - 150,000
Remote DevOps & SRE Engineer — Scale & Reliability
Remote DevOps & SRE Engineer — Scale & Reliability

Bright Vision Technologies • Chapel Hill (NC)

On-site
USD 100,000 - 150,000
Remote Site Reliability Engineer - Automate & Scale
Remote Site Reliability Engineer - Automate & Scale

Jobot • Little Rock (AR)

Remote
USD 100,000 - 120,000
Remote Senior Site Reliability Engineer-Scale & Resilience
Remote Senior Site Reliability Engineer-Scale & Resilience

Far Coder • Northern (KY)

Hybrid
USD 25,000 - 40,000
Remote-Optional Senior Site Reliability Engineer
Remote-Optional Senior Site Reliability Engineer

Multi Media, LLC • United States

On-site
USD 169,000 - 215,000
Fully Remote Optional
Health, Vision, Dental, Life Insurance
Unlimited PTO
+4
Remote Senior Site Reliability Engineer — Reliability Lead
Remote Senior Site Reliability Engineer — Reliability Lead

Priority Technology Holdings, Inc. • Alpharetta (GA)

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
Medical, dental, and vision coverage
+1