Remote SRE — Scale, Automation & Observability

Bright Vision Technologies

United States

Remote

USD 100,000 - 150,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an experienced Site Reliability Engineer to ensure availability, performance, and operational excellence of large-scale distributed systems in production.

As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil.

Qualifications

  • Bachelor’s degree in computer science, engineering or a related technical discipline.
  • 5+ years in SRE, DevOps or production engineering for large-scale distributed systems.
  • Strong programming skills in Python, Go or Java and automation tooling.
  • Deep Linux expertise including networking and performance tuning.
  • Production Kubernetes experience and container-based workloads.
  • Strong observability experience with Prometheus, Grafana, OpenTelemetry, ELK/EFK or equivalents.
  • Experience designing and operating CI/CD pipelines for infra and apps.
  • Understanding of distributed systems concepts, including consistency and failure semantics.
  • Experience leading incident response and post-incident reviews.
  • Excellent communication and documentation skills.

Responsibilities

  • Define SLOs, SLIs, and error budgets for services and drive priorities.
  • Lead production incident response and post-incident reviews to improve reliability.
  • Design and implement monitoring, logging, and tracing strategies with modern tools.
  • Build on-call runbooks and escalation paths to reduce toil and protect engineers.
  • Automate operational toil with production-grade tooling in Python, Go or Bash.
  • Architect and operate large-scale Kubernetes clusters and workloads.
  • Design CI/CD pipelines for reliable releases with tests and feature flags.
  • Plan capacity and performance with models and load testing.
  • Collaborate with application teams to embed reliability in design.
  • Advance platform resilience through chaos engineering and fault injection.

Skills

Python
Go
Java
Linux
Kubernetes
Observability
CI/CD
Incident response
Communication

Education

Bachelor’s degree in Computer Science or Engineering

Tools

Prometheus
Grafana
OpenTelemetry
ELK/EFK
Datadog
Chaos Monkey
Gremlin
Litmus

Job description

Bright Vision Technologies is seeking an experienced Site Reliability Engineer to ensure availability, performance, and operational excellence of large-scale distributed systems in production.

As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote SRE: Scale & Reliability
Remote SRE: Scale & Reliability

Bright Vision Technologies • Tampa (FL), Virginia (MN)

Remote
USD 100,000 - 150,000
Senior Remote Site Reliability Engineer (SRE)
Senior Remote Site Reliability Engineer (SRE)

Bright Vision Technologies • Woodbridge Township (NJ)

Remote
USD 100,000 - 180,000
Remote DevOps & SRE Engineer — Cloud Reliability
Remote DevOps & SRE Engineer — Cloud Reliability

Bright Vision Technologies • Columbus (OH), Powell (OH), New Albany (OH), Hilliard (OH)

Remote
USD 100,000 - 150,000
Remote Senior Platform Reliability Engineer (SRE)
Remote Senior Platform Reliability Engineer (SRE)

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
Health insurance
Remote Reliability Engineer for Large-Scale Systems
Remote Reliability Engineer for Large-Scale Systems

Bright Vision Technologies • Yuba City (CA)

Remote
USD 75,000 - 95,000
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Remote DevOps & SRE Engineer for Scalable Systems
Remote DevOps & SRE Engineer for Scalable Systems

Bright Vision Technologies • United States

Remote
USD 100,000 - 150,000
SRE Technical Lead - Automation & Observability
SRE Technical Lead - Automation & Observability

Bright Vision Technologies • Chesterfield (MO)

On-site
USD 100,000 - 150,000
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: Scale Systems with Automation & Observability
Senior SRE: Scale Systems with Automation & Observability

Tata Consultancy Services • Englewood Cliffs (NJ)

On-site
USD 110,000 - 125,000