Systems Reliability Engineer

Bright-Vision-Technologies

United States

Remote

USD 100,000 - 150,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking a Senior Site Reliability Engineer to ensure the availability and performance of large-scale distributed systems in production. You will apply software engineering principles to infrastructure, automating operations and reducing toil.

The ideal candidate has deep systems knowledge and strong programming skills, with a track record of designing reliable services and leading incident reviews.

Qualifications

  • Bachelor’s degree in CS, Engineering, or related field.
  • Five+ years in SRE/DevOps/production engineering for large distributed systems.
  • Proficient in Python/Go/Java with automation tooling.
  • Extensive Linux experience including networking, tuning, troubleshooting.
  • Hands-on Kubernetes and container workload management.
  • Strong observability skills with Prometheus, Grafana, OpenTelemetry, ELK/EFK.
  • Experience designing and operating CI/CD pipelines for infra and apps.
  • Understanding of distributed systems concepts and failure semantics.
  • Proven incident response leadership and post-incident reviews.
  • Excellent writing and communication abilities.

Skills

Python/Go/Java
Linux at scale
Incident response
Documentation skills
Strong communication

Education

Bachelor’s degree in CS/Engineering or related

Tools

Kubernetes
Prometheus
Grafana
OpenTelemetry
ELK/EFK
CI/CD tooling
AWS/Azure/GCP

Job description

Systems Reliability Engineer - Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: Systems Reliability Engineer

Location: 100% Remote (United States)

Position Type: Full-time, Direct W2

Salary Range: $100,000-150,000 Annually

Experience: 6+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

We are seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern.

Required Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
  • Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
  • Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling.
  • Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting.
  • Production experience operating Kubernetes and container-based workloads.
  • Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents.
  • Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications.
  • Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics.
  • Demonstrated experience leading incident response and conducting effective post-incident reviews.
  • Excellent communication and documentation skills.
Preferred Qualifications
  • Experience defining and operationalizing SLOs and error budgets in real production environments.
  • Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP).
  • Background in capacity planning, performance engineering, or large-scale load testing.
  • Familiarity with service mesh technologies such as Istio, Linkerd, or Consul.
Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Reliability Engineer
Systems Reliability Engineer

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Reliability Engineer
Reliability Engineer

Bright Vision Technologies • Milpitas (CA)

On-site
USD 75,000 - 95,000
Reliability Monitoring Engineer
Reliability Monitoring Engineer

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Backend Solutions Architect
Backend Solutions Architect

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Senior Remote SRE - Reliability, Automation & Scale
Senior Remote SRE - Reliability, Automation & Scale

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior Devops Engineer
Senior Devops Engineer

Bright Vision Technologies • Flower Mound (TX)

Remote
USD 155,000 - 175,000
Lead Backend Engineer
Lead Backend Engineer

Socket.dev • Sunnyvale (CA)

On-site
USD 100,000 - 150,000
Backend Systems Engineer
Backend Systems Engineer

United States Digital Space LLC • United States

Remote
USD 60,000 - 150,000
Senior Server-Side Engineer
Senior Server-Side Engineer

Bright Vision Technologies • Raleigh (NC)

On-site
USD 100,000 - 150,000
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

Quantum Technologies. LLC • Atlanta (GA)

On-site
USD 124,000 - 220,000