Systems Reliability Engineer

United States Digital Space LLC

United States

Remote

USD 100,000 - 150,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking a Systems Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. This remote US-based role requires strong software engineering, automation, and reliability mindset.

The ideal candidate has 6+ years of experience, proficiency in Python/Go/Java, Linux, Kubernetes, and observability tooling, plus CI/CD expertise and excellent communication skills.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
  • Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
  • Strong programming skills in Python, Go, or Java with automation tooling.
  • Deep Linux experience at scale including networking and performance tuning.
  • Production experience with Kubernetes and container workloads.
  • Knowledge of observability tooling (Prometheus, Grafana, OpenTelemetry, ELK/EFK).
  • Experience designing and operating CI/CD pipelines for infra and apps.
  • Solid understanding of distributed systems design, consistency models, partitioning, failure semantics.
  • Experience leading incident response and post-incident reviews.
  • Excellent communication and documentation skills.

Responsibilities

  • Ensure availability, performance, and reliability of large-scale distributed systems in production.
  • Apply software engineering principles to infrastructure and operations problems.
  • Design, automate, and operate complex services to reduce toil and improve reliability.
  • Lead incident response and conduct post-incident reviews.
  • Communicate clearly and document systems and processes.

Skills

Python
Go
Java
Linux
Kubernetes
Observability
CI/CD
Distributed systems
Incident response
Communication

Education

Bachelor's degree in Computer Science or Engineering

Job description

Systems Reliability Engineer – Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: Systems Reliability EngineerLocation: 100% Remote (United States)Position Type: Full-time, Direct W2Salary Range: $100,000–$150,000 AnnuallyExperience: 6+ yearsSponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

We are seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern.

* Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.

  • Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
  • Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling.
  • Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting.
  • Production experience operating Kubernetes and container-based workloads.
  • Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents.
  • Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications.
  • Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics.
  • Demonstrated experience leading incident response and conducting effective post-incident reviews.
  • Excellent communication and documentation skills.
Preferred Qualifications

* Experience defining and operationalizing SLOs and error budgets in real production environments.

  • Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP).
  • Background in capacity planning, performance engineering, or large-scale load testing.
  • Familiarity with service mesh technologies such as Istio, Linkerd, or Consul.

Bright Vision Technologies is an Equal Opportunity Employer.

Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (the company) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

the company expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Reliability Engineer
Systems Reliability Engineer

Bright-Vision-Technologies • United States

Remote
USD 100,000 - 150,000
Reliability Engineer
Reliability Engineer

Bright Vision Technologies • Milpitas (CA)

On-site
USD 75,000 - 95,000
Reliability Monitoring Engineer
Reliability Monitoring Engineer

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior Remote SRE - Reliability, Automation & Scale
Senior Remote SRE - Reliability, Automation & Scale

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

Quantum Technologies. LLC • Atlanta (GA)

On-site
Backend Systems Engineer
Backend Systems Engineer

United States Digital Space LLC • United States

Remote
USD 60,000 - 150,000
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

Cloud Hybrid Technologies, LLC • Atlanta (GA)

On-site
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

Robotics Prcocess Automation, LLC • Atlanta (GA)

On-site
USD 100,000 - 150,000
Site Reliability Engineering (SRE) Architect
Site Reliability Engineering (SRE) Architect

Cloud Analytics Technologies, LLC • Atlanta (GA)

On-site
USD 140,000 - 210,000
Backend Solutions Architect
Backend Solutions Architect

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000