Reliability Engineer

Bright Vision Technologies

Milpitas (CA)

On-site

USD 75,000 - 95,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking an experienced Reliability Engineer to ensure the availability and performance of large-scale distributed systems in production. You will apply software engineering principles to infrastructure and operations problems, pushing the platform toward higher reliability with automation and reduced toil.

The ideal candidate has deep systems knowledge, strong programming skills, and a disciplined approach to designing, automating, and operating complex services as

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related technical discipline.
  • 6+ years of SRE/DevOps/production engineering experience for large-scale distributed systems.
  • Strong programming skills in Python, Go, or Java with robust automation tooling.
  • Deep Linux experience at scale, including networking and troubleshooting.
  • Production experience operating Kubernetes and container workloads.
  • Observability tooling knowledge (Prometheus, Grafana, OpenTelemetry, ELK/EFK).
  • Experience designing CI/CD pipelines for infrastructure and applications.
  • Understanding of distributed system design, consistency, partitioning, and failure semantics.
  • Experience leading incident response and post-incident reviews.
  • Excellent communication and documentation skills.

Responsibilities

  • Ensure the availability, performance, and operational excellence of large-scale distributed systems.
  • Apply software engineering principles to infrastructure and operations challenges.
  • Automate and operate complex services to reduce toil and improve reliability.
  • Lead incident response and conduct post-incident reviews to drive improvements.

Skills

Programming in Python/Go/Java
Linux at scale
Incident response leadership
Communication skills

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

Kubernetes
Prometheus
Grafana
OpenTelemetry
ELK/EFK
CI/CD pipelines
AWS/Azure/GCP

Job description

Reliability Engineer – Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title:

Reliability Engineer

Location:

100% Remote (U.S.)

Position Type:

Full-time, Direct W2

Salary Range:

$75,000–$95,000 Annually

Experience Required:

6+ years

Sponsorship:

U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

We are seeking an experienced Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern.

Required Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
  • Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
  • Strong programming skills in at least one of Python, Go, or Java, with the ability to build robust automation and tooling.
  • Deep, hands-on experience operating Linux at scale, including networking, performance tuning, and systems-level troubleshooting.
  • Production experience operating Kubernetes and container-based workloads.
  • Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents.
  • Hands-on experience designing and operating CI/CD pipelines for both infrastructure and applications.
  • Solid understanding of distributed system design, including consistency models, partitioning, and failure semantics.
  • Demonstrated experience leading incident response and conducting effective post-incident reviews.
  • Excellent communication and documentation skills.
Preferred Qualifications
  • Experience defining and operationalizing SLOs and error budgets in real production environments.
  • Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP).
  • Background in capacity planning, performance engineering, or large-scale load testing.
  • Familiarity with service mesh technologies such as Istio, Linkerd, or Consul.
Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Bright Vision Technologies is an Equal Opportunity Employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Reliability Engineer
Systems Reliability Engineer

Bright-Vision-Technologies • United States

Remote
USD 100,000 - 150,000
Systems Reliability Engineer
Systems Reliability Engineer

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Reliability Monitoring Engineer
Reliability Monitoring Engineer

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior Remote SRE - Reliability, Automation & Scale
Senior Remote SRE - Reliability, Automation & Scale

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Lead Backend Engineer
Lead Backend Engineer

Socket.dev • Sunnyvale (CA)

On-site
USD 100,000 - 150,000
Remote Reliability Engineer — SRE for Cloud Platforms
Remote Reliability Engineer — SRE for Cloud Platforms

Bright Vision Technologies • Milpitas (CA)

On-site
USD 75,000 - 95,000
Backend Solutions Architect
Backend Solutions Architect

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Senior Devops Engineer
Senior Devops Engineer

Bright Vision Technologies • Flower Mound (TX)

Remote
USD 155,000 - 175,000
Remote Reliability Engineer: Scalable Systems & SRE
Remote Reliability Engineer: Scalable Systems & SRE

Visa Hunt • United States

On-site
USD 75,000 - 95,000
Senior Platform Reliability Engineer - Remote
Senior Platform Reliability Engineer - Remote

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000