Site Reliability Engineer (SRE)

JPS Tech Solutions

Colorado

On-site

USD 160,000 - 230,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPS Tech Solutions in Colorado seeks an experienced Site Reliability Engineer with a strong Java background to lead reliability initiatives and ensure stability, scalability, and performance of mission-critical systems.

The role blends hands-on engineering with leadership, ownership, and proactive operations, partnering with development, platform, and operations teams to implement SLOs/SLIs and CI/CD best practices.

Qualifications

  • 12+ years of IT experience in SRE, DevOps, or Production Engineering.
  • Strong Java development experience (Java 17+, Spring Boot Microservices, Spring Web).
  • Hands-on experience with OpenShift (OCP), Kubernetes, and Docker.
  • Strong expertise in MongoDB (data modeling, design, optimization).
  • Experience with Apache Kafka and event-driven architectures
  • Working knowledge of Oracle Database
  • Familiarity with BDD practices
  • Solid experience with CI/CD, automation, and IaC (Terraform, Ansible)
  • Exposure to AI-assisted development tools (e.g., GitHub Copilot)
  • Excellent troubleshooting skills in high-pressure production environments
  • Strong communication, collaboration, and ownership mindset

Responsibilities

  • Design, build, and maintain highly reliable, scalable, and fault-tolerant systems in production environments.
  • Embed reliability best practices (SLOs, SLIs, error budgets) into the software development lifecycle.
  • Work closely with development teams on Java Spring Boot microservices to improve operability and resilience.
  • Automate operational workflows to reduce manual effort and improve system efficiency.
  • Monitor system health, performance, and availability; proactively identify risks and bottlenecks.
  • Lead incident management, on-call support, and root cause analysis for production issues.
  • Drive continuous improvement initiatives focused on availability, scalability, and performance.
  • Support and oversee release and deployment activities, including after-hours support when required.
  • Champion best practices around CI/CD, infrastructure as code, and cloud-native operations.
  • Mentor engineers and provide technical leadership across SRE and development teams.
  • Collaborate with stakeholders to align reliability goals with business priorities.

Skills

Java development
SRE leadership
Cloud-native
CI/CD
IaC
OpenShift
Kubernetes
Docker
MongoDB
Kafka
JVM

Tools

OpenShift
Kubernetes
Docker
Terraform
Ansible
Prometheus
Grafana
ELK stack
Apache Kafka
Oracle DB

Job description

Job Description:

We are seeking a highly experienced Site Reliability Engineer (SRE) with a strong Java development background to lead reliability initiatives and ensure the stability, scalability, and performance of mission-critical systems. This role blends deep hands-on engineering with leadership, ownership, and a proactive approach to reliability and operations.


The ideal candidate is someone who has evolved from a strong developer into an SRE/DevOps leader, understands production systems deeply, and can partner effectively with development, platform, and operations teams.


Key Responsibilities:


  • Design, build, and maintain highly reliable, scalable, and fault-tolerant systems in production environments.

  • Embed reliability best practices (SLOs, SLIs, error budgets) into the software development lifecycle.

  • Work closely with development teams on Java Spring Boot microservices to improve operability and resilience.

  • Automate operational workflows to reduce manual effort and improve system efficiency.

  • Monitor system health, performance, and availability; proactively identify risks and bottlenecks.

  • Lead incident management, on-call support, and root cause analysis for production issues.

  • Drive continuous improvement initiatives focused on availability, scalability, and performance.

  • Support and oversee release and deployment activities, including after-hours support when required.

  • Champion best practices around CI/CD, infrastructure as code, and cloud-native operations.

  • Mentor engineers and provide technical leadership across SRE and development teams.

  • Collaborate with stakeholders to align reliability goals with business priorities.


Required Qualifications


  • 12+ years of IT experience in SRE, DevOps, or Production Engineering

  • Strong Java development experience (Java 17+, Spring Boot Microservices, Spring Web)

  • Hands-on experience with OpenShift (OCP), Kubernetes, and Docker

  • Strong expertise in MongoDB (data modeling, design, optimization)

  • Experience with Apache Kafka and event-driven architectures

  • Working knowledge of Oracle Database

  • Familiarity with BDD practices

  • Solid experience with CI/CD, automation, and IaC (Terraform, Ansible)

  • Exposure to AI-assisted development tools (e.g., GitHub Copilot)

  • Excellent troubleshooting skills in high-pressure production environments

  • Strong communication, collaboration, and ownership mindset


Preferred Qualifications:


  • Experience with monitoring and observability tools such as Prometheus, Grafana, and the ELK stack.

  • Knowledge of security best practices, compliance standards, and production hardening.

  • Prior experience leading or mentoring SRE teams or guiding engineers in reliability practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

Hybrid
USD 110,000 - 170,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

TechDigital Group • Fairfax (VA)

On-site
USD 100,000 - 130,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Site Reliability Engineer
Site Reliability Engineer

Shya Workforce Solutions • Town of Florida (NY)

On-site
USD 100,000 - 140,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000