Site Reliability Engineer (Lead)

Boulder Connect

O’Fallon (MO)

On-site

USD 120,000 - 180,000

Full time

4 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Boulder Connect is seeking a Lead Site Reliability Engineer to design resilient cloud-native architectures for mission-critical financial services platforms. You will build and optimize CI/CD pipelines and implement comprehensive monitoring, alerting, and incident response practices to maintain reliability.

You will collaborate with software engineering, security, and operations teams to strategize and improve system reliability, and lead blameless postmortems while promoting SRE best practices

Qualifications

  • Extensive experience in cloud-native architectures and scalable infrastructure.
  • Proficiency with CI/CD pipelines and automation tools.
  • Experience with monitoring, alerting and incident management.
  • Strong collaboration and communication across technical teams.
  • Knowledge of SRE principles and reliability best practices.

Responsibilities

  • Design and implement resilient cloud-native architectures for mission-critical financial services platforms.
  • Build and enhance CI/CD pipelines to ensure smooth deployment processes.
  • Develop comprehensive monitoring, alerting, and incident response practices to maintain system reliability.
  • Collaborate with software engineering, security, and operations teams to strategize and improve system reliability.
  • Lead blameless postmortems and promote SRE best practices across the organization.

Skills

Cloud-native architectures
CI/CD pipelines
Monitoring & incident management
Collaborative communication
SRE principles & best practices

Job description

Lead Site Reliability Engineer
Responsibilities
  • Design and implement resilient cloud-native architectures for mission-critical financial services platforms.
  • Build and enhance CI/CD pipelines to ensure smooth deployment processes.
  • Develop comprehensive monitoring, alerting, and incident response practices to maintain system reliability.
  • Collaborate with software engineering, security, and operations teams to strategize and improve system reliability.
  • Lead blameless postmortems and promote SRE best practices across the organization.
Qualifications & Skills
  • Strong expertise in cloud-native architectures and scalable infrastructure.
  • Proficiency in CI/CD pipeline development and automation tools.
  • Experience with monitoring, alerting, and incident management systems.
  • Excellent collaboration and communication skills to work across technical teams.
  • Knowledge of SRE principles and best practices to enhance system performance and reliability.
Get your free, confidential resume review.

or drag and drop your file here.