Senior SRE

Selby Jennings

Austin (TX)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Job summary

Selby Jennings is seeking a Senior Site Reliability Engineer to scale and support critical workflow orchestration and automation platforms across the organization. The role sits in Platform Engineering, delivering highly available and resilient infrastructure for business-critical workloads.

You will blend hands-on operational support with long-term engineering improvements, collaborating with engineering, data, infrastructure, and security teams to automate processes and modernize the

Qualifications

  • Bachelor's degree or equivalent experience required.
  • 5+ years in Site Reliability, Platform Engineering or DevOps roles.
  • Strong Airflow production experience and distributed deployments (Celery/Kubernetes).
  • Experience with Automic/UC4 and enterprise job scheduling is a plus.
  • Proficiency with Linux; Windows is a plus; AWS cloud experience.

Responsibilities

  • Own operational support for workflow orchestration platforms and scheduling.
  • Lead incident escalation, troubleshoot production issues, and perform RCAs.
  • Improve reliability standards, monitoring strategies, and runbooks.
  • Automate tasks and build tooling to reduce manual effort.
  • Design observability with metrics, dashboards, alerts, and logs.
  • Support platform lifecycle, upgrades, patching, and security remediation.
  • Contribute to modernization: cloud, containers, migrations, CI/CD pipelines.
  • Develop IaC using Terraform, Helm, Ansible; participate in architecture discussions.

Skills

SRE
DevOps
Platform Engineering
Automation
Cloud

Education

Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent experience

Tools

Airflow
Celery
Kubernetes
Terraform
Helm
Ansible
AWS
Python

Job description

A leading investment management firm is seeking a Senior Site Reliability Engineer to help scale and support critical workflow orchestration and automation platforms across the organization. This role sits within a Platform Engineering team responsible for delivering highly available, resilient, and scalable infrastructure that powers business-critical workloads and data processes.

The ideal candidate will have a strong background in SRE, DevOps, or Platform Engineering and enjoy balancing hands-on operational support with long-term engineering improvements. You'll work closely with engineering, data, infrastructure, and security teams to enhance platform reliability, automate manual processes, and drive modernization efforts across the environment.

Responsibilities
  • Provide operational ownership of workflow orchestration and enterprise scheduling platforms, including Apache Airflow and similar technologies.
  • Act as an escalation point for platform-related incidents, troubleshooting complex production issues and driving root cause analysis through resolution.
  • Partner with application and data teams to resolve workflow failures, dependency issues, scheduling conflicts, and performance bottlenecks.
  • Develop and maintain reliability standards, service objectives, monitoring strategies, and operational best practices.
  • Build automation and tooling that reduce manual effort and improve the overall user experience for engineering teams.
  • Design and implement observability solutions utilizing metrics, dashboards, alerting, logging, and performance monitoring.
  • Support platform lifecycle management, including upgrades, patching, configuration management, and security remediation.
  • Contribute to infrastructure modernization initiatives involving cloud services, containerization, platform migrations, and deployment automation.
  • Develop and maintain Infrastructure-as-Code solutions using tools such as Terraform, Helm, Ansible, and related technologies.
  • Participate in architectural discussions, platform roadmap planning, and engineering standards development.
  • Maintain operational documentation, procedures, and on-call readiness for supported environments.
Required Experience
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent experience.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or DevOps-focused roles.
  • Strong production experience supporting Apache Airflow environments.
  • Experience with distributed Airflow deployments, including Celery and/or Kubernetes executors.
  • Experience supporting enterprise workload automation and job scheduling platforms such as Automic/UC4, Control-M, or comparable technologies.
  • Strong Linux administration skills with working knowledge of Windows-based environments.
  • Experience supporting cloud infrastructure, preferably within AWS environments.
  • Proficiency with Python and scripting for automation, tooling, and operational efficiency.
  • Experience working with Kubernetes, Docker, CI/CD pipelines, and modern deployment methodologies.
  • Strong understanding of monitoring, logging, tracing, and observability concepts.
  • Experience with tools such as Grafana, Prometheus, ELK, or comparable monitoring platforms.
  • Proven ability to manage production incidents and communicate effectively during high-priority situations.
  • Strong automation mindset with a focus on improving efficiency and reducing operational overhead.
Preferred Qualifications
  • Hands-on experience administering Broadcom Automic/UC4.
  • Experience with managed Airflow platforms such as AWS MWAA, Cloud Composer, or Astronomer.
  • Exposure to modern data ecosystems, including technologies such as Kafka, dbt, and Snowflake.
  • Experience migrating workloads from legacy scheduling platforms to cloud-native orchestration solutions.
  • Familiarity with SLOs, SLIs, error budgets, and reliability engineering best practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Workflow Orchestration Engineer (Airflow & Scheduling Platforms)
Senior Workflow Orchestration Engineer (Airflow & Scheduling Platforms)

Benton Partners • Chicago (IL), New York (NY)

On-site
USD 140,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

State of Wisconsin Investment Board • Madison (WI)

On-site
USD 150,000 - 190,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Senior SRE (Contract/Hybrid)
Senior SRE (Contract/Hybrid)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

SCIGON • Naperville (IL)

Hybrid
USD 110,000 - 170,000