Site Reliability Engineer

ICE Clear Europe Limited

Atlanta, Northern (GA, KY)

Hybrid

USD 120,000 - 150,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Intercontinental Exchange, Inc. seeks a Site Reliability Engineer II with 3+ years of hands-on SRE experience to enhance platform reliability and automation in ICE's 24x7 production environment.

You will partner with Product and Engineering to plan releases, design proactive monitoring, and mentor junior engineers while contributing to incident response and runbook automation.

Qualifications

  • 3+ years of site reliability or production engineering experience.
  • Strong teamwork and mentoring abilities.
  • Ability to prioritize and manage work with minimal supervision.
  • Familiarity with ICE core competencies.

Responsibilities

  • Use advanced troubleshooting and root-cause analysis to improve availability, performance, and security of services.
  • Collaborate with Product and Engineering to plan releases with quality gates.
  • Build and evolve shared services for platform teams.
  • Design proactive monitoring, alerting, and self-healing automation.
  • Resolve defects and operational issues with increasing independence.
  • Implement automated tests, deployments, and tooling across the SRE toolchain.
  • Ensure 24x7 availability and operational readiness of services.
  • Lead projects and provide status updates to management.
  • Mentor junior engineers and contribute to team knowledge sharing.
  • Build health checks post-deployment and during incidents to speed MTTR.
  • Develop AI-assisted diagnosis to correlate signals across monitoring.
  • Create automation pipelines (Rundeck, Jenkins) with AI/LLM tooling.
  • Tune AWS CloudWatch metrics, dashboards, and observability visualizations.

Skills

SRE experience
Incident response
Automation & CI/CD
Mentoring
Team collaboration
Observability
OpenTelemetry
Grafana
Jenkins
AWS CloudWatch
Cloud monitoring

Education

Bachelor's degree in Computer Science, Engineering, or equivalent experience

Tools

Terraform
Chef
Ansible
Rundeck
Jenkins
AWS CloudWatch
Grafana
OpenTelemetry/Alloy
Splunk
BigPanda
PagerDuty

Job description

Overview
Job Purpose

At Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing houses that connect companies around the world to global capital and derivative markets. With a leading-edge approach to developing technology platforms, we have built market infrastructure in all major trading centers, offering customers the ability to manage risk and make informed decisions globally. By leveraging our core strengths in technology, we continue to identify new ways to serve our customers and transform global markets. We're looking for motivated, results-oriented people to join our team.

We are seeking a Site Reliability Engineer II to bring 3+ years of hands-on experience to our SRE team, operating with significant autonomy to improve platform reliability, drive automation, and mentor junior engineers. The ideal candidate contributes meaningfully to platform release cycles, leads smaller projects, and actively shapes the team's approach to observability, incident response, and service design in ICE's 24x7 production environment.

Responsibilities
  • Employ advanced troubleshooting and root-cause analysis to improve availability, performance, and security of IMT and platform services
  • Collaborate with Product and Engineering teams to plan and deploy product releases with operational rigor and quality gates
  • Work with Engineering leadership to build and evolve shared services meeting the requirements of platform and application teams
  • Design and implement proactive monitoring, alerting, trend analysis, and self-healing automation
  • Resolve product and service defects, infrastructure issues, and operational changes with increasing independence
  • Implement automated tests, automated deployments, and operational tooling across the SRE toolchain
  • Ensure services are designed with 24x7 availability and operational readiness and rigor
  • Lead smaller projects and provide status updates to management and stakeholders
  • Mentor SRE I engineers and contribute actively to team training and knowledge-sharing
  • Partner with application and platform teams to identify critical workflows and build automated health checks that run post-deployment and during incidents to accelerate root-cause identification
  • Design and build AI-assisted automated diagnosis jobs that correlate signals across monitoring and alerting platforms to reduce Mean Time to Resolution (MTTR) for production incidents
  • Build and maintain automation pipelines (e.g., Rundeck, Jenkins) that integrate with AI/LLM tooling to drive efficiency gains in observability, runbook execution, and incident triage
  • Develop and tune AWS CloudWatch metrics, alarms, and dashboards, instrument services using OpenTelemetry/Alloy, and build observability visualizations in Grafana; integrate alerting and event correlation workflows across PagerDuty, BigPanda, and Splunk to ensure timely, actionable incident notification
Knowledge and Experience
  • Bachelor's degree in Computer Science, Engineering, or equivalent experience
  • 3+ years of experience in a site reliability, production engineering, or software operations role
  • Proven technical skills with strong personal initiative and consistent delivery of important work
  • Excellent teamwork with active involvement in training and mentoring
  • Ability to prioritize and execute without direct management guidance
  • Strong understanding of ICE Core Competencies
Preferred Knowledge and Experience
  • Experience in financial services technology, mortgage platforms, or exchange infrastructure
  • Familiarity with SRE principles including SLI, SLO, and error budget management
  • Exposure to Terraform, Chef, Ansible, or equivalent infrastructure automation frameworks
  • Hands-on experience with AWS observability services, CloudWatch, Grafana, OpenTelemetry/Alloy, Splunk, BigPanda, PagerDuty, and job orchestration/automation platforms such as Rundeck and Jenkins
  • Practical experience integrating AI/LLM-based tooling into operational workflows to automate diagnosis, reduce manual triage, and improve incident response efficiency

#LI-JM1

Intercontinental Exchange, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to legally protected characteristics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

ICE • Atlanta (GA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Intercontinental Exchange Holdings, Inc. • Atlanta (GA)

On-site
USD 120,000 - 150,000
SRE Engineer: AI‑Driven Reliability & Automation
SRE Engineer: AI‑Driven Reliability & Automation

ICE • Atlanta (GA)

On-site
USD 140,000 - 210,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

ICE Clear Europe Limited • Jacksonville (FL)

On-site
USD 120,000 - 170,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE) – II
Site Reliability Engineer (SRE) – II

Huntington National Bank • Columbus (OH)

Hybrid
USD 90,000 - 120,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

ICE • Jacksonville (FL)

On-site
USD 140,000 - 190,000