Senior Engineer, Platform & Site Reliability

ICE Clear Europe Limited

Jacksonville (FL)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intercontinental Exchange, Inc. seeks a Senior Engineer to own the cloud-native platform and reliability tooling for MSP DX (IMT). You will manage EKS clusters across regions, implement GitOps, and contribute to Java microservices and React frontends that run on the platform.

Responsibilities include building CI/CD pipelines, designing observability solutions, and ensuring secure, scalable, highly available systems. Strong collaboration with product teams is essential.

Qualifications

  • Bachelor's degree or equivalent experience in a relevant field.
  • 5+ years of site reliability, platform, DevOps, or software engineering work experience.
  • Hands-on experience operating Kubernetes and cloud-native tech in production, preferably EKS on AWS.
  • Experience with infrastructure as code and GitOps tools.
  • Experience building and maintaining CI/CD pipelines.
  • Proficiency in Java or TypeScript/JavaScript for automation and tooling.

Responsibilities

  • Provision, upgrade, and operate EKS clusters across multiple regions.
  • Manage cloud infrastructure declaratively with Crossplane and implement GitOps with ArgoCD.
  • Operate and upgrade Istio service mesh, including canary rollouts and TLS.
  • Design and operate monitoring, logging, and observability using Prometheus, Grafana, Jaeger, OTEL, and Kiali.

Skills

Kubernetes
AWS
GitOps
ArgoCD
Crossplane
Istio
Observability
Java
TypeScript
Platform Engineer

Education

Bachelor's Degree

Tools

Azure DevOps
Jira
OpenShift
Git
Prometheus
Grafana
Jaeger
Kiali
DMS
MSK

Job description

Overview

As a key player within Intercontinental Exchange's (ICE) innovative servicing technology division, our team is dedicated to deliveringcutting-edgemortgage processing solutions on a resilient, cloud-native platform. This role is pivotal to our platform engineering and site reliability initiatives, owning the Amazon EKS (Kubernetes) foundation on which our product teams build and run. The Senior Engineer, Platform & Site Reliability will apply deepexpertisein AWS,GitOpswithArgoCD,Crossplane, the Istioservice mesh, and modern observability to keep our systems secure, scalable, andhighly availableacross multiple regions and environments. Java (Spring) and React (TypeScript) development remain part of the role, ensuring you can build platform tooling and contribute to the applications youoperate. By joining our team, you will directly shape the reliability, performance, and operability of the platform that powers our business.

Designs, builds, andoperates the cloud-native platform and reliability tooling for the MSP DX (IMT) with an emphasis on a secure, scalable, andhighly availablefoundation. Our engineers manage Amazon EKS clusters across multiple AWS regions and environments in an Agile SDLC, delivered entirely throughGitOps. Responsible for the provisioning and lifecycle of Kubernetes clusters and cloud infrastructure, CI/CD pipelines, service mesh, observability frameworks, and secrets management, while also contributing to the Java microservices andReactmicro frontends that run on the platform.

Responsibilities
  • Provisions, upgrades, andoperates Amazon EKS (Kubernetes) clusters across multiple AWS regions and environments (UAT, stable, production, and chaos), ensuring a secure, scalable, andhighly availableplatform.

  • Manages cloud infrastructure declaratively through Kubernetes usingCrossplane, and implementsGitOpspractices withArgoCDto deliver all infrastructure as code.

  • Operates and upgrades the Istio service mesh, including canary rollouts, along with Envoy and ingress, for traffic management, routing, and mutual TLS.

  • Designs andoperatesmonitoring, logging, and observability solutions using Prometheus, Grafana, Jaeger,OpenTelemetry(OTEL), Kiali, and Fluent Bit, and defines SLOs, alerting, and dashboards.

  • Improves system reliability through capacity planning, autoscaling, incident response, on-call practices, and chaos engineering.

  • Administers platform services such as cert-manager, external-dns, external-secrets, sealed-secrets, and the AWS Load Balancer Controller, and integrates with AWS services including DMS and MSK.

  • Builds andmaintainsCI/CD pipelines (Azure DevOps) and automated security scanning (for example, SonarQube) to deliver changes safely and repeatably.

  • Provides full-stack Java (Spring) and React (TypeScript) development for platform tooling and for the microservices and micro frontends that run on the platform.

  • Designs and develops APIs and automation that support platform capabilities and self-service for product teams.

  • Participates in reliability and architecture design ceremonies and analyzes system needs todeterminetechnical requirements.

  • Writes technical specifications and operational runbooks based on conceptual design and stated business and reliability requirements.

  • Develops and/or reviews automated tests and reliability validation before release, with an emphasis on Unit,Component, and Scenario tests.

  • Troubleshootsoperational failures in both test and production environments andleads rootcause analysis.

  • Mentors or guides the work of less experienced site reliability and software engineers.

  • Remains current on industry standards in cloud, DevOps, SRE, and web development disciplines.

  • Performsadditionalrelated duties as assigned.

Knowledge and Experience
  • Bachelor's Degree or the equivalent combination of education, training, or work experience.

  • 5+ years of site reliability, platform, DevOps, or software engineering work experience.

  • Hands-on experience operating Kubernetes and cloud-native technologies in production, preferably Amazon EKS on AWS.

  • Experience with infrastructure as code andGitOpspractices and tools.

  • Experience developing andmaintainingCI/CD pipelines.

  • Workingproficiencyin at least one general-purpose language, such as Java or TypeScript/JavaScript,sufficientto build automation and platform tooling.

Preferred Knowledge and Experience
  • Experience with Kubernetes,ArgoCD,Crossplane, Istio, Envoy, Jaeger, Prometheus, Grafana, Kiali, or similar technologies.

  • Experience running workloads with cloud providers (preferably AWS) and/or OpenShift, and with the Java JVM.

  • Experience with modern observability frameworks, including SLO/SLI definition and alerting.

  • Experienceoperatinga service mesh, including upgrades and canary rollouts.

  • Experience with secrets management and certificate automation, such as external-secrets, sealed-secrets, and cert-manager.

  • Experience with chaos engineering and resilience testing.

  • Experience with RESTful service development and working with microservices applications.

  • Experience with Postgres SQL Databases and PL/SQL.

  • Experience with modern JavaScript frameworks such as React.

  • Familiarity with source code management tools such as Azure DevOps, TFS, Jira, or Git.

  • Proficiencywith development techniques such as Test-Driven Development (TDD and BDD), Unit Tests,ComponentTests, and/or Scenario Tests.

  • Familiarity working in a Software Development Life Cycle (SDLC)leveragingAgile principles.

Intercontinental Exchange, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to legally protected characteristics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

Intercontinental Exchange Holdings, Inc. • Jacksonville (FL)

On-site
USD 140,000 - 180,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

ICE • Jacksonville (FL)

On-site
USD 140,000 - 190,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

ICE • Atlanta (GA)

On-site
USD 150,000 - 190,000
Principal Engineer, Full Stack
Principal Engineer, Full Stack

ICE Clear Europe Limited • Jacksonville (FL)

On-site
USD 180,000 - 240,000
Principal Engineer, Full Stack
Principal Engineer, Full Stack

ICE • Jacksonville (FL)

On-site
USD 150,000 - 210,000
Principal Engineer, Full Stack
Principal Engineer, Full Stack

Intercontinental Exchange Holdings, Inc. • Jacksonville (FL)

On-site
USD 150,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ICE • Jacksonville (FL)

On-site
USD 110,000 - 140,000
Engineer, Platform Engineering
Engineer, Platform Engineering

Intercontinental Exchange (ICE) • Atlanta (GA)

On-site
USD 80,000 - 100,000
Senior Platform & SRE Engineer — AWS/EKS
Senior Platform & SRE Engineer — AWS/EKS

ICE • Atlanta (GA)

On-site
USD 150,000 - 190,000
Senior Platform & SRE Engineer – Cloud‑Native, Kubernetes
Senior Platform & SRE Engineer – Cloud‑Native, Kubernetes

Intercontinental Exchange Holdings, Inc. • Jacksonville (FL)

On-site
USD 140,000 - 180,000