Senior Engineer, Platform & Site Reliability

Intercontinental Exchange (ICE)

Jacksonville (FL)

On-site

USD 140,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Intercontinental Exchange in Jacksonville, FL is seeking a Senior Engineer for Platform & Site Reliability to own the cloud-native platform and reliability tooling, focusing on EKS clusters, GitOps, and secure, scalable operations across regions.

You will help build platform tooling with Java Spring and React, contribute to CI/CD pipelines, observability, and incident response, while mentoring teammates and shaping architecture in a fast-paced Agile environment.

Qualifications

  • Bachelor’s degree or the equivalent combination of education, training, or work experience.
  • 5+ years of site reliability, platform, DevOps, or software engineering work experience.
  • Hands-on experience operating Kubernetes and cloud-native technologies in production, preferably Amazon EKS on AWS.
  • Experience with infrastructure as code and GitOps practices and tools.
  • Experience developing and maintaining CI/CD pipelines.
  • Working proficiency in Java or TypeScript/JavaScript to build automation and platform tooling.

Responsibilities

  • Provisions, upgrades, and operates Kubernetes clusters across multiple AWS regions and environments.
  • Manages cloud infrastructure declaratively through Kubernetes using Crossplane and GitOps via ArgoCD.
  • Operates Istio service mesh, canary rollouts, and Envoy ingress for traffic management.
  • Designs and operates monitoring, logging, and observability solutions using Prometheus, Grafana, Jaeger, OpenTelemetry, Kiali; defines SLOs and dashboards.
  • Improves system reliability through capacity planning, autoscaling, incident response, on-call, and chaos engineering.
  • Maintains platform services such as cert-manager, external-dns, external-secrets, sealed-secrets, and AWS Load Balancer Controller; integrates with AWS services.
  • Builds and maintains CI/CD pipelines (Azure DevOps) and automated security scanning (e.g., SonarQube).
  • Provides full-stack Java (Spring) and React tooling for platform and microservices running on the platform.
  • Designs and develops APIs and automation to support platform capabilities and self-service for product teams.
  • Participates in reliability and architecture design ceremonies; analyzes system needs to determine requirements.
  • Writes technical specifications and runbooks based on design and reliability requirements.
  • Develops and reviews automated tests and reliability validation; emphasizes Unit/Component/Scenario tests.
  • Troubleshoots failures and leads root cause analysis; mentors junior engineers.

Skills

AWS
GitOps
Kubernetes
Java
React
CI/CD
Observability

Education

Bachelor’s Degree in Computer Science or related

Tools

ArgoCD
Crossplane
Istio
Envoy
Prometheus
Grafana
Jaeger
Kiali
Azure DevOps
SonarQube

Job description

Job Purpose

As a key player within Intercontinental Exchange's (ICE) innovative servicing technology division, our team is dedicated to delivering cutting-edge mortgage processing solutions on a resilient, cloud-native platform. This role is pivotal to our platform engineering and site reliability initiatives, owning the Amazon EKS (Kubernetes) foundation on which our product teams build and run. The Senior Engineer, Platform & Site Reliability will apply deep expertise in AWS, GitOps with ArgoCD, Crossplane, the Istio service mesh, and modern observability to keep our systems secure, scalable, and highly available across multiple regions and environments. Java (Spring) and React (TypeScript) development remain part of the role, ensuring you can build platform tooling and contribute to the applications you operate. By joining our team, you will directly shape the reliability, performance, and operability of the platform that powers our business.

Designs, builds, and operates the cloud-native platform and reliability tooling for the MSP DX (IMT) with an emphasis on a secure, scalable, and highly available foundation. Our engineers manage Amazon EKS clusters across multiple AWS regions and environments in an Agile SDLC, delivered entirely through GitOps. Responsible for the provisioning and lifecycle of Kubernetes clusters and cloud infrastructure, CI/CD pipelines, service mesh, observability frameworks, and secrets management, while also contributing to the Java microservices and React micro frontends that run on the platform.

Responsibilities
  • Provisions, upgrades, and operates Amazon EKS (Kubernetes) clusters across multiple AWS regions and environments (UAT, stable, production, and chaos), ensuring a secure, scalable, and highly available platform.
  • Manages cloud infrastructure declaratively through Kubernetes using Crossplane, and implements GitOps practices with ArgoCD to deliver all infrastructure as code.
  • Operates and upgrades the Istio service mesh, including canary rollouts, along with Envoy and ingress, for traffic management, routing, and mutual TLS.
  • Designs and operates monitoring, logging, and observability solutions using Prometheus, Grafana, Jaeger, OpenTelemetry (OTEL), Kiali, and Fluent Bit, and defines SLOs, alerting, and dashboards.
  • Improves system reliability through capacity planning, autoscaling, incident response, on-call practices, and chaos engineering.
  • Administers platform services such as cert-manager, external-dns, external-secrets, sealed-secrets, and the AWS Load Balancer Controller, and integrates with AWS services including DMS and MSK.
  • Builds and maintains CI/CD pipelines (Azure DevOps) and automated security scanning (for example, SonarQube) to deliver changes safely and repeatably.
  • Provides full-stack Java (Spring) and React (TypeScript) development for platform tooling and for the microservices and micro frontends that run on the platform.
  • Designs and develops APIs and automation that support platform capabilities and self-service for product teams.
  • Participates in reliability and architecture design ceremonies and analyzes system needs to determine technical requirements.
  • Writes technical specifications and operational runbooks based on conceptual design and stated business and reliability requirements.
  • Develops and/or reviews automated tests and reliability validation before release, with an emphasis on Unit, Component, and Scenario tests.
  • Troubleshoots operational failures in both test and production environments and leads root cause analysis.
  • Mentors or guides the work of less experienced site reliability and software engineers.
  • Remains current on industry standards in cloud, DevOps, SRE, and web development disciplines.
  • Performs additional related duties as assigned.
Knowledge and Experience
  • Bachelor's Degree or the equivalent combination of education, training, or work experience.
  • 5+ years of site reliability, platform, DevOps, or software engineering work experience.
  • Hands-on experience operating Kubernetes and cloud-native technologies in production, preferably Amazon EKS on AWS.
  • Experience with infrastructure as code and GitOps practices and tools.
  • Experience developing and maintaining CI/CD pipelines.
  • Working proficiency in at least one general-purpose language, such as Java or TypeScript/JavaScript, sufficient to build automation and platform tooling.
Preferred Knowledge and Experience
  • Experience with Kubernetes, ArgoCD, Crossplane, Istio, Envoy, Jaeger, Prometheus, Grafana, Kiali, or similar technologies.
  • Experience running workloads with cloud providers (preferably AWS) and/or OpenShift, and with the Java JVM.
  • Experience with modern observability frameworks, including SLO/SLI definition and alerting.
  • Experience operating a service mesh, including upgrades and canary rollouts.
  • Experience with secrets management and certificate automation, such as external-secrets, sealed-secrets, and cert-manager.
  • Experience with chaos engineering and resilience testing.
  • Experience with RESTful service development and working with microservices applications.
  • Experience with Postgres SQL Databases and PL/SQL.
  • Experience with modern JavaScript frameworks such as React.
  • Familiarity with source code management tools such as Azure DevOps, TFS, Jira, or Git.
  • Proficiency with development techniques such as Test-Driven Development (TDD and BDD), Unit Tests, Component Tests, and/or Scenario Tests.
  • Familiarity working in a Software Development Life Cycle (SDLC) leveraging Agile principles.

Intercontinental Exchange, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to legally protected characteristics.

Get your free, confidential resume review.

or drag and drop your file here.