Lead Performance & Observability Engineer

ICE

Atlanta (GA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading financial services company in Atlanta is looking for a Lead Performance Engineer to define performance engineering strategies and lead investigations on complex Java platforms. The ideal candidate has over 8 years of experience in performance testing and is skilled in JVM internals, Java, and Prometheus. Strong communication skills are essential, as the role requires collaboration across various teams to ensure system performance and reliability standards are met.

Qualifications

  • 8+ years of experience in performance engineering or Java development.
  • Experience with load generation frameworks (JMeter, Gatling).
  • Expertise in tuning multi-threaded Java applications.

Responsibilities

  • Define and own performance engineering strategy for critical platforms.
  • Lead performance investigations on JVM and hardware levels.
  • Design and build robust test harnesses for performance regressions.

Skills

Deep expertise in JVM internals
Proficiency with Java and scripting languages
Performance testing
Hands-on experience with Prometheus
Excellent verbal and written communication skills

Education

Bachelor's Degree in Computer Science or Engineering

Tools

Valgrind
Prometheus
Grafana
JMeter

Job description

Job Purpose

Intercontinental Exchange, Inc. (ICE), the owner of the New York Stock Exchange (NYSE), is seeking a results‑oriented, self‑motivated individual for its Capacity and Performance Management team in Atlanta. This individual will serve as a technical lead within a team of software architects and performance engineers, operating in a cutting‑edge technology environment responsible for running critical financial sector exchanges and clearinghouses. The successful candidate will drive performance engineering strategy across multiple platforms, mentor peers, and deliver deep‑dive analysis on the most complex and time‑sensitive systems in the organization.

Overview

You must be technically authoritative, outcome‑focused, and capable of thriving in a mission‑critical environment where end‑of‑day processing windows are measured in minutes. This role requires close collaboration with software architects, developers, quant library owners, infrastructure teams, and project managers to ensure our platforms meet the highest standards of reliability, scalability, and throughput.

As a Lead Performance Engineer you will own the end‑to‑end performance strategy for complex, event‑driven Java platforms. You will define testing methodologies, build durable test harnesses, drive observability initiatives, and act as the authoritative voice on performance trade‑offs across application, JVM, OS, and hardware layers.

Responsibilities
  • Define and own the performance engineering strategy across multiple critical platforms; set standards for testing approach, tooling, and reporting across the team.
  • Lead deep‑dive performance investigations on CPU‑bound, multi‑threaded Java systems—including analysis at the JVM, OS, and hardware (NUMA, hyper‑threading) levels.
  • Profile and diagnose performance at JNI boundaries between Java and native C++ libraries; identify and quantify overhead introduced at language‑crossing layers.
  • Design and build robust test harnesses that accurately measure version‑to‑version performance regressions for compute‑intensive components.
  • Tune JVM thread pools, garbage collection, and heap allocation for high‑throughput, latency‑sensitive processing pipelines.
  • Analyze multi‑threaded concurrency—thread contention, core utilization, and parallel subgroup scheduling—to optimize throughput on dedicated hardware and horizontally scaled worker architectures.
  • Drive scalability analysis: model how system performance scales linearly (or non‑linearly) as data volume and instrument counts grow; produce capacity projections to guide architecture decisions.
  • Lead critical path segregation analysis: identify which operations are time‑constrained, propose architectural solutions (e.g., head‑start strategies, out‑of‑band processing), and validate their impact.
  • Build and maintain Performance Engineering KPI dashboards; drive adoption of Prometheus instrumentation and Grafana visualization across platform components.
  • Certify system reliability and failover behavior under stress conditions; validate that hot/cold and distributed worker architectures perform correctly and within SLA under load.
  • Create and maintain automation scripts and tooling to simplify repeatable performance analysis tasks.
  • Act as a technical liaison between performance, development, infrastructure, and operations teams; translate performance findings into actionable recommendations with clear data backing.
  • Champion AI‑augmented workflows within the performance team: identify where AI tools accelerate root‑cause analysis, reduce boilerplate in test harnesses, and surface anomalies in benchmark data — while establishing team standards for validating AI‑generated output before it reaches production.
Knowledge and Experience
  • Bachelor's Degree or equivalent in Computer Science, Engineering, or a related field.
  • 8+ years of experience in performance engineering, performance testing, or Java development in high‑volume, low‑latency, transactional systems.
  • Familiarity with tools such as Valgrind, perf, gprof, or equivalent for analyzing CPU‑bound native code.
  • Deep expertise in JVM internals: heap dump analysis, thread dump analysis, GC log interpretation, and memory profiling.
  • Experience tuning multi‑threaded Java applications on dedicated hardware—thread pool sizing, core affinity, NUMA awareness, hyper‑threading trade‑offs.
  • Proficiency with Java and scripting languages such as Python, Groovy, or Linux shell for building test harness foundation and automation.
  • Experience with event‑driven, message‑based architectures (Kafka or equivalent); ability to test and profile throughput and latency across event pipelines.
  • Hands‑on experience with Prometheus metrics instrumentation and Grafana dashboard construction for operational and performance monitoring.
  • Experience with load generation frameworks (JMeter, Gatling, or custom harnesses) and the ability to design workloads representative of real production traffic patterns.
  • Proven ability to perform scalability analysis—projecting how systems behave as data volumes grow and validating linearity assumptions with empirical evidence.
  • Strong grasp of critical path analysis: ability to identify time‑constrained operations, separate them from non‑critical work, and validate architectural changes that improve throughput within a fixed time window.
  • Excellent verbal and written communication skills; able to present findings clearly to both technical architects and non‑technical stakeholders.
  • Ability to work across teams with varying levels of performance domain knowledge and build collaborative relationships with application owners.
  • Demonstrated proficiency with AI coding and analysis tools (GitHub Copilot, Claude, Cursor, or equivalent) for performance‑specific tasks: generating profiling scripts, writing Prometheus queries, analyzing GC logs, and building test harness scaffolding.
  • Proven habit of validating AI‑generated output: identifying hallucinated APIs, incorrect performance assumptions, or subtly broken concurrency logic before it reaches production.

Intercontinental Exchange, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to legally protected characteristics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Engineer
Performance Engineer

LeoForce • Chicago (IL)

On-site
USD 200,000 - 260,000
Lead Performance Engineer: High-Throughput Java Systems
Lead Performance Engineer: High-Throughput Java Systems

ICE • Atlanta (GA)

On-site
USD 120,000 - 150,000
Java Performance Engineer
Java Performance Engineer

krg technology inc • Orlando (FL)

On-site
USD 80,000 - 100,000
Lead Performance Engineer
Lead Performance Engineer

PineQ Lab Technology • Sunnyvale (CA)

On-site
USD 150,000 - 230,000
Senior Java Developer
Senior Java Developer

ICE • Atlanta (GA)

On-site
USD 120,000 - 180,000
Principal, Software Engineering: Software Development Test (SDET)
Principal, Software Engineering: Software Development Test (SDET)

theocc • United States

On-site
USD 170,000 - 230,000
Performance Engineer
Performance Engineer

Compunnel, Inc. • Town of Texas (WI)

On-site
USD 100,000 - 120,000
Performance Engineering
Performance Engineering

Two95 International Inc. • Richfield (MN)

On-site
USD 90,000 - 120,000
Principal Performance Engineer
Principal Performance Engineer

Apexon • St. Louis (MO)

Hybrid
USD 100,000 - 130,000
Performance Engineer
Performance Engineer

Veriipro • Boston (MA)

On-site
USD 100,000 - 140,000