Lead Site Reliability Engineer | Production Infrastructure

P2P

Chicago (IL)

On-site

USD 175,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Discretionary bonus eligibility
Medical, dental, and vision insurance
Paid vacation plus paid holidays
Retirement plan with employer match
Paid parental leave

Job summary

P2P is looking for a Lead Site Reliability Engineer to join our CORE team in Chicago, Illinois. In this role, you will manage and mentor engineers while collaborating with various teams to ensure the resilience and efficiency of our production systems.

Your responsibilities include designing high-performance monitoring systems, overseeing incident management processes, and tackling complex reliability challenges. The position offers a competitive salary and comprehensive benefits package.

Qualifications

  • Proven leadership experience having managed people across distributed teams.
  • Demonstrated history of solving reliability challenges in large-scale production environments.
  • Strong programming skills in Python, Go, or equivalent.

Responsibilities

  • Manage and mentor engineers across teams and contribute directly to key projects.
  • Architect and implement high-performance monitoring and alerting systems.
  • Oversee and improve incident management and change management processes.
  • Identify and eliminate sources of operational toil through automation.
  • Investigate low-level performance issues across complex software stacks.
  • Influence the strategic direction of production tooling and infrastructure scaling.

Skills

Leadership experience
Solving reliability challenges
Strategic thinking
Programming skills in Python/Go

Job description

Jump Trading Group is committed to world class research. We empower exceptional talents in Mathematics, Physics, and Computer Science to seek scientific boundaries, push through them, and apply cutting edge research to global financial markets. Our culture is unique. Constant innovation requires fearlessness, creativity, intellectual honesty, and a relentless competitive streak. We believe in winning together and unlocking unique individual talent by incenting collaboration and mutual respect. At Jump, research outcomes drive more than superior risk adjusted returns. We design, develop, and deploy technologies that change our world, fund start-ups across industries, and partner with leading global research organizations and universities to solve problems.

CORE (Central Ops and Reliability Engineering) is the Production Infrastructure team responsible for operating and improving Jump’s production trading environment. The team combines deep operational ownership with software and reliability engineering practices to support production systems, drive incident and change management, improve observability and deployment workflows, and reduce operational toil across a fast-moving global trading platform.

What You’ll Do:

As Lead Site Reliability Engineer in CORE, you will both manage and mentor engineers across teams and contribute directly to key projects, balancing leadership responsibilities with hands‑on work.

  • Design & Build: Architect and implement high‑performance monitoring and alerting systems, real‑time packet/flow analysis tooling, and automation frameworks for managing Jump’s global production footprint.
  • Lead Operational Maturity: Oversee and improve incident management, change management, and post‑incident review processes to increase resilience and reduce downtime.
  • Drive Efficiency: Identify and eliminate sources of operational toil through automation and tooling.
  • Collaborate Globally: Partner with engineering, networking, and trading teams in multiple regions to align technical priorities with business objectives.
  • Debug Deeply: Investigate low‑level performance issues across complex software stacks, optimizing for ultra‑low latency and high throughput.
  • Shape the Roadmap: Influence the strategic direction of production tooling, infrastructure scaling, and vendor partnerships.
Skills You’ll Need:
  • Proven leadership experience having managed people across distributed teams.
  • Demonstrated history of solving reliability challenges in large‑scale production environments.
  • Previous experience demonstrating strategic thinking skills and maturity in tackling complex problems, dealing with people, technology and processes.
  • Strong programming skills in Python, Go, or equivalent.
Benefits
  • Discretionary bonus eligibility
  • Medical, dental, and vision insurance
  • HSA, FSA, and Dependent Care options
  • Employer Paid Group Term Life and AD&D Insurance
  • Voluntary Life & AD&D insurance
  • Paid vacation plus paid holidays
  • Retirement plan with employer match
  • Paid parental leave
  • Wellness Programs
Annual Base Salary Range

$175,000 — $200,000 USD

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Trading Production Engineer
Trading Production Engineer

Socket.dev • New York (NY)

On-site
USD 200,000 - 250,000
Discretionary bonus
Medical insurance
Dental insurance
+8
HPC Data Center Infrastructure Planning Lead
HPC Data Center Infrastructure Planning Lead

P2P • New York (NY)

On-site
USD 125,000 - 150,000
Discretionary bonus eligibility
Medical, dental, and vision insurance
Paid vacation plus paid holidays
+2
Software Engineer
Software Engineer

P2P • New York (NY)

On-site
USD 200,000 - 250,000
Discretionary bonus eligibility
Medical, dental, and vision insurance
HSA, FSA, and Dependent Care options
+6
Software Engineer | Market Data Systems
Software Engineer | Market Data Systems

Jump Trading • Chicago (IL)

On-site
USD 200,000 - 250,000
Discretionary bonus eligibility
Medical, dental, and vision insurance
HSA, FSA, and Dependent Care options
+6
Software Engineer | Market Data Systems
Software Engineer | Market Data Systems

P2P • Chicago (IL)

On-site
USD 200,000 - 250,000
Discretionary bonus eligibility
Medical, dental, and vision insurance
HSA, FSA, and Dependent Care options
+6
Campus Software Engineer (Full-Time)
Campus Software Engineer (Full-Time)

P2P • Chicago (IL)

On-site
USD 250,000
Private Medical
Vision and Dental Insurance
Travel Medical Insurance
+4
Campus Software Engineer (Full-Time)
Campus Software Engineer (Full-Time)

Jump Trading • Chicago (IL)

On-site
USD 250,000
Private Medical Insurance
Travel Medical Insurance
Group Pension Scheme
+3
HPC Data Center Technician
HPC Data Center Technician

P2P • Carrollton (TX)

On-site
USD 70,000 - 90,000
Private Medical, Vision and Dental Insurance
Travel Medical Insurance
Group Pension Scheme
+2
Software Engineer | Trading Team
Software Engineer | Trading Team

Jump Trading • Chicago (IL)

On-site
USD 175,000 - 225,000
Discretionary bonus eligibility
Medical, dental, and vision insurance
Paid vacation plus paid holidays
+1
Campus Systems Engineer (Intern)
Campus Systems Engineer (Intern)

P2P • Chicago (IL)

On-site
USD 200,000 - 250,000