Lead Site Reliability Engineer (SRE)

Optimal Market Technologies

New York (NY)

On-site

USD 175,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Optimal Market Technologies is seeking a Lead Site Reliability Engineer to own production systems administration across colocation and cloud. This hands-on leadership role focuses on automating operations, publishing standards, and reducing key-person risk while partnering with the development team to keep systems reliable and fast.

You will mentor staff, implement Infrastructure as Code, manage incidents, and oversee recovery drills.

Qualifications

  • Experience leading production systems in an engineering-driven environment.
  • Strong scripting and automation skills.
  • Hands-on Linux and network administration.
  • Experience with Infrastructure as Code.
  • Experience with AI tools (Claude Code preferred).
  • Proven track record in production support and incident response.
  • Experience managing and developing technical staff.
  • Ability to introduce structure, standards and strategy to a growing function.
  • Strong communication with senior stakeholders.

Responsibilities

  • Own how production runs across colocation and the cloud: deployment, capacity, and failover.
  • Build and lead the systems administration function.
  • Set and publish engineering standards and strategy for running production.
  • Hands-on Linux and network administration; automate with Infrastructure as Code.
  • Manage vendors and service agreements; advise on build-vs-contract-out.
  • Own infrastructure security: hardening, access control, recoverable backups, and security incident response.
  • Assist first-line production support, reducing reliance on the development team.
  • Be accountable for production stability: track what breaks and why, and automate to prevent it.
  • Own incident response, on-call, and post-incident review; coverage is market-hours plus a support rotation.
  • Own recovery runbooks, and recovery drills.
  • Automate client self-service for common issues and access to their own data.
  • Partner with the development team on deployments, and on performance tracking and capacity planning.

Skills

Scripting & automation
Linux administration
Network administration
Infrastructure as Code
AI tools (Claude Code)
Production support
Team leadership
Strategic standards
Stakeholder communication

Job description

Position Overview

We're seeking a Lead Site Reliability Engineer to oversee production systems administration. We are moving from an individual‑knowledge work style to an engineering‑run discipline: automated, reliable, and built on published standards. You will build and lead our systems administration function, professionalize how we run our infrastructure, and reduce key‑person risk, partnering with the development team to keep the firm running reliably and moving fast. This is a hands‑on role; you will build and operate the systems you put in place. You will report to the CTO.

Our Environment

We run an automated trading system with single‑digit millisecond latency requirements on bare‑metal Linux. We use Azure for development environments, storage, and offline studies, not for production execution. Systems are written in C++, Python, and SQL. We are actively modernizing, upgrading technologies (e.g. CentOS 7 to RHEL 9), and have legacy and new systems running in parallel. You will lead the rollout of a stream of technology changes.

Primary Responsibilities
  • Own how production runs across colocation and the cloud: deployment, capacity, and failover.
  • Build and lead the systems administration function: mentor existing staff, set how the function works, and hire as we grow.
  • Set and publish the engineering standards and strategy for how we run production.
  • Hands‑on Linux and network administration; automate routine work through Infrastructure as Code.
  • Manage vendors and service agreements; advise on build‑vs‑contract‑out.
  • Own infrastructure security: hardening, access control, recoverable backups, and security incident response.
Production Support, Incident Response, Resilience, and Performance
  • Assist first‑line production support, reducing reliance on the development team.
  • Be accountable for production stability: track what breaks and why, and turn repeat firefighting into automation that prevents it.
  • Own incident response, on‑call, and post‑incident review; coverage is market‑hours plus a support rotation.
  • Own recovery runbooks, and recovery drills.
  • Automate client self‑service for common issues and access to their own data, reducing manual support work.
  • Partner with the development team on deployments, and on performance tracking and capacity planning.
What We're Looking For
  • Strong scripting and automation skills.
  • Strong hands‑on Linux and network administration.
  • Expertise with Infrastructure as Code (we are open on which tools).
  • Experience using AI tools, ideally Claude Code.
  • A track record owning production support and incident response.
  • Experience managing and developing technical staff.
  • The ability to bring structure, standards, and strategy to a function as it grows and matures.
  • Strong communication; effective with senior stakeholders and a small team.
Preferred Qualifications
  • Experience at a start‑up, or building a new line or function inside a larger firm; comfortable under resource constraints and automation‑first by instinct.
  • Experience in real‑time critical systems.
  • Familiarity with any of: FIX Protocol, PostgreSQL, middleware (ideally AERON), observability and monitoring tooling.
  • Microsoft Azure, including Azure Virtual Desktop (AVD).
  • Vendor management.
Salary Range

$175,000 USD - $200,000 USD

Optimal is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

optimal sp. z o.o. • Chicago (IL)

Hybrid
USD 150,000 - 250,000
Lead Sysadmin/SRE
Lead Sysadmin/SRE

Optimal Market Technologies, LLC • New York (NY), Chicago (IL)

Hybrid
USD 150,000 - 250,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

Optimal Market Technologies • Chicago (IL)

Hybrid
USD 175,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Longbridge Securities • Town of Texas (WI)

On-site
USD 100,000 - 130,000
Competitive compensation package
Growth opportunities
Senior Site Reliability Engineer - Banking & Finance
Senior Site Reliability Engineer - Banking & Finance

Hamilton Barnes Associates Limited • New York (NY)

Hybrid
USD 360,000 - 440,000
Strong compensation and bonus potential
Collaborative engineering culture
Work on mission-critical systems
Site Reliability Engineer
Site Reliability Engineer

Engtal • Chicago (IL)

On-site
USD 250,000 - 350,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000