Lead Site Reliability Engineer - Architect & Own Production

Optimal Market Technologies

New York (NY)

On-site

USD 175,000 - 200,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Optimal Market Technologies is seeking a Lead Site Reliability Engineer to own production systems administration across colocation and cloud. This hands-on leadership role focuses on automating operations, publishing standards, and reducing key-person risk while partnering with the development team to keep systems reliable and fast.

You will mentor staff, implement Infrastructure as Code, manage incidents, and oversee recovery drills.

Qualifications

  • Experience leading production systems in an engineering-driven environment.
  • Strong scripting and automation skills.
  • Hands-on Linux and network administration.
  • Experience with Infrastructure as Code.
  • Experience with AI tools (Claude Code preferred).
  • Proven track record in production support and incident response.
  • Experience managing and developing technical staff.
  • Ability to introduce structure, standards and strategy to a growing function.
  • Strong communication with senior stakeholders.

Responsibilities

  • Own how production runs across colocation and the cloud: deployment, capacity, and failover.
  • Build and lead the systems administration function.
  • Set and publish engineering standards and strategy for running production.
  • Hands-on Linux and network administration; automate with Infrastructure as Code.
  • Manage vendors and service agreements; advise on build-vs-contract-out.
  • Own infrastructure security: hardening, access control, recoverable backups, and security incident response.
  • Assist first-line production support, reducing reliance on the development team.
  • Be accountable for production stability: track what breaks and why, and automate to prevent it.
  • Own incident response, on-call, and post-incident review; coverage is market-hours plus a support rotation.
  • Own recovery runbooks, and recovery drills.
  • Automate client self-service for common issues and access to their own data.
  • Partner with the development team on deployments, and on performance tracking and capacity planning.

Skills

Scripting & automation
Linux administration
Network administration
Infrastructure as Code
AI tools (Claude Code)
Production support
Team leadership
Strategic standards
Stakeholder communication

Job description

Optimal Market Technologies is seeking a Lead Site Reliability Engineer to own production systems administration across colocation and cloud. This hands-on leadership role focuses on automating operations, publishing standards, and reducing key-person risk while partnering with the development team to keep systems reliable and fast.

You will mentor staff, implement Infrastructure as Code, manage incidents, and oversee recovery drills.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer - Cloud Reliability & DR
Lead Site Reliability Engineer - Cloud Reliability & DR

Peraton • Northern (KY)

Hybrid
USD 112,000 - 179,000
Senior Site Reliability Engineer — Remote Production Reliability
Senior Site Reliability Engineer — Remote Production Reliability

Fingerprint • Chicago (IL)

Remote
USD 152,000 - 205,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Ll Oefentherapie • Nashville (TN)

On-site
USD 140,000 - 180,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000
Production Site Reliability Engineering Lead
Production Site Reliability Engineering Lead

WebHosting • Northern (KY)

Hybrid
USD 140,000 - 160,000
401(k) match up to 4%
Paid holidays & PTO
Professional development reimbursement
+1
Lead Site Reliability Engineer — Cloud & Automation
Lead Site Reliability Engineer — Cloud & Automation

Optimum • Plano (TX)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer — PubTech Scale
Senior Site Reliability Engineer — PubTech Scale

Amazon • Seattle (WA)

On-site
USD 151,000 - 205,000
Health insurance
401(k) matching
Paid time off
+1
Staff Site Reliability Engineer: Lead Reliability at Global Scale
Staff Site Reliability Engineer: Lead Reliability at Global Scale

Attentive • United States

On-site
USD 180,000 - 240,000
Health benefits
Equity
Flexible work
Senior Site Reliability Lead - Cloud & Automation
Senior Site Reliability Lead - Cloud & Automation

NetApp • Morrisville (NC)

Hybrid
USD 170,000 - 253,000
Health Insurance
Life Insurance
Retirement or Pension Plans
+5