Manager, Software Engineering DevOps

The Options Clearing Corporation

Chicago (IL)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work up to 2 days remote
Tuition Reimbursement
Student Loan Repayment Assistance
Technology Stipend
Generous PTO and Parental leave
401k Employer Match
Competitive health benefits

Job summary

The Options Clearing Corporation (OCC) is seeking a Manager, Environment Operations to lead L1–L2 production support across deployments, middleware, and platform infrastructure. You will drive incident response, SLA governance, alert tuning, and MTTR improvements while ensuring ready, well-documented runbooks and handoffs.

You will manage 6–10 engineers, collaborate across Platform, Security, and App Dev teams, and report on incident trends and operational metrics to leadership.

Qualifications

  • Minimum 5 years of hands-on environment/production support experience across multiple domains.
  • Experience across deployment pipelines, monitoring/observability, incident management, and automation.
  • Familiarity with SLA frameworks in tiered support models (L1/L2/L3).
  • Experience in financial services or regulated environments is a plus.

Responsibilities

  • Lead L1 and L2 incident response across production environments and platforms.
  • Coordinate cross-team resolution and RCA reporting after incidents.
  • Define, publish, and enforce SLA targets across severity levels and publish monthly reports.
  • Lead alert tuning and MTTR reduction initiatives to improve reliability.
  • Own runbooks and ensure documentation is updated for every incident.

Skills

Team leadership
Incident management
SLA governance
MTTR improvement
Monitoring & observability
Communication skills
Cross-team collaboration

Education

Experience in environment operations

Tools

Kubernetes
Jenkins
HashiCorp Vault
Splunk
Datadog
Prometheus/Grafana
ServiceNow
Terraform

Job description

About OCC

The Options Clearing Corporation (OCC) is the world's largest equity derivatives clearing organization. Founded in 1973, OCC is dedicated to promoting stability and market integrity by delivering clearing and settlement services for options, futures and securities lending transactions. As a Systemically Important Financial Market Utility (SIFMU), OCC operates under the jurisdiction of the U.S. Securities and Exchange Commission (SEC), the U.S. Commodity Futures Trading Commission (CFTC), and the Board of Governors of the Federal Reserve System. OCC has more than 100 clearing members and provides central counterparty (CCP) clearing and settlement services to 19 exchanges and trading platforms. More information about OCC is available at www.theocc.com.

Benefits

  • A hybrid work environment, up to 2 days per week of remote work
  • Tuition Reimbursement to support your continued education
  • Student Loan Repayment Assistance
  • Technology Stipend allowing you to use the device of your choice to connect to our network while working remotely
  • Generous PTO and Parental leave
  • 401k Employer Match
  • Competitive health benefits including medical, dental and vision
Position Overview

The Manager, Environment Operations (EnvOps) will lead and optimize the L1 and L2 support functions responsible for full production-level environment support across application deployments, middleware, and platform infrastructure. This role drives incident response quality, implements SLA governance, leads alert reduction and MTTR improvement initiatives, and delivers actionable operational metrics to leadership. The Manager ensures the team operates at a high standard of readiness — triaging, resolving, and escalating environment issues with speed, accuracy, and full documentation discipline.

Primary Duties and Responsibilities
  • Incident Management & Environment Support: Lead L1 and L2 support engineers in all incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting. Oversee technical analysis of environment incidents across application deployments, middleware, and platform layers while coordinating response activities with internal engineering, platform, and application development teams. Serve as Tier 3 escalation point for complex incidents beyond L2 capability — triaging, directing, and driving resolution across Platform (k8s, Kafka, TFE), S&I (deployment, middleware, storage, network), Security (Vault, certs, secrets), and App Dev teams. Own the full incident lifecycle — from first alert through to RCA documentation and permanent fix or accepted workaround. Drive post-incident reviews for all P1 and P2 incidents, ensuring root cause is identified, documented, and actioned — not filed.
  • SLA Governance & Operational Standards: Define, publish, and enforce SLA targets across all severity levels: P1 (Critical): 15-minute response, 2-hour resolution; P2 (High): 30-minute response, 4-hour resolution; P3 (Medium): 2-hour response, 8-hour resolution; P4 (Low): 4-hour response, 24-hour resolution. Monitor SLA compliance in real time; escalate breaches immediately and report trends to leadership on a sprint cadence. Ensure incident priority classifications are accurate, consistent, and applied at point of triage — not revised post-resolution. Publish monthly SLA compliance reports to leadership with trend analysis and improvement actions.
  • Alert Tuning & MTTR Improvement: Lead alert tuning and noise reduction initiatives across monitoring toolsets — on-call engineers are paged for situations requiring human judgement, not system noise. Track and publish the alert-to-incident ratio each sprint; hold the team accountable to a visible and improving trend. Drive continuous improvement on Mean Time to Resolve (MTTR) — analyzed by environment, severity, and team tier; reported quarterly with commentary on trend and action. Identify recurring incident patterns and open Problem records; drive permanent fixes, not repeated workarounds. Lead automation and tooling initiatives to reduce toil, accelerate triage, and eliminate manual steps from the support workflow.
  • Documentation & Knowledge Management: Ensure accurate, complete documentation for every incident — symptoms, steps taken, diagnostics, resolution, and RCA where applicable. Own the runbook library — every novel resolution produces a runbook published to L1 before the incident is closed; coverage gaps are tracked and closed sprint-on-sprint. Maintain environment configuration documentation and operational procedures in a current, accessible, and team-reviewed state.
Supervisory Responsibilities
  • Lead a team of 6–10 L1 and L2 support engineers and contingent labor within the Environment Operations function. Manage team scheduling to ensure full coverage of production support windows including on-call rotations, shift handoffs, and escalation availability for 24×7 support responsibilities.
  • Perform all talent management functions including performance reviews, direct and timely feedback, goal setting, and administrative functions as required. Confer with and advise team members on operational policies and procedures, technical priorities, escalation paths, and resolution methods.
  • Promote employee development through structured career-planning sessions, identification of training opportunities, and scheduling of relevant conferences, certifications, and skill-building programs. Build and maintain a clear L1 to L2 career progression path — internal promotion is the first option for L2 vacancies. Review scheduling, training plans, and team capacity with senior leadership on a regular cadence to ensure the team is resourced and developed for the demands of the role.
Qualifications
  • Proven team leadership experience — taking initiative, driving follow-through, and holding a team accountable to defined standards across varying incident types and urgencies.
  • Demonstrated ability to operate as a strong team player — working across L1, L2, Platform, Security, and App Dev teams on both long-running improvement programs and rapid incident response under tight deadlines.
  • Deep experience in production environment supports hands-on knowledge of application deployment pipelines, container orchestration, messaging platforms, and middleware infrastructure.
  • Ability to create, tune, and maintain monitoring alerts and operational runbooks independently.
  • Effective and excellent oral and written communication skills — able to translate technical incident detail into clear, concise leadership reporting without loss of accuracy.
  • Strong analytical, judgement, and consultation skills — able to triage ambiguous situations, make sound decisions under pressure, and consult effectively across technical and non-technical stakeholders.
  • Ability to work independently and manage multiple parallel priorities with strong organizational discipline.
Technical Skills
  • Deployment & Pipeline tooling: Harness, Jenkins, GitHub
  • Container & orchestration platforms: Kubernetes
  • Messaging & streaming platforms: Apache Kafka
  • Secrets & configuration management: HashiCorp Vault
  • Monitoring & observability: Splunk, Dynatrace, Datadog, AppDynamics, Prometheus/Grafana
  • Middleware platforms: application servers, messaging brokers, storage integrations
  • Incident and ticketing platforms: ServiceNow or equivalent ITSM tooling
  • MTTR and operational metrics: dashboards and reports
Education and/or Experience

Minimum 5 years of hands-on environment operations, production support, or infrastructure operations experience, including interdisciplinary experience across four or more of the following: application deployment pipelines, container platform operations, middleware support, incident management, monitoring and observability, configuration management, release engineering, platform operations, or scripting and automation. Technical experience and comprehensive knowledge of production environment failure modes — including deployment failures, configuration drift, platform instability, and integration breakdowns — and the methodologies used to diagnose and resolve them. Demonstrated experience defining and enforcing SLA frameworks in a tiered support model (L1/L2/L3 or equivalent). Familiarity with financial services or other regulated-industry production environments is a strong advantage — understanding of change governance, audit requirements, and production access controls. Industry knowledge of current and emerging practices in environment operations, platform reliability, and support automation. Shift work and on-call availability required — including 24×7 on-call response capacity and availability during planned and emergency maintenance windows. Previous people management or team lead experience required; formal people management experience strongly preferred.

Certifications

The following are considered advantageous and will be recognized in the hiring process: Certified Kubernetes Administrator (CKA) or equivalent platform certification ITIL Foundation or above — demonstrating grounding in incident, problem, and change management frameworks HashiCorp Vault Associate Harness Certified Continuous Delivery Architect Relevant cloud platform certification (AWS, Azure, or GCP) at associate level or above

About Our Recruitment Process

The text above has been refined for clarity and formatting while maintaining core details. For full original content and protections, please refer to the source job posting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, Software Engineering DevOps
Manager, Software Engineering DevOps

theocc • United States

On-site
USD 120,000 - 160,000
Environment Ops Lead - Incident & MTTR Champion
Environment Ops Lead - Incident & MTTR Champion

theocc • United States

On-site
USD 120,000 - 160,000
Principal, Software Engineering: Java
Principal, Software Engineering: Java

theocc • United States

On-site
USD 180,000 - 260,000
Executive Director, Platform Governance & Strategy
Executive Director, Platform Governance & Strategy

The Options Clearing Corporation • Chicago (IL)

Hybrid
USD 150,000 - 200,000
Hybrid work environment
Tuition reimbursement
Student loan repayment assistance
+4
Lead Associate Principal, Cloud Engineering
Lead Associate Principal, Cloud Engineering

The Options Clearing Corporation (OCC) • Illinois

Hybrid
USD 143,000 - 229,000
Hybrid work environment, up to 2 days remote
Tuition Reimbursement
Student Loan Repayment Assistance
+4
Systems Operations and Engineering Manager
Systems Operations and Engineering Manager

LOOP • Greenville (SC), Spartanburg (SC), Anderson (SC)

On-site
USD 120,000 - 160,000
Manager, Cyber Defense
Manager, Cyber Defense

The Options Clearing Corporation (OCC) • Illinois

Hybrid
USD 132,000 - 219,000
Hybrid work
Tuition reimbursement
Student loan repayment
+4
Manager-Cloud Operations
Manager-Cloud Operations

WellSpan Health • York

On-site
USD 90,000 - 120,000
Comprehensive health benefits
Retirement savings plan
Paid time off (PTO)
+3
Executive Director, Platform Architecture
Executive Director, Platform Architecture

National Black MBA Association • Chicago (IL)

Hybrid
USD 198,000 - 347,000
Hybrid work environment
Tuition Reimbursement
Student Loan Repayment Assistance
+4
Manager of Infrastructure
Manager of Infrastructure

OculusIT • Newton (MA)

On-site
USD 120,000 - 170,000