Manager, Software Engineering DevOps

theocc

United States

On-site

USD 120,000 - 160,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

theocc is seeking a Manager, Environment Operations to lead L1/L2 support for production environments across deployments, middleware, and platform infrastructure. You will drive incident response quality, SLA governance, alert reduction, and MTTR improvement while delivering actionable metrics to leadership.

You will supervise 6-10 engineers, manage 24x7 support coverage, and ensure documentation discipline, runbooks, and knowledge management for ongoing reliability.

Responsibilities

  • Lead L1 and L2 support engineers in incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting.
  • Oversee technical analysis of environment incidents across deployments, middleware, and platform layers, coordinating with internal teams.
  • Serve as Tier 3 escalation point for complex incidents across Platform, S&I, Security, and App Dev teams.
  • Own the full incident lifecycle from first alert to RCA documentation and fix or workaround.
  • Drive post-incident reviews for all P1 and P2 incidents to identify root causes and action items.

Job description

What You'll Do:

The Manager, Environment Operations (EnvOps) will lead and optimize the L1 and L2 support functions responsible for full production-level environment support across application deployments, middleware, and platform infrastructure. This role drives incident response quality, implements SLA governance, leads alert reduction and MTTR improvement initiatives, and delivers actionable operational metrics to leadership. The Manager ensures the team operates at a high standard of readiness - triaging, resolving, and escalating environment issues with speed, accuracy, and full documentation discipline.

Primary Duties and Responsibilities:

To perform this job successfully, an individual must be able to perform each primary duty satisfactorily.

Incident Management & Environment Support:
  • Lead L1 and L2 support engineers in all incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting.
  • Oversee technical analysis of environment incidents across application deployments, middleware, and platform layers while coordinating response activities with internal engineering, platform, and application development teams.
  • Serve as Tier 3 escalation point for complex incidents beyond L2 capability - triaging, directing, and driving resolution across Platform (k8s, Kafka, TFE), S&I (deployment, middleware, storage, network), Security (Vault, certs, secrets), and App Dev teams.
  • Own the full incident lifecycle - from first alert through to RCA documentation and permanent fix or accepted workaround.
  • Drive post-incident reviews for all P1 and P2 incidents, ensuring root cause is identified, documented, and actioned - not filed.
SLA Governance & Operational Standards:
  • Define, publish, and enforce SLA targets across all severity levels:
  • P1 (Critical): 15-minute response, 2-hour resolution
  • P2 (High): 30-minute response, 4-hour resolution
  • P3 (Medium): 2-hour response, 8-hour resolution
  • P4 (Low): 4-hour response, 24-hour resolution
  • Monitor SLA compliance in real time; elevate breaches immediately and report trends to leadership on a sprint cadence.
  • Ensure incident priority classifications are accurate, consistent, and applied at point of triage - not revised post-resolution.
  • Publish monthly SLA compliance reports to leadership with trend analysis and improvement actions.
Alert Tuning & MTTR Improvement:
  • Lead alert tuning and noise reduction initiatives across monitoring toolsets - on-call engineers are paged for situations requiring human judgement, not system noise.
  • Track and publish the alert-to-incident ratio each sprint; hold the team accountable to a visible and improving trend.
  • Drive continuous improvement on Mean Time to Resolve (MTTR) - analyzed by environment, severity, and team tier; reported quarterly with commentary on trend and action.
  • Identify recurring incident patterns and open Problem records; drive permanent fixes, not repeated workarounds.
  • Lead automation and tooling initiatives to reduce toil, accelerate triage, and eliminate manual steps from the support workflow.
Documentation & Knowledge Management:
  • Ensure accurate, complete documentation for every incident - symptoms, steps taken, diagnostics, resolution, and RCA where applicable.
  • Own the runbook library - every novel resolution produces a runbook published to L1 before the incident is closed; coverage gaps are tracked and closed sprint-on-sprint.
  • Maintain environment configuration documentation and operational procedures in a current, accessible, and team-reviewed state.
Supervisory Responsibilities:
  • Lead a team of 6-10 L1 and L2 support engineers and contingent labor within the Environment Operations function.
  • Manage team scheduling to ensure full coverage of production support windows including on-call rotations, shift handoffs, and escalation availability for 24×7 support responsibilities.
  • Perform all talent management functions including performance reviews, direct and timely feedback, goal setting, and administrative functions as required.
  • Confer with and advise team members on operational policies and procedures, technical priorities, escalation paths, and resolution methods.
  • Promote employee development through structured career-planning sessions, identification of training opportunities, and scheduling of relevant conferences, certifications, and skill-building programs.
  • Build and maintain a clear L1 to L2 career progression path - internal promotion is the first option for L2 vacancies.
  • Review scheduling, training plans, and team capacity with senio
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, Software Engineering DevOps
Manager, Software Engineering DevOps

The Options Clearing Corporation • Chicago (IL)

Hybrid
USD 120,000 - 160,000
Hybrid work up to 2 days remote
Tuition Reimbursement
Student Loan Repayment Assistance
+4
Environment Ops Lead - Incident & MTTR Champion
Environment Ops Lead - Incident & MTTR Champion

theocc • United States

On-site
USD 120,000 - 160,000
Environment Management Engineer
Environment Management Engineer

Veriipro • Boston (MA)

On-site
USD 70,000 - 90,000
EnvOps Manager: Production Reliability & Incident Response
EnvOps Manager: Production Reliability & Incident Response

The Options Clearing Corporation • Chicago (IL)

Hybrid
USD 120,000 - 160,000
Hybrid work up to 2 days remote
Tuition Reimbursement
Student Loan Repayment Assistance
+4
Manager, Hosting Service Delivery
Manager, Hosting Service Delivery

Jobtailor • Wayne (PA)

On-site
USD 110,000 - 170,000
Systems Operations and Engineering Manager
Systems Operations and Engineering Manager

LOOP • Greenville (SC), Spartanburg (SC), Anderson (SC)

On-site
USD 120,000 - 160,000
Manager of IT Operations
Manager of IT Operations

Midland Industries • Kansas City (MO)

On-site
USD 120,000 - 180,000
Environment Manager
Environment Manager

Compunnel, Inc. • Atlanta (GA)

On-site
USD 100,000 - 130,000
Technical Platform Operations Support, Manager
Technical Platform Operations Support, Manager

Jobtailor • Town of Texas (WI)

On-site
USD 120,000 - 180,000
Manager, DevOps
Manager, DevOps

1 O.C. Tanner Company • Salt Lake City (UT)

On-site
USD 120,000 - 150,000