Manager, DevSecOps Production Support

Request Technology, LLC

Chicago (IL)

On-site

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Request Technology, LLC is seeking an experienced Platform Operations Lead in Chicago to steer a 6-10-person L1/L2 team responsible for production support, incident management, and SLA enforcement. The role oversees runbooks, alert tuning, automation initiatives, and cross-team incident resolution to ensure reliable service delivery.

The position requires strong experience in regulated environments, deployment pipelines, Kubernetes, Kafka, and Vault, with a focus on reducing toil and improving

Qualifications

  • 5+ years in hands-on environment operations or production support.
  • Experience across deployment pipelines, container platforms, middleware, incident management, monitoring, release engineering, or scripting.
  • Defined and enforced SLA frameworks in tiered support models (L1/L2/L3).
  • Familiar with regulated production environments, change governance, and access controls.

Responsibilities

  • Lead a team of 6-10 L1 and L2 support engineers.
  • Manage scheduling for 24x7 production support coverage and on-call rotations.
  • Conduct talent management: performance reviews, feedback, and goals.
  • Ensure incident documentation: symptoms, diagnostics, resolution, RCA.
  • Maintain runbook library and publish resolutions to L1.
  • Lead alert tuning and noise reduction across monitoring tools.
  • Drive automation and tooling to reduce toil and speed triage.
  • Define, publish, and enforce SLA targets.
  • Monitor SLA compliance in real time and escalate breaches.
  • Lead incident response across deployments, middleware, and platform.
  • Escalate complex incidents as Tier 3 across Platform, S&I, Security, App Dev.

Skills

Production operations
Incident management
SLA frameworks
Regulated environments

Tools

Harness
Jenkins
GitHub
Kubernetes
Kafka
HashiCorp Vault
Splunk
Dynatrace
Datadog
AppDynamics
Prometheus
Grafana

Job description

*We are unable to provide sponsorship for this role*

Qualifications
  • 5+ years of hands-on environment operations, production support, or infrastructure operations experience
  • Experience across: application deployment pipelines, container platform operations, middleware support, incident management, monitoring and observability, configuration management, release engineering, platform operations, or scripting and automation.
  • Demonstrated experience defining and enforcing SLA frameworks in a tiered support model (L1/L2/L3 or equivalent).
  • Familiarity with financial services or other regulated-industry production environments including knowledge of change governance, audit requirements, and production access controls.
Technical skillset
  • Harness (continuous delivery pipelines, deployment verification, rollback automation), Jenkins (CI/CD pipeline management, job configuration, build troubleshooting), GitHub (branching strategies, pull request workflows, pipeline integration).
  • Kubernetes (k8s) pod lifecycle management, namespace operations, log retrieval, resource troubleshooting, and coordination with Platform teams on cluster-level issues.
  • Apache Kafka topic management, consumer group monitoring, lag analysis, and escalation to Platform for broker-level issues.
  • HashiCorp Vault - secrets retrieval, token/lease troubleshooting, policy review, and escalation to Security teams for certificate and secrets rotation.
  • Proficiency in at least two production monitoring toolsets (e.g. Splunk, Dynatrace, Datadog, AppDynamics, Prometheus/Grafana) alert triage, dashboard interpretation, log analysis, and tuning requests.
  • Working knowledge of middleware infrastructure including application servers, messaging brokers, storage integrations, and network-layer dependencies sufficient to triage, gather diagnostics, and route correctly to L3.
Responsibilities
  • Lead a team of 6-10 L1 and L2 support engineers
  • Manage team scheduling to ensure full coverage of production support windows including on-call rotations, shift handoffs, and escalation availability for 24x7 support responsibilities.
  • Perform all talent management functions including performance reviews, direct and timely feedback, goal setting, and administrative functions as required.
  • Ensure accurate, complete documentation for every incident — symptoms, steps taken, diagnostics, resolution, and RCA where applicable.
  • Own the runbook library - every novel resolution produces a runbook published to L1 before the incident is closed; coverage gaps are tracked and closed sprint-on-sprint.
  • Lead alert tuning and noise reduction initiatives across monitoring toolsets - on-call engineers are paged for situations requiring human judgement, not system noise.
  • Lead automation and tooling initiatives to reduce toil, accelerate triage, and eliminate manual steps from the support workflow.
  • Define, publish, and enforce SLA targets
  • Monitor SLA compliance in real time; escalation breaches immediately and report trends to leadership on a sprint cadence.
  • Lead L1 and L2 support engineers in all incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting.
  • Oversee technical analysis of environment incidents across application deployments, middleware, and platform layers while coordinating response activities with internal engineering, platform, and application development teams.
  • Serve as Tier 3 escalation point for complex incidents beyond L2 capability - triaging, directing, and driving resolution across Platform (k8s, Kafka, TFE), S&I (deployment, middleware, storage, network), Security (Vault, certs, secrets), and App Dev teams.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager of DevOps Operational Support
Manager of DevOps Operational Support

Request Technology, LLC • Chicago (IL)

On-site
USD 140,000 - 190,000
Bonus eligible
Manager - Software Engineer DevOps
Manager - Software Engineer DevOps

Request Technology, LLC • Chicago (IL)

Hybrid
USD 140,000 - 200,000
Manager - Software Engineering DevOps
Manager - Software Engineering DevOps

Request Technology, LLC • Chicago (IL)

Hybrid
USD 170,000 - 200,000
Manager, Software Engineering DevOps
Manager, Software Engineering DevOps

The Options Clearing Corporation • Chicago (IL)

Hybrid
USD 120,000 - 160,000
Hybrid work up to 2 days remote
Tuition Reimbursement
Student Loan Repayment Assistance
+4
L2/L3 Run Engineer
L2/L3 Run Engineer

Veriipro • Dallas (TX)

On-site
USD 80,000 - 120,000
IT CONSULTANT SR
IT CONSULTANT SR

First Horizon Corp. • Memphis (TN), Northern (KY)

Hybrid
USD 120,000 - 160,000
Senior DevOps Incident & Reliability Manager
Senior DevOps Incident & Reliability Manager

Request Technology, LLC • Chicago (IL)

On-site
USD 140,000 - 190,000
Bonus eligible
Production Support Engineer
Production Support Engineer

NLB Services • Atlanta (GA)

On-site
USD 85,000 - 100,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Software Engineering Group Manager – Site Reliability
Software Engineering Group Manager – Site Reliability

Jobtailor • Alabama

On-site
USD 120,000 - 180,000