Manager - Software Engineering DevOps

Request Technology, LLC

Chicago (IL)

Hybrid

USD 170,000 - 200,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Request Technology, LLC in Chicago is seeking a Manager, Software Engineering DevOps to lead a team of 6-10 engineers and oversee 24/7 support across L1–L3. You will manage incident response, deployment pipelines, and platform operations within a hybrid on-site/remote work model.

The role commands a salary of $170,000–$200,000 with a 15% bonus and requires no sponsorship or OPT. Hybrid schedule: 3 days onsite, 2 days remote.

Qualifications

  • Minimum 5 years of hands-on environment operations, production support, or infrastructure operations experience.
  • Experience across deployment pipelines, container platforms, middleware, incident management, monitoring, automation, and platform operations.

Responsibilities

  • Lead L1 and L2 support engineers in incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting.
  • Oversee analysis of environment incidents across deployments, middleware, and platform layers and coordinate responses with internal teams.
  • Serve as Tier 3 escalation point for complex incidents across Platform (k8s, Kafka, TFE), S&I, Security, and App Dev teams.
  • Own the full incident lifecycle from first alert to RCA and permanent fix or workaround.
  • Drive post-incident reviews for all P1 and P2 incidents ensuring root cause is identified and addressed.

Skills

Leadership
Incident Management
SLA Frameworks
Team Leadership

Tools

Kubernetes
Kafka
Vault
Splunk
Dynatrace
Datadog
Jenkins
GitHub
ServiceNow

Job description

NO SPONSORSHIP - NO OPT

Manager, Software Engineering DevOps

Salary:$170k - $200k plus 15% Bonus

LOCATION: Chicago, IL

Hybrid 3 days onsite and 2 days remote

You will lead a team of 6-10, 24/7 support, L1, L2 support. Application deployment container platform ops middleware support incident management monitoring observability configuration release engineering platform operations scripting automation SLA frameworks harness Jenkins Kubernetes apache kakfa splunk Dynatrace datadog.

  • Lead L1 and L2 support engineers in all incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting.
  • Oversee technical analysis of environment incidents across application deployments, middleware, and platform layers while coordinating response activities with internal engineering, platform, and application development teams.
  • Serve as Tier 3 escalation point for complex incidents beyond L2 capability triaging, directing, and driving resolution across Platform (k8s, Kafka, TFE), S&I (deployment, middleware, storage, network), Security (Vault, certs, secrets), and App Dev teams.
  • Own the full incident lifecycle from first alert through to RCA documentation and permanent fix or accepted workaround.
  • Drive post-incident reviews for all P1 and P2 incidents, ensuring root cause is identified, documented, and actioned not filed.
Technical Skills:
  • Deployment & Pipeline tooling: Harness (continuous delivery pipelines, deployment verification, rollback automation), Jenkins (CI/CD pipeline management, job configuration, build troubleshooting), GitHub (branching strategies, pull request workflows, pipeline integration).
  • Container & orchestration platforms: Kubernetes (k8s) pod lifecycle management, namespace operations, log retrieval, resource troubleshooting, and coordination with Platform teams on cluster-level issues.
  • Messaging & streaming platforms: Apache Kafka topic management, consumer group monitoring, lag analysis, and escalation to Platform for broker-level issues.
  • Secrets & configuration management: HashiCorp Vault secrets retrieval, token/lease troubleshooting, policy review, and escalation to Security teams for certificate and secrets rotation.
  • Monitoring & observability: Proficiency in at least two production monitoring toolsets (e.g. Splunk, Dynatrace, Datadog, AppDynamics, PrometheGrafana) alert triage, dashboard interpretation, log analysis, and tuning requests.
  • Middleware platforms: Working knowledge of middleware infrastructure including application servers, messaging brokers, storage integrations, and network-layer dependencies sufficient to triage, gather diagnostics, and route correctly to L3.
  • Incident and ticketing platforms: ServiceNow or equivalent ITSM tooling incident creation, SLA tracking, problem record management, and reporting.
  • MTTR and operational metrics: Ability to build and maintain operational dashboards and reports covering MTTR, SLA compliance, alert-to-incident ratio, repeat incident rate, and deployment success rate.
Education and/or Experience:
  • Minimum 5 years of hands-on environment operations, production support, or infrastructure operations experience, including interdisciplinary experience across four or more of the following: application deployment pipelines, container platform operations, middleware support, incident management, monitoring and observability, configuration management, release engineering, platform operations, or scripting and automation.
  • Technical experience and comprehensive knowledge of production environment failure modes including deployment failures, configuration drift, platform instability, and integration breakdowns and the methodologies used to diagnose and resolve them.
  • Demonstrated experience defining and enforcing SLA frameworks in a tiered support model (L1/L2/L3 or equivalent).
  • Familiarity with financial services or other regulated-industry production environments is a strong advantage understanding of change governance, audit requirements, and production access controls.
  • Industry knowledge of current and emerging practices in environment operations, platform reliability, and support automation.
  • Shift work and on-call availability required including 247 on-call response capacity and availability during planned and emergency maintenance windows.
  • Previous people management or team lead experience required; formal people management experience
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager - Software Engineer DevOps
Manager - Software Engineer DevOps

Request Technology, LLC • Chicago (IL)

Hybrid
USD 140,000 - 200,000
Manager of DevOps Operational Support
Manager of DevOps Operational Support

Request Technology, LLC • Chicago (IL)

On-site
USD 140,000 - 190,000
Bonus eligible
Manager, DevSecOps Production Support
Manager, DevSecOps Production Support

Request Technology, LLC • Chicago (IL)

On-site
USD 120,000 - 180,000
Manager, Software Engineering DevOps
Manager, Software Engineering DevOps

The Options Clearing Corporation • Chicago (IL)

Hybrid
USD 120,000 - 160,000
Hybrid work up to 2 days remote
Tuition Reimbursement
Student Loan Repayment Assistance
+4
DevOps Platform Manager — Lead 24/7 Incident Response
DevOps Platform Manager — Lead 24/7 Incident Response

Request Technology, LLC • Chicago (IL)

Hybrid
USD 170,000 - 200,000
IT CONSULTANT SR
IT CONSULTANT SR

First Horizon Corp. • Memphis (TN), Northern (KY)

Hybrid
USD 120,000 - 160,000
Hybrid DevOps Manager: Incident Lead & Platform Reliability
Hybrid DevOps Manager: Incident Lead & Platform Reliability

Request Technology, LLC • Chicago (IL)

Hybrid
USD 140,000 - 200,000
Manager, DevOps
Manager, DevOps

1 O.C. Tanner Company • Salt Lake City (UT)

On-site
USD 120,000 - 150,000
Software Engineering Group Manager – Site Reliability
Software Engineering Group Manager – Site Reliability

Jobtailor • Alabama

On-site
USD 120,000 - 180,000
Lead Platform Engineer
Lead Platform Engineer

The Sherwin-Williams Company • Cleveland (OH)

On-site
USD 110,000 - 140,000