Service Reliability & Incident Analyst

Oakridge Staffing

New York (NY)

On-site

USD 110,000 - 140,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Oakridge Staffing partners with a US-based multi-strategy hedge fund to build an Technology Operations Center team. The role leads operational monitoring, incident management, and service recovery, reporting to senior stakeholders during disruptions.

You will own runbooks, escalation matrices, and KPI/SLA compliance, driving continual operational improvements across engineering and support teams. This is not an engineering role but a hands-on operations leadership position.

Qualifications

  • Experience in IT Service Operations, Production Support, or NOC.
  • Knowledge of ITIL or other IT service management frameworks.
  • Experience with observability platforms such as Datadog.

Responsibilities

  • Own operational monitoring, incident management, and service recovery processes.
  • Lead major incident response and resolution activities across engineering and support teams.
  • Manage escalation paths and maintain clear, timely communications with senior stakeholders and business partners during service disruptions.
  • Define, track, and report on service availability, operational KPIs, and SLA compliance.
  • Review recurring incidents to identify root causes and drive operational improvement initiatives.
  • Ensure runbooks, escalation matrices, and operational procedures are current and accurate.

Skills

IT Service Operations
Production Support
NOC experience

Tools

Datadog

Job description

Our multi-strategy hedge fund client has come to us exclusively to build a US-based Technology Operations Center team for their growing Systematic Trading business. The team will start with a lead and two team members. This is not an engineering role. Beyond the technical skills listed below, much of what we’re looking for in this hire is intangible. We seek candidates who demonstrate strong attention to detail, a sense of urgency, and excellent communication skills.

Big picture of what you will do:

• Own operational monitoring, incident management, and service recovery processes.

• Lead major incident response and resolution activities across engineering and support teams.

• Manage escalation paths and maintain clear, timely communications with senior stakeholders and business partners during service disruptions.

• Define, track, and report on service availability, operational KPIs, and SLA compliance.

• Review recurring incidents to identify root causes and drive operational improvement initiatives.

• Ensure runbooks, escalation matrices, and operational procedures are current and accurate.

Skills you should come in with:

• Experience in a role such as IT Service Operations, Production Support, or Network Operations Center (NOC).

• Knowledge of ITIL or other IT service management frameworks.

• Experience with observability platforms such as Datadog.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Systems Operations Analyst
Lead Systems Operations Analyst

Oakridge Staffing • New York (NY)

On-site
USD 110,000 - 140,000
Incident Manager
Incident Manager

Insight Global • Jersey City (NJ)

On-site
USD 110,000 - 140,000
Senior Manager, Systems Operations
Senior Manager, Systems Operations

ICE Clear Europe Limited • Atlanta (GA)

On-site
USD 150,000 - 210,000
Operations Engineer - Triage Operations
Operations Engineer - Triage Operations

Perennial Resources International • Totowa (NJ)

On-site
USD 120,000 - 160,000
Senior Systems Operations Engineer - Incident & Reliability
Senior Systems Operations Engineer - Incident & Reliability

thetradedesk • Chicago (IL)

On-site
USD 120,000 - 170,000
Healthcare (medical, dental, vision)
401k plan with company match
Employee stock purchase plan
+1
Operations Analyst
Operations Analyst

Encore Technologies • Cincinnati (OH)

On-site
USD 70,000 - 90,000
Data Center Incident Response Project Coordinator
Data Center Incident Response Project Coordinator

Astreya • Austin (TX)

On-site
USD 90,000 - 130,000
Site Reliability Engineer – AWS, Azure, IaC, Typescript, .NET
Site Reliability Engineer – AWS, Azure, IaC, Typescript, .NET

Stott and May • New York (NY)

Hybrid
USD 120,000 - 140,000
Technology Support Lead - Major Incident Management
Technology Support Lead - Major Incident Management

JPMorgan Chase & Co. • Columbus (OH)

On-site
USD 110,000 - 170,000
Senior Manager, Systems Operations
Senior Manager, Systems Operations

ICE • Atlanta (GA)

On-site
USD 140,000 - 190,000