Incident Support Engineer

ClifyX

California (MO)

On-site

USD 110,000 - 150,000

Full time

34 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

ClifyX in Irvine, CA is seeking an Incident Support Engineer to own end‑to‑end incident handling in a 24/7 IoT/connected car operations program. You will monitor health using Dynatrace, Datadog, Grafana, and MaxGauge, elevate anomalies, run incident bridges, and drive cross‑functional teams to resolution.

The role requires a strong background in Telematics, Remote Services, and OEM collaborations, with ITIL practices and solid communication to keep SLAs and customers informed.

Qualifications

  • 5+ years in Operations/Service Management with 4+ years leading Incident Management in large-scale, 24/7 environments.
  • Demonstrated experience running bridges for P1/P0 incidents with incident commander skills.
  • Strong background in Telematics/Connected Car domain and vehicle remote services concepts.
  • Proficiency with observability & monitoring tools: Dynatrace, Datadog, Grafana, MaxGauge.
  • Working knowledge of ITIL practices (Incident, Problem, Change, SLA Management).
  • Hands-on experience with JIRA, Confluence, Jenkins, XMatters.
  • Excellent communication and stakeholder management across global teams.

Responsibilities

  • Monitor health and availability of production systems in Connected Car/Telematics ecosystem.
  • Use monitoring tools to identify abnormal behavior, degradation, errors, and outages.
  • Analyze dashboards, metrics, and logs to scope incidents and impacts.
  • Run proactive monitoring and validate alerts to reduce noise.
  • Lead incident bridges/war rooms and coordinate cross-functional responders.
  • Ensure adherence to IM SLAs and post-incident governance (RCA, CAPA).
  • Define and maintain IM dashboards and executive reports.

Skills

Incident management
Observability
Dynatrace
Datadog
Grafana
MaxGauge
ITIL
JIRA
Confluence
Jenkins
XMatters
Stakeholder management
Communication

Education

Bachelor’s degree in engineering or related field

Tools

Jira
Confluence
Jenkins
XMatters
ServiceNow

Job description

Position Title: Incident Support Engineer with an automotive background.

Location: Irvine, CA

Duration: 12+Months Contract

Role Summary

We are seeking Support Engineer role for a 24/7 Incident Management (IM) operations program supporting a key player in the automotive industry's connected car space. This role will primarily focus on monitoring the system health using different data monitoring tools such as Datadog, Grafana, MaxGauge, Elastic, Dynatrace. Once an anomaly or unusual pattern is observed, elevate it to the respective teams, initiate an incident bridge, drive the bridge, take incident timeline, facilitate the teams to bring an ongoing incident to closure and ensuring adherence to contractual SLAs. The IM Support Engineer will orchestrate end-to-end incident handling—from anomaly detection to bridge initiation, stakeholder communications, resolution, and post-incident governance—with a strong background in Telematics and Connected Vehicle systems (including Remote Services).

Key Responsibilities
  • Monitor the health and availability of production systems supporting the Connected Car / Telematics ecosystem.
  • Proactively monitor application, infrastructure, API, database, and service health using tools such as:
  • Application and system logs
  • Organization-specific monitoring and observability tools
  • Identify abnormal system behavior, performance degradation, errors, latency, service failures, and potential outages.
  • Analyze dashboards, alerts, metrics, logs, and application behavior to identify the initial scope and impact of incidents.
  • Perform proactive monitoring to identify issues before they impact customers.
  • Validate alerts and distinguish between genuine production incidents and false positives/noise.
  • Adhere to the IM SLAs (MTTD, MTTA, MTTR, communication SLAs); ensure measurement, reporting, and continuous improvement.
  • Should be proficient in handling the incidents ranging from high priority ones (P1, P2) to the low priority incidents (P3, P4).
  • Initiate and run incident bridges/war rooms; coordinate cross-functional responders (application, infrastructure, network, OEM partners, and third parties).
  • Oversee proactive detection through data monitoring tools (Datadog, Dynatrace, Grafana, MaxGauge) and ensure alert quality, runbooks, and signal-to-noise optimization.
  • Establish and enforce SOPs for incident declaration, severity classification, response roles and decision logs.
  • Drive disciplined stakeholder communications: timely updates to product, operations, OEM/customer contacts, leadership, and impacted regions; maintain comms cadence and channels.
  • Ensure post-incident governance: facilitate RCA, document contributing causes, corrective and preventive actions (CAPA), and circulate the RCA across teams for sign-off.
  • Define and maintain IM dashboards, KPIs, and executive reports; present weekly/monthly service reviews with trend analysis and action plans.
  • Collaborate with Product/Engineering to influence reliability roadmaps (resiliency patterns, observability, capacity, release safeguards).
  • Ensure compliance with information security, data privacy, and OEM contractual obligations during incident handling and communications.
  • Continuously refine IM playbooks, runbooks, and training; conduct simulations/game days and readiness audits across onsite–offshore teams.
  • Hands‑on knowledge of Telematics Control Unit (TCU), eSIM/OTA provisioning, backend telematics platforms, and data flows between vehicle, cloud, and mobile apps.
  • Understanding on Internet of Things including but not limited to Messaging Queues, bulk provisioning, notification systems, API Gateways, Load Balancers, Mobile applications & web applications.
  • Familiarity with Remote Services (e.g., remote lock/unlock, start/stop, charge control, climate pre‑conditioning), geo‑services, and safety/assist features.
  • Understanding of service dependencies: identity/auth, messaging, device management, CAN bus signals, firmware/OTA update orchestration, and regional compliance.
  • Experience coordinating incidents across OEM partners, Tier‑1 suppliers, cloud providers, and customer support operations.
Required Qualifications
  • 5+ years in Operations/Service Management with 4+ years leading Incident Management in large‑scale, 24/7 environments.
  • Demonstrated experience running bridges for P1/P0 incidents; proven incident commander skills and decision‑making under pressure.
  • Strong background in Telematics/Connected Car domain and vehicle remote services concepts.
  • Proficiency with observability and monitoring tools: Dynatrace, Datadog, Grafana, MaxGauge (dashboards, alerting, traces, logs).
  • Working knowledge of ITIL practices (Incident, Problem, Change, Service Level Management).
  • Should have good hands on experience in using tools like JIRA, Confluence, Jenkins, XMatters.
  • Excellent communication (written/verbal), executive presence, and stakeholder management across onsite–offshore teams.
  • Ability to analyse telemetry and time‑series data to drive root cause hypotheses and corrective actions.
  • Experience defining SLAs/SLOs and building KPI dashboards (MTTD/MTTA/MTTR, incident volume, recurrence, comms SLA, customer impact).
Tools & Technologies
  • Incident & Comms: ticketing (Jira/ServiceNow), chat/bridge tools (Teams/Zoom), status pages.
  • Data: time‑series analysis, log aggregation, tracing, alerting policies and noise reduction.
Education & Experience
  • Bachelor’s degree in engineering, Computer Science, or related field (or equivalent practical experience).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

24/7 Incident Commander — Connected Car & Telematics
24/7 Incident Commander — Connected Car & Telematics

ClifyX • California (MO)

On-site
USD 110,000 - 150,000
Incident Support Engineer (Automotive Domain) role
Incident Support Engineer (Automotive Domain) role

Pyramid Consulting, Inc • Fort Mill (SC)

On-site
USD 143,270,000 - 148,781,000
Health insurance
401(k) plan
Paid sick leave
Sr Observability Engineer
Sr Observability Engineer

IT Associates • Irvine (CA)

Hybrid
USD 150,000 - 210,000
Senior Observability Engineer — AI-Driven Reliability
Senior Observability Engineer — AI-Driven Reliability

IT Associates • Irvine (CA)

On-site
Incident Manager
Incident Manager

Insight Global • United States

On-site
USD 34,000 - 36,000
EOC Incident Manager - Watch Officer
EOC Incident Manager - Watch Officer

Dunhill Professional Search & Government Solutions • Ashburn (VA)

On-site
USD 110,000 - 150,000
Incident Response Lead: Automotive Telematics
Incident Response Lead: Automotive Telematics

Pyramid Consulting, Inc • Fort Mill (SC)

On-site
USD 143,270,000 - 148,781,000
Health insurance
401(k) plan
Paid sick leave
Application Engineer View role →
Application Engineer View role →

NRnP Technology • Northern (KY)

Hybrid
USD 90,000 - 140,000
Operations Admin Specialist - Senior
Operations Admin Specialist - Senior

Spectraforce Technologies • Southlake (TX)

Hybrid
USD 68,000 - 102,000
Application Monitoring Engineer | RazorPe Innovations LLC
Application Monitoring Engineer | RazorPe Innovations LLC

Razorpe • Arlington Heights (IL)

On-site
USD 90,000 - 120,000