Platform Reliability & Incident Response Engineer

Vizient

Missouri

On-site

USD 77,000 - 135,000

Full time

38 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Incentive eligible

Job summary

Vizient in Missouri seeks a production support professional to provide technical triage, diagnostics, and operational support across Vizient's core platforms and business domains. You will improve incident response, restoration speed, and stakeholder communication while partnering with Engineering, Product, and Operations teams to diagnose issues and strengthen platform reliability.

The role requires 2+ years in application/production support, strong SQL/log analysis, cloud-hosted service

Qualifications

  • Relevant degree preferred.
  • 2 or more years of relevant experience required.
  • Experience in application support, production support engineering, software engineering, DevOps, Site Reliability Engineering (SRE), technical systems analysis, monitoring and observability instrumentation, or technical product support required.
  • Strong troubleshooting, SQL/data analysis, log analysis, stack trace analysis, monitoring review, and technical documentation skills required.
  • Experience supporting enterprise applications, cloud-hosted services, integrations, APIs, reporting platforms, or data-driven systems required.
  • Strong verbal and written communication skills with the ability to communicate technical issues clearly to technical and non-technical audiences required.
  • Experience with Databricks, Azure, Spark, data pipelines, observability tools, ServiceNow, Azure DevOps, Python, Splunk, Datadog, or Application Insights preferred.
  • Familiarity with ITSM/ITIL processes, Incident Management, Problem Management, Root Cause Analysis (RCA), corrective action preventative action, or known error management practices preferred.
  • Experience leveraging AI-assisted tools and technologies to improve operational support, incident diagnostics, troubleshooting efficiency, documentation generation, or workflow automation preferred.
  • Experience supporting healthcare technology, analytics platforms, enterprise SaaS platforms, or data-driven systems preferred.

Responsibilities

  • Investigate production issues involving logs, SQL/data patterns, failed jobs, alerts, workflows, APIs, integrations, reporting, access, configuration, and application behavior.
  • Analyze platform dependencies, workflows, monitoring events, and system health indicators to identify suspected causes and resolution paths.
  • Identify workarounds, ownership paths, escalation readiness, and supportable resolution options.
  • Perform Tier 2 triage for production incidents.
  • Create evidence-based escalation summaries and handoffs for Product Engineering, Platform Engineering, or Data Engineering teams.
  • Improve incident intake quality, classification accuracy, ownership clarity, and stakeholder communication.
  • Develop and maintain runbooks, Standard Operating Procedures (SOPs), known error articles, diagnostic checklists, and support documentation.
  • Support Root Cause Analysis (RCA) documentation, Problem Records, corrective action tracking, and recurring incident reduction initiatives.
  • Partner with Product, Engineering, Operations, Support teams, and business stakeholders to improve operational stability and support processes.
  • Reduce avoidable escalations and recurring production issues through improved diagnostics, documentation, and operational follow-through.

Skills

Troubleshooting
SQL/data analysis
Log analysis
Monitoring review
Technical documentation
Communication

Education

Bachelor's degree

Tools

Databricks
Azure
Spark
Datadog
Splunk
ServiceNow
Azure DevOps
Python
Application Insights

Job description

Vizient in Missouri seeks a production support professional to provide technical triage, diagnostics, and operational support across Vizient's core platforms and business domains. You will improve incident response, restoration speed, and stakeholder communication while partnering with Engineering, Product, and Operations teams to diagnose issues and strengthen platform reliability.

The role requires 2+ years in application/production support, strong SQL/log analysis, cloud-hosted service

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Support Engineer: Platform Reliability & Incidents
Production Support Engineer: Platform Reliability & Incidents

Vizient • Chicago (IL)

On-site
USD 77,000 - 135,000
Development Support Engineer
Development Support Engineer

Vizient • Chicago (IL)

On-site
USD 77,000 - 135,000
Development Support Engineer
Development Support Engineer

Vizient • Missouri

On-site
USD 77,000 - 135,000
Incentive eligible
Infrastructure Support Engineer I
Infrastructure Support Engineer I

Vizient, Inc • Chicago (IL)

On-site
USD 60,000 - 101,000
Comprehensive benefits
Operations Engineer: Production & Platform Reliability
Operations Engineer: Production & Platform Reliability

Veriipro • New York (NY)

On-site
USD 85,000 - 115,000
Senior Platform Engineer – Production Support & SQL
Senior Platform Engineer – Production Support & SQL

Beacon Hill • Orlando (FL)

On-site
USD 110,000 - 160,000
Production Support and QA Specialist
Production Support and QA Specialist

Agile Resources, Inc. • United States

Remote
USD 80,000 - 100,000
Onsite SRE II: Production Reliability & Incident Leader
Onsite SRE II: Production Reliability & Incident Leader

hudsonmanpower • Cincinnati (OH)

On-site
USD 110,000 - 140,000
Incident Command Lead – Production Reliability
Incident Command Lead – Production Reliability

State Street • Boston (MA)

On-site
USD 70,000 - 119,000
IT Production Support & Incident Management Expert
IT Production Support & Incident Management Expert

Jobtailor • Franklin (TN)

On-site
USD 90,000 - 120,000