Production Support Engineer

Compunnel, Inc.

Aliso Viejo, Northern (CA, KY)

Hybrid

USD 120,000 - 160,000

Full time

9 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Compunnel, Inc. is seeking a hands-on Lead - Production Support Engineer to drive reliability and observability across Datadog, SQL Server, Oracle, and related tools.

You will lead incident response, RCA efforts, and escalation governance while partnering with engineering, database, and vendor teams to accelerate operational improvements. The role emphasizes automation, runbooks, and scalable processes, with a remote-friendly setup within the United States.

Qualifications

  • 4+ years in Production Support, IT Operations, or related fields.
  • Experience leading technical teams and incident response.
  • Strong knowledge of monitoring and observability practices.

Responsibilities

  • Participate in daily production and operational support activities.
  • Assist engineers with troubleshooting and technical escalations.
  • Support production incidents (P0/P1/P2) and SWAT bridges.
  • Lead RCA activities and post-incident reviews.
  • Support deployments, migrations, and change implementations.
  • Define monitoring standards and governance for Datadog.
  • Create and maintain Datadog monitors and dashboards.
  • Drive alert optimization and observability maturity.
  • Partner with cross-functional teams for reliability improvements.
  • Own KPI/SLA reviews and operational reporting.

Skills

Datadog
SQL Server
Oracle
SSIS
SSRS
Informatica
Observability
Incident management
Runbooks
Leadership
Stakeholder management

Tools

Datadog

Job description

The Lead - Production Support Engineer is a hands-on technical leadership role responsible for the operational health, reliability, observability, automation, and continuous improvement of APM and EnvOps services. The role requires active participation in production incidents, SWAT engagements, deployments, troubleshooting, migrations, monitoring onboarding, automation initiatives, and operational escalations while also providing technical leadership, coaching, governance, and stakeholder management. The engineer will support services including Datadog, SQL Server, Oracle, SSIS, SSRS, Informatica, CTU, APM services, and Precise and TeamQuest tools.

Key Responsibilities
  • Participate in daily production and operational support activities.
  • Assist engineers with troubleshooting, issue resolution, and technical escalations.
  • Support P0, P1, and P2 production incidents and SWAT bridges.
  • Lead technical investigations and Root Cause Analysis (RCA) activities.
  • Support deployments, change implementations, migrations, and operational escalations.
  • Perform monitoring onboarding, configuration, and ongoing optimization.
  • Define monitoring standards and governance for Datadog and observability services.
  • Create and maintain Datadog monitors and dashboards.
  • Drive alert optimization, noise reduction, monitoring coverage, and observability maturity.
  • Partner with application, database, engineering, and vendor teams on observability and platform improvements.
  • Assess business and operational impact during incidents and coordinate technical response activities.
  • Provide leadership and stakeholder updates during critical incidents and ensure appropriate follow-up actions.
  • Lead RCA efforts, reduce recurring incidents and alert fatigue, and improve service reliability.
  • Identify technical debt and reliability improvement opportunities and drive corrective action plans through closure.
  • Identify and prioritize operational automation opportunities.
  • Partner with automation engineers on operational workflows and support AI-enabled operational solutions.
  • Promote automation-first thinking to reduce manual effort and improve scalability.
  • Coach and mentor engineers and facilitate cross-training and knowledge sharing.
  • Reduce single points of failure and develop backup coverage models across the team.
  • Drive technical capability development across the engineering team.
  • Own KPI/SLA reviews, operational reporting, governance dashboards, and service metrics.
  • Conduct documentation and runbook reviews and ensure operational knowledge is maintained.
  • Track risks, issues, action items, and service improvements.
  • Support executive and stakeholder reporting.
  • Support strategic migration activities and observability improvement initiatives.
  • Contribute to automation transformation roadmaps and initiatives focused on operational efficiency and scalability.
Required Qualifications
  • 4+ years of experience in Production Support, Application Support, IT Operations, Database Operations, Observability, or Infrastructure Support.
  • 1+ years of experience leading technical teams, incident response, operational governance, stakeholder management, and service ownership.
  • Strong experience with SQL Server administration and production support.
  • Experience with Oracle support and operations.
  • Experience with Datadog administration and observability.
  • Experience with SSIS, SSRS, and Informatica.
  • Strong knowledge of incident and problem management processes.
  • Experience with performance monitoring and tuning.
  • Experience supporting production environments and participating in SWAT engagements.
  • Strong Root Cause Analysis and troubleshooting skills.
  • Experience with monitoring and alert management.
  • Strong documentation and runbook management skills.
  • Strong leadership, coaching, stakeholder management, and communication skills.
Preferred Qualifications
  • Experience with AI and automation platforms.
  • Experience with Power Automate.
  • Experience with Playwright.
  • Experience developing or supporting workflow automation.
  • Experience with Monitoring as Code.
  • Knowledge of reliability engineering practices.
  • Experience with cloud monitoring and observability.

Notes:

Production Support Engg

Experience: 4+ years in Production Support, Application Support, IT Operations, Database Operations, Observability, or Infrastructure Support.

Leadership: 1+ years leading technical teams, incident response, operational governance, stakeholder management, and service ownership.

Need a person with good attitude/should be passionate and should have good communication skills.

Position is remote in US.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Support Engineer
Production Support Engineer

Value Innovation Labs • Alpharetta (GA)

On-site
USD 120,000 - 150,000
Production Support
Production Support

Tata Consultancy Services • Phoenix (AZ)

On-site
USD 90,000 - 100,000
Production Support Engineer
Production Support Engineer

ASM Tech Solutions • Town of Florida (NY)

Hybrid
USD 90,000 - 120,000
Production Support and QA Specialist
Production Support and QA Specialist

Agile Resources, Inc. • United States

Remote
USD 80,000 - 100,000
Production Support Engineer
Production Support Engineer

ASM Tech Solutions • United States

Hybrid
USD 90,000 - 150,000
Mid Level AWS Production Support Engineer
Mid Level AWS Production Support Engineer

System One • Birmingham (AL)

On-site
USD 80,000 - 120,000
Vice President, Production Services Application Support
Vice President, Production Services Application Support

BNY Mellon • Pittsburgh

On-site
USD 90,000 - 150,000
Remote Production Support & Observability Lead
Remote Production Support & Observability Lead

Compunnel, Inc. • Aliso Viejo (CA), Northern (KY)

Hybrid
USD 120,000 - 160,000
Application Engineer View role →
Application Engineer View role →

NRnP Technology • Northern (KY)

Hybrid
USD 90,000 - 140,000
Technology Support Lead - AI and Production Management Tools
Technology Support Lead - AI and Production Management Tools

JPMorgan Chase & Co. • Columbus (OH)

On-site
USD 180,000 - 240,000