Production Support And Incident Response Lead

Nimbyx

Taguig

On-site

PHP 1,200,000 - 1,800,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

EviSmart is seeking a Global Continuity Lead to own incident response and platform continuity during night operations. You will be the lead incident response person on shift, investigate issues, determine safe restoration paths, and coordinate with Engineering and DevOps to minimize disruption for 2,000+ dental labs on our platform.

The role emphasizes business continuity, proactive monitoring, and clear communication with stakeholders during active incidents.

Qualifications

  • Experience in production support and incident management.
  • Ability to investigate under pressure during night shifts.
  • Strong understanding of logs, dashboards, APIs, and observability tools.

Responsibilities

  • Lead the response to production incidents during the night shift.
  • Continuously monitor platform health and act on warning signs.
  • Perform first-line investigations using logs and monitoring tools.
  • Assess impact and source, escalating when needed.
  • Develop safe workarounds or restorations when fixes aren’t immediate.
  • Escalate intelligently and coordinate with Engineering/DevOps.
  • Communicate clearly with stakeholders during incidents.
  • Document incidents, root causes, and follow-up actions.
  • Ensure handoff to daytime team with no loss of information.
  • Develop a backup capable of handling routine night triage.

Skills

Incident response
Log analysis
Root cause analysis
Cross-team coordination
Communication during incidents

Tools

Monitoring tools

Job description

Job Description:

Company Description

Why EviSmart

  • 300 people. Two hubs: Vancouver HQ and Manila operations.
  • 145% year-over-year SaaS growth - the market is responding.
  • 28 countries. One platform. The dental industry's Autopilot.
  • An in-house AI model research and development team building proprietary intelligence.

How We Work We ship before we're 100% certain. We write things down because we have two offices and memory is lossy. We debate loudly and move without resentment. We treat the customer's real problem as more important than an elegant internal process. If you've spent time waiting for permission to try something obvious — you'll notice the difference here immediately.

Job Description

Production Support & Incident Response Lead

When something goes wrong in production overnight, you are the person leading the response.

EviSmart is looking for a Global Continuity Lead to own incident response and platform continuity during our night operations.

This is not a role where you simply monitor dashboards, create a ticket, and wait for Engineering.

You will be the lead incident response person on shift. You are expected to investigate first, understand what is happening, determine the safest way to restore operations, bring in the right technical people when necessary, and remain accountable for the incident until the platform is stable.

Our platform supports 2,000+ dental labs, so an issue in production can quickly become a real business problem for our customers. The goal is simple: keep cases moving and minimize disruption.

What you'll own:

You will be responsible for the health and continuity of the platform during your coverage.

That means:

  • Lead the response to production incidents during the night shift
  • Continuously monitor platform health and act on warning signs before they become customer-impacting problems
  • Personally perform first-line investigation using logs, dashboards, monitoring tools, diagnostic commands and available system access
  • Determine the impact and likely source of an issue before escalating
  • Look for safe workarounds or restoration options when a permanent fix is not immediately available
  • Decide when an issue can be handled at your level and when Engineering, DevOps or another specialist needs to be brought in
  • Command the incident even after technical teams become involved: keep people aligned, decisions moving and communication clear
  • Keep Application Support and other stakeholders informed during active incidents
  • Document incidents, root causes, workarounds and follow-up actions
  • Make sure recurring issues don't simply become accepted problems
  • Provide a complete handoff to the daytime team with nothing dropped overnight
  • Develop one Application Support teammate into a reliable backup who can eventually handle routine night triage independently

The role's ownership of monitoring, incident command, proactive customer communication, post-mortems and backup development is explicit in the operating playbook.

What this role is NOT

This is not a traditional Service Delivery Manager or ITIL governance position.

It is also not a pure DevOps or Software Engineering role.

You don't need to be the person who writes the permanent code fix for every problem. But you do need enough technical depth to investigate intelligently before asking someone else to solve it.

If your normal incident process is:Alert → Create ticket → Escalate → Wait

this probably isn't the right role.

We're looking for someone whose instinct is closer to:

Detect → Investigate → Isolate → Restore or Work Around → Escalate Intelligently → Command Through Resolution → Prevent Recurrence

The kind of person we're looking for

You may currently be an:

Senior Application Support Engineer, L2/L3 Application Support Engineer, Production Support Engineer, Application Operations Engineer, Technical Operations Engineer, Platform Support Engineer, or similar.

More important than your current title is how you operate.

You are someone who:

  • Has personally supported live production applications
  • Can investigate an unfamiliar production problem without immediately needing someone to tell you what to check
  • Is comfortable working with logs, dashboards, APIs, databases and monitoring/observability tools
  • Understands enough infrastructure and application behavior to distinguish between likely application, database, API/connectivity and infrastructure problems
  • Thinks about business continuity, not only technical resolution
  • Can make sensible decisions with incomplete information
  • Knows when a workaround is safer and faster than waiting for the perfect fix
  • Stays calm when customers are affected and several teams are involved
  • Communicates clearly during incidents without creating noise
  • Can challenge or direct technical teams when an incident needs movement
  • Notices patterns and asks why the same problem keeps happening
  • Doesn't need constant hand-holding
  • Is comfortable being accountable when they are the most senior incident-response person available

Here's a good way to know whether you'll enjoy this role.

It's 2 AM.A production issue is preventing a customer from processing cases. The permanent fix requires an engineer who isn't immediately available.

What do you do?

We're looking for someone who doesn't stop at "I'll escape it."

We want someone who starts asking:

What's actually broken?
What's the business impact?
What changed?
What can I verify myself?
Can I safely restore the previous working state?
Is there another way to keep the customer's operation moving?
Who genuinely needs to be involved?
What can we do now instead of waiting until morning?

That's the mindset we're hiring for.

What success looks like

Your first 90 days are designed to progressively prove that we can trust you with the night.

First 30 days: Learn the platform and prove you can troubleshoot real issues and identify the correct workaround without being walked through every step.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Production Support & Incident Response Lead
Production Support & Incident Response Lead

EviSmart™ • Taguig

On-site
PHP 1,200,000 - 2,000,000
Production Support & Incident Response Lead
Production Support & Incident Response Lead

Nimbyx • Makati

On-site
PHP 1,339,000 - 2,009,000
Lead Application Support Engineer (SAAS)
Lead Application Support Engineer (SAAS)

EviSmart • Taguig

On-site
PHP 900,000 - 1,200,000
Lead Application Support Engineer SAAS
Lead Application Support Engineer SAAS

EviSmart™ • Taguig

On-site
PHP 600,000 - 1,000,000
Restaurant discounts (if any)
Site Reliability Engineer (SAAS)
Site Reliability Engineer (SAAS)

EviSmart • Taguig

On-site
PHP 800,000 - 1,100,000
Principal Enterprise Integration Technical Engineer
Principal Enterprise Integration Technical Engineer

Ingram Micro • Taguig

Hybrid
PHP 1,800,000 - 3,200,000
IT Support Specialist
IT Support Specialist

EviSmart™ • Taguig

On-site
PHP 335,000 - 603,000
Principal, Enterprise Integration (Technical Engineer)
Principal, Enterprise Integration (Technical Engineer)

Ingram Micro • Taguig

Hybrid
PHP 2,100,000 - 3,600,000
Principal, Enterprise Integration (Technical Engineer)
Principal, Enterprise Integration (Technical Engineer)

Ingram Micro Philippines BPO LLC • Pateros

Hybrid
PHP 1,800,000 - 2,600,000
Principal, Enterprise Integration (Technical Engineer)
Principal, Enterprise Integration (Technical Engineer)

Ingram Micro, Inc. • Philippines

Hybrid
PHP 2,000,000 - 4,000,000