Stand out for this role — generate a tailored resume and cover letter in about a minute.
Evi Smart in Taguig, Philippines is seeking a Production Support & Incident Response Lead to own the night shift incident response and platform continuity. You will investigate using logs, dashboards, APIs, and databases, determine safe workarounds, and coordinate with Eng/DevOps as needed while keeping stakeholders informed to minimize disruption.
This role emphasizes business continuity, post-mortems, and proactive incident handling to prevent recurring issues.
BGC, Taguig, Philippines
Why Evi Smart
How We Work • We ship before we're 100% certain. We write things down because we have two offices and memory is lossy. We debate loudly and move without resentment. We treat the customer's real problem as more important than an elegant internal process. If you've spent time waiting for permission to try something obvious — you'll notice the difference here immediately.
EviSmart is looking for a Production Support & Incident Response Leadto own incident response and platform continuity during our night operations.
This is not a role where you simply monitor dashboards, create a ticket, and wait for Engineering.
You will be thelead incident response person on shift. You are expected to investigate first, understand what is happening, determine the safest way to restore operations, bring in the right technical people when necessary, and remain accountable for the incident until the platform is stable.
Our platform supports2,000+ dental labs, so an issue in production can quickly become a real business problem for our customers. The goal is simple:keep cases moving and minimize disruption.
What you'll own • You will be responsible for the health and continuity of the platform during your coverage.
That means:
Lead the response to production incidents during the night shift
Continuously monitor platform health and act on warning signs before they become customer-impacting problems
Personally perform first-line investigation using logs, dashboards, monitoring tools, diagnostic commands and available system access
Determine the impact and likely source of an issue before escalating
Look for safe workarounds or restoration options when a permanent fix is not immediately available
Decide when an issue can be handled at your level and when Engineering, DevOps or another specialist needs to be brought in
Command the incident even after technical teams become involved: keep people aligned, decisions moving and communication clear
Keep Application Support and other stakeholders informed during active incidents
Document incidents, root causes, workarounds and follow-up actions
Make sure recurring issues don't simply become accepted problems
Provide a complete handoff to the daytime team with nothing dropped overnight
Develop one Application Support teammate into a reliable backup who can eventually handle routine night triage independently
The role's ownership of monitoring, incident command, proactive customer communication, post-mortems and backup development is explicit in the operating playbook.
It is also not a pure DevOps or Software Engineering role. You don't need to be the person who writes the permanent code fix for every problem. But youdo need enough technical depth to investigate intelligently before asking someone else to solve it.
If your normal incident process is:Alert → Create ticket → Escalate → Wait ...then this probably isn't the right role.
We're looking for someone whose instinct is closer to:
The kind of person we're looking for
You may currently be a: Senior Application Support Engineer, L2/L3 Application Support Engineer, Production Support Engineer, Application Operations Engineer, Technical Operations Engineer, Platform Support Engineer, or similar.
More important than your current title is how you operate.
You are someone who:
Has personally supported live production applications
Can investigate an unfamiliar production problem without immediately needing someone to tell you what to check
Is comfortable working with logs, dashboards, APIs, databases and monitoring/observability tools
Understands enough infrastructure and application behavior to distinguish between likely application, database, API/connectivity and infrastructure problems
Thinks aboutbusiness continuity, not only technical resolution
Can make sensible decisions with incomplete information
Knows when a workaround is safer and faster than waiting for the perfect fix
Stays calm when customers are affected and several teams are involved
Communicates clearly during incidents without creating noise
Can challenge or direct technical teams when an incident needs movement
Notices patterns and asks why the same problem keeps happening
Doesn't need constant hand-holding
Is comfortable being accountable when they are the most senior incident-response person available
Here's a good way to know whether you'll enjoy this role.
It's 2 AM.A production issue is preventing a customer from processing cases. The permanent fix requires an engineer who isn't immediately available.
What do you do?
We're looking for someone who doesn't stop at"I'll stop at"
We want someone who starts asking:
What's actually broken?
What's the business impact?
What changed?
What