Описание:
Xsolla is a global commerce company that provides tools and services to help video game developers fund, distribute, market, and monetize their games. It operates as a merchant of record and supports game developers around the world.
Задачи:
- Serve as Incident Commander for major incidents, coordinate cross-functional response teams, drive investigations, make escalation decisions, and ensure incidents are resolved within SLA targets
- Own incident communications by preparing timely updates for senior leadership, Customer Success, and partner and customer contacts, and managing customer-facing status page updates
- Facilitate blameless Post-Incident Reviews (PIRs), lead root cause identification, assign corrective actions with owners and deadlines, and track them to closure
- Analyze incident trends, recurring issues, and production bugs; identify patterns, create Problem tickets, and regularly report findings and recommendations to product and engineering teams
- Enforce the incident management framework, including the severity model, priority matrix, SLA targets, escalation procedures, and deployment readiness gates
- Oversee and mentor the Operations Engineer on the shift, coach on triage, investigation, runbook execution, and documentation quality, and conduct knowledge-transfer sessions
- Produce shift handoff reports and operational reporting on incident trends, KPI performance (MTTD, MTTA, MTTR), SLA adherence, proactive detection rates, and repeat incident analysis
- Regularly audit service catalogue completeness and govern JIRA Service Management workflows for incident, PIR, and problem management
- Cover the Operations Engineer role during absences, breaks, or surge incidents
- Participate in a weekend on-call rotation for major incidents
Требования:
- Previous experience working at a gaming company
- 6+ Years of experience in incident management, SRE, NOC leadership, or technical operations in a production environment supporting high-availability, high-transaction systems
- Proven experience coordinating multi-team incident responses, making real-time escalation decisions, and communicating with executive stakeholders under pressure
- Excellent written and verbal English communication skills, including drafting executive updates under pressure, facilitating blameless PIRs, presenting operational metrics to senior leadership, and communicating incident status to customers and partners
- Strong ITIL foundation and practical experience with incident, problem, and change management lifecycles and ITIL-aligned workflows
- Technical knowledge of observability tools; ability to interpret logs, traces, and metrics in Datadog or equivalent tools such as Grafana, Splunk, or New Relic
- Understanding of APM, SLOs, error budgets, burn-rate alerting, and synthetic monitoring
- Hands-on experience with Datadog, PagerDuty or OpsGenie, JIRA or JIRA Service Management, Slack, and Confluence
- Ability to identify trends and recurring issues in incident data and turn them into recommendations for product and engineering teams
- Experience with SLA/SLO-driven operations measuring, reporting, and improving MTTD, MTTA, and MTTR
- Comfortable with 24x7 shift-based operations in a follow-the-sun model with handoff overlaps
- Будет плюсом: customer/partner-facing incident communications and status page management, AI/ML-assisted operations, JIRA Service Management administration, Datadog Service Catalog, scorecards and SLOs, building an operations function from scratch, Kubernetes, cloud infrastructure (GCP preferred), microservices architecture, distributed systems, ITIL certification
Условия:
- Unlimited Flexible Time Off
- Gym membership and monthly train ticket
- Personalized career roadmap, training, and educational opportunities
- Weekend on-call rotation for critical severities is required
- Background checks may be conducted after the final interview stage, where permitted by law and in compliance with local regulations