technical service operations tso

HireHi

United States

Remote

USD 120,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Unlimited Flexible Time Off
Gym membership
Monthly train ticket

Job summary

Xsolla — глобальная платформа для поддержки разработки, распространения и монетизации игр. В должности Инженера по инцидент-менеджменту вы будете руководить ответами на крупные инциденты, обеспечивать своевременные обновления для руководства и клиентов, а также проводить PIR и корень причин с командами Product и Engineering.

Требуется опыт в индустрии игр, владение ITIL и умение работать в режиме 24x7. Мы предлагаем гибкий график, развитие навыков и возможность карьерного роста.

Qualifications

  • 6+ лет опыта в управлении инцидентами, SRE или технической эксплуатации в продуктивной среде.
  • Опыт координации межкомандных инцидент-ответов и коммуникаций с руководством.
  • Продвинутый уровень владения ITIL и жизненными циклами инцидентов, изменений и проблем.
  • Опыт работы с инструментами наблюдаемости и мониторинга.

Responsibilities

  • Выступать как Инцидент Командир для крупных инцидентов и обеспечение SLA.
  • Владение коммуникациями по инцидентам для руководства, Success и партнеров.
  • Ведение без blame-разборов (PIR) и назначение корректирующих действий.
  • Анализ тенденций инцидентов, проблем и выпусков, формирование рекомендаций.
  • Обеспечение рамок управления инцидентами, уровни важности и эскалации.
  • Ментор Operations Engineer, обучение триажу и работе с Runbook.
  • Создание отчетов по сменам и KPI по инцидентам и SLA.
  • Аудит каталога услуг и управление JIRA Service Management.
  • Замещение Operations Engineer при отсутствии, перерасходах.
  • Участие в очередной дежурности на выходных для критических ситуаций.

Skills

Incident management
SRE leadership
Executive communication
ITIL framework
Observability tools
JIRA Service Management
Datadog
PagerDuty
Grafana
Kubernetes

Tools

Datadog
Grafana
Splunk
New Relic
JIRA Service Management
Slack
Confluence
OpsGenie

Job description

Описание:

Xsolla is a global commerce company that provides tools and services to help video game developers fund, distribute, market, and monetize their games. It operates as a merchant of record and supports game developers around the world.

Задачи:
  • Serve as Incident Commander for major incidents, coordinate cross-functional response teams, drive investigations, make escalation decisions, and ensure incidents are resolved within SLA targets
  • Own incident communications by preparing timely updates for senior leadership, Customer Success, and partner and customer contacts, and managing customer-facing status page updates
  • Facilitate blameless Post-Incident Reviews (PIRs), lead root cause identification, assign corrective actions with owners and deadlines, and track them to closure
  • Analyze incident trends, recurring issues, and production bugs; identify patterns, create Problem tickets, and regularly report findings and recommendations to product and engineering teams
  • Enforce the incident management framework, including the severity model, priority matrix, SLA targets, escalation procedures, and deployment readiness gates
  • Oversee and mentor the Operations Engineer on the shift, coach on triage, investigation, runbook execution, and documentation quality, and conduct knowledge-transfer sessions
  • Produce shift handoff reports and operational reporting on incident trends, KPI performance (MTTD, MTTA, MTTR), SLA adherence, proactive detection rates, and repeat incident analysis
  • Regularly audit service catalogue completeness and govern JIRA Service Management workflows for incident, PIR, and problem management
  • Cover the Operations Engineer role during absences, breaks, or surge incidents
  • Participate in a weekend on-call rotation for major incidents
Требования:
  • Previous experience working at a gaming company
  • 6+ Years of experience in incident management, SRE, NOC leadership, or technical operations in a production environment supporting high-availability, high-transaction systems
  • Proven experience coordinating multi-team incident responses, making real-time escalation decisions, and communicating with executive stakeholders under pressure
  • Excellent written and verbal English communication skills, including drafting executive updates under pressure, facilitating blameless PIRs, presenting operational metrics to senior leadership, and communicating incident status to customers and partners
  • Strong ITIL foundation and practical experience with incident, problem, and change management lifecycles and ITIL-aligned workflows
  • Technical knowledge of observability tools; ability to interpret logs, traces, and metrics in Datadog or equivalent tools such as Grafana, Splunk, or New Relic
  • Understanding of APM, SLOs, error budgets, burn-rate alerting, and synthetic monitoring
  • Hands-on experience with Datadog, PagerDuty or OpsGenie, JIRA or JIRA Service Management, Slack, and Confluence
  • Ability to identify trends and recurring issues in incident data and turn them into recommendations for product and engineering teams
  • Experience with SLA/SLO-driven operations measuring, reporting, and improving MTTD, MTTA, and MTTR
  • Comfortable with 24x7 shift-based operations in a follow-the-sun model with handoff overlaps
  • Будет плюсом: customer/partner-facing incident communications and status page management, AI/ML-assisted operations, JIRA Service Management administration, Datadog Service Catalog, scorecards and SLOs, building an operations function from scratch, Kubernetes, cloud infrastructure (GCP preferred), microservices architecture, distributed systems, ITIL certification
Условия:
  • Unlimited Flexible Time Off
  • Gym membership and monthly train ticket
  • Personalized career roadmap, training, and educational opportunities
  • Weekend on-call rotation for critical severities is required
  • Background checks may be conducted after the final interview stage, where permitted by law and in compliance with local regulations
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

IT Operations Manager
IT Operations Manager

HireHi • United States

Remote
USD 70,000 - 110,000
Remote work
Internal R&D opportunities
Paid time off
+1
Team Lead Backend Engineer
Team Lead Backend Engineer

HireHi • United States

Remote
USD 140,000 - 210,000
Unlimited time off
Private health insurance
Career roadmap
+5
support engineer for B2B iGaming platforms
support engineer for B2B iGaming platforms

HireHi • United States

Remote
USD 55,000 - 65,000
Annual Benefits Budget
Maternity / Paternity leave
20+ Vacation days
+1
Backend Engineer Xsolla ID
Backend Engineer Xsolla ID

HireHi • United States

Remote
USD 120,000 - 180,000
Unlimited flexible time off
Private health insurance with dental —
Personalized career roadmap
+3
platform operations engineer
platform operations engineer

HireHi • United States

Remote
USD 120,000 - 170,000
Vacation days
English lessons
Team Lead Web3
Team Lead Web3

HireHi • United States

Remote
USD 180,000 - 260,000
Unlimited Flexible Time Off
Private health insurance with dental
Career roadmap and growth
+5
devops engineer for hybrid infrastructure
devops engineer for hybrid infrastructure

HireHi • United States

Remote
USD 120,000 - 160,000
AWS/GCP stack
Microsoft 365 stack
support engineer in infrastructure operations
support engineer in infrastructure operations

HireHi • United States

Remote
USD 75,000 - 95,000
devops engineer in Linux infrastructure
devops engineer in Linux infrastructure

HireHi • United States

Remote
USD 140,000 - 210,000
Гибкий график
Оплачиваемый отпуск 24 дня в год
Медицинское страхование
support engineer
support engineer

HireHi • United States

On-site
USD 90,000 - 120,000