Operations Engineer, Germany

DUDE CHEM

Berlin

Hybrid

EUR 70.000 - 110.000

Vollzeit

Vor 2 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Unlimited Flexible Time Off
Gym membership
Monthly train ticket
Career roadmap

Zusammenfassung

Xsolla is seeking an Operations Engineer to join the Global Technical Operations (GTO) team. The role focuses on monitoring and investigating production issues across a global platform, improving incident detection and response, and communicating with partners and stakeholders during incidents.

Ideal candidates will have strong troubleshooting skills, observability platform experience (Datadog), and scripting ability, with experience in SRE/DevOps or production operations supporting

Qualifikationen

  • Experience in SRE, DevOps, production operations, or NOC environments.
  • Strong troubleshooting and incident management skills.
  • Excellent written and verbal English communication during incidents and handoffs.

Aufgaben

  • Monitor production issues across a global platform using observability tools.
  • Triage incidents and create incident tickets in JIRA Service Management.
  • Lead end-to-end resolution for lower-severity incidents and escalate when needed.

Kenntnisse

Observability
Troubleshooting
Communication
SRE
DevOps
Incident Management

Tools

Datadog
JIRA Service Management
Kubernetes
Slack
CI/CD tooling

Jobbeschreibung

We are looking for an Operations Engineer who is technically curious, detail-oriented, a strong communicator, and proactive to join our Global Technical Operations (GTO) team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to monitor and investigate production issues across a global platform, help improve how we detect and respond to incidents, analyze trends and patterns in production data, and contribute to better communication with partners and stakeholders during incidents.

Strong troubleshooting skills, observability platform experience, and scripting ability are essential, along with experience in SRE, DevOps, production operations, or NOC environments supporting high-availability platforms in the gaming industry. The ability to communicate clearly and effectively in English — both written and verbal — when writing incident updates, shift handoffs, and status page communications will be key to your success in this role.

If you’re passionate about keeping critical systems running and continuously improving operational processes and love being the first to spot issues and the one who drives them to resolution for game developers and players worldwide, we would love to hear from you!

ABOUT US

Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.

Serve as the primary dashboard monitor during your shift — continuously watch the GTO Operational Dashboard in Datadog, detect anomalies by correlating signals across APM, logs, metrics, synthetic tests, and Real User Monitoring, and determine whether alerts warrant an incident ticket or can be resolved through immediate investigation.

Triage and investigate production incidents — create incident tickets in JIRA Service Management, perform initial technical investigation using Datadog (traces, logs, infrastructure and application metrics), determine blast radius and likely root cause domain, and route to the correct team (Product SRE, Infrastructure SRE, or Engineering) using the smart routing model.

Own lower-severity incidents end-to-end from detection through resolution — diagnose, execute runbook procedures, and resolve without escalation where possible. Escalate promptly when an incident is unresolved within defined thresholds or requires a code-level fix.

Support the TSO Lead during major incidents as the technical right hand in the war room — surface real-time data (error rates, impact scope, deployment history, related alerts), maintain the incident ticket with live timeline entries and linked evidence, and execute mitigation actions as directed.

Draft incident communications under TSO Lead direction, including internal Slack updates, stakeholder notifications, and customer-facing status page updates (status.xsolla.com). Support clear, timely communication throughout the incident lifecycle.

During non-incident periods, analyze incident trends, recurring issues, and production bugs — compile data from Datadog, JIRA, and Slack, identify patterns, and contribute findings to regular reports for product and engineering teams.

Compile incident timelines and draft initial PIR documents for Post-Incident Review preparation. Track PIR action items post-session and flag overdue items to the TSO Lead.

Build and maintain operational automation (alert enrichment scripts, incident templates, Slack workflows, dashboard widgets) and contribute to runbook development — documenting new resolution procedures so they can be repeated by any Operations Engineer on any shift.

Conduct structured shift handoffs covering active incidents, at-risk services, upcoming deployments, and follow-up items. Participate in knowledge transfer sessions with SREs to continuously expand independent resolution capability.

Cover for the TSO Lead during vacations, absences, or emergencies — including severity classification, escalation decisions, stakeholder communications, and basic Incident Commander functions.

Publish health reports of critical apps periodically.

ABOUT YOU

We are looking for an Operations Engineer who is technically curious, detail-oriented, a strong communicator, and proactive to join our Global Technical Operations (GTO) team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to monitor and investigate production issues across a global platform, help improve how we detect and respond to incidents, analyze trends and patterns in production data, and contribute to better communication with partners and stakeholders during incidents.

Strong troubleshooting skills, observability platform experience, and scripting ability are essential, along with experience in SRE, DevOps, production operations, or NOC environments supporting high-availability platforms in the gaming industry. The ability to communicate clearly and effectively in English — both written and verbal — when writing incident updates, shift handoffs, and status page communications will be key to your success in this role.

If you’re passionate about keeping critical systems running and continuously improving operational processes and love being the first to spot issues and the one who drives them to resolution for game developers and players worldwide, we would love to hear from you!

ABOUT US

Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.

For more information, visit xsolla.com .

Responsibilities
  • Serve as the primary dashboard monitor during your shift — continuously watch the GTO Operational Dashboard in Datadog, detect anomalies by correlating signals across APM, logs, metrics, synthetic tests, and Real User Monitoring, and determine whether alerts warrant an incident ticket or can be resolved through immediate investigation.
  • Triage and investigate production incidents — create incident tickets in JIRA Service Management, perform initial technical investigation using Datadog (traces, logs, infrastructure and application metrics), determine blast radius and likely root cause domain, and route to the correct team (Product SRE, Infrastructure SRE, or Engineering) using the smart routing model.
  • Own lower-severity incidents end-to-end from detection through resolution — diagnose, execute runbook procedures, and resolve without escalation where possible. Escalate promptly when an incident is unresolved within defined thresholds or requires a code-level fix.
  • Support the TSO Lead during major incidents as the technical right hand in the war room — surface real-time data (error rates, impact scope, deployment history, related alerts), maintain the incident ticket with live timeline entries and linked evidence, and execute mitigation actions as directed.
  • Draft incident communications under TSO Lead direction, including internal Slack updates, stakeholder notifications, and customer-facing status page updates (status.xsolla.com). Support clear, timely communication throughout the incident lifecycle.
  • During non-incident periods, analyze incident trends, recurring issues, and production bugs — compile data from Datadog, JIRA, and Slack, identify patterns, and contribute findings to regular reports for product and engineering teams.
  • Compile incident timelines and draft initial PIR documents for Post-Incident Review preparation. Track PIR action items post-session and flag overdue items to the TSO Lead.
  • Build and maintain operational automation (alert enrichment scripts, incident templates, Slack workflows, dashboard widgets) and contribute to runbook development — documenting new resolution procedures so they can be repeated by any Operations Engineer on any shift.
  • Conduct structured shift handoffs covering active incidents, at-risk services, upcoming deployments, and follow-up items. Participate in knowledge transfer sessions with SREs to continuously expand independent resolution capability.
  • Cover for the TSO Lead during vacations, absences, or emergencies — including severity classification, escalation decisions, stakeholder communications, and basic Incident Commander functions.
  • Publish health reports of critical apps periodically.
Qualifications
  • Previous experience working at a gaming company is required — you understand the pace, player expectations, live operations dynamics, and the operational demands of the gaming industry.
  • 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment.
  • Strong troubleshooting and investigation skills — ability to take an alert or user-reported symptom and methodically trace it through the stack: application logs, APM traces, infrastructure metrics, database queries, and network paths.
  • Hands-on experience with Datadog (or equivalent observability platform: Grafana, Splunk, New Relic, Elastic) — navigating APM, building log queries, reading infrastructure dashboards, interpreting SLO burn rates, and configuring monitors and alerts.
  • Proficiency in at least one scripting language: Python, Go, or Bash. You will write automation scripts, build operational tooling, and work with APIs.
  • Clear written and verbal communication skills in English — ability to write incident tickets, investigation notes, Slack updates, shift handoff reports, status page communications, and PIR drafts that are clear, concise, and useful to both technical and non-technical audiences.
  • Working knowledge of Kubernetes and cloud infrastructure (GCP preferred, AWS/Azure acceptable) — understanding of pods, deployments, services, ingress, node health, and how to investigate Kubernetes-related production issues.
  • Understanding of SLOs, error budgets, and burn-rate alerting — knowing what a multi-window burn-rate alert means, how error budgets deplete, and how SLO breaches translate into incident severity.
  • Experience with incident management tooling: JIRA or JIRA Service Management, PagerDuty or OpsGenie, Slack, and Confluence.
  • Experience with or strong interest in AI/ML-assisted operations: anomaly detection, alert correlation, predictive monitoring, or automated remediation.
  • Comfort with 24x7 shift-based operations as part of a follow-the-sun model with handoff overlaps. Weekend on-call (rotating) is required.
Nice to Have
  • Familiarity with Datadog Service Catalog, synthetic monitoring, and RUM (Real User Monitoring).
  • Experience with distributed systems debugging: tracing failures across microservices, understanding cascading failures, and reading distributed traces end-to-end.
  • Exposure to database operations (MySQL, PostgreSQL, Redis, Kafka) at a level sufficient to investigate connection pool exhaustion, replication lag, slow queries, or queue backlogs during incidents.
  • Familiarity with CI/CD pipelines and deployment tooling (GitLab CI, ArgoCD, Helm) — enough to correlate recent deployments with production issues and identify rollback targets.
  • JIRA Service Management administration experience: workflows, automation rules, SLA timers, and queues.
  • ITIL Foundation certification is a plus but not required — practical experience matters more.
Benefits

We are passionate about fostering a supportive environment for our team, so we prioritize the physical, mental, and emotional well-being of our employees through a comprehensive Benefits Program. This includes unlimited Flexible Time Off, Gym membership, monthly train ticket and a personalized career roadmap for each employee. By investing in professional development through training and educational opportunities, we ensure that our team thrives both personally and professionally.

Together, we’re not just building a business; we’re cultivating a community that values creativity, collaboration, and the transformative power of play.

The duties and responsibilities of this position may evolve over time to support the organization’s goals and individual growth. This job description is intended to outline the general nature and level of work being performed and is not intended to be an exhaustive list of all duties, responsibilities, and qualifications required.

By submitting your application, you consent to Xsolla conducting background checks, where permitted by law, after the final interview stage. All checks will comply with local regulations, and your information will be handled confidentially.

Xsolla takes your privacy seriously and will not sell or externally distribute any personal data received during the hiring process. In accordance with applicable data protection laws, Xsolla is committed to protecting your personal information and respecting your privacy.

For any inquiries related to data privacy, please contact: careers@xsolla.com

Explore more opportunities at: https://xsolla.com/careers

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Operations Engineer, Germany
Operations Engineer, Germany

Xsolla • Berlin

Vor Ort
EUR 50.000 - 70.000
Unlimited Flexible Time Off
Gym membership
Monthly train ticket
+1
Technical Service Operations Lead (TSO Lead), Germany
Technical Service Operations Lead (TSO Lead), Germany

Xsolla • Berlin

Vor Ort
EUR 70.000 - 90.000
Unlimited Flexible Time Off
Gym membership
Monthly train ticket
+1
Tech Lead — LiveOps Team
Tech Lead — LiveOps Team

Xsolla • Deutschland

Remote
EUR 70.000 - 90.000
Medical, dental, and vision
PTO
Personalized career roadmap
DevOps Engineer
DevOps Engineer

Xsolla • Deutschland

Vor Ort
EUR 70.000 - 110.000
Medical benefits
Dental benefits
Vision benefits
+2
platform engineer
platform engineer

Enfint • Deutschland

Remote
EUR 65.000 - 95.000
Comprehensive benefits program
Medical, dental, and vision
PTO
+1
Senior Database Administrator
Senior Database Administrator

DUDE CHEM • Berlin

Hybrid
EUR 65.000 - 98.000
Unlimited Flexible Time Off
Gym membership
Monthly train ticket
+1
Junior UA/Growth Manager
Junior UA/Growth Manager

Xsolla • Deutschland

Remote
EUR 45.000 - 60.000
100% company-paid medical, dental, and vision plans
Unlimited Flexible Time Off
Personalized career roadmap
Senior Database Administrator
Senior Database Administrator

Xsolla • Berlin

Vor Ort
EUR 55.000 - 75.000
100% company-paid medical, dental, and vision plans
Unlimited Flexible Time Off
Personalized career roadmap
Account Executive — Gaming/Adtech Sales
Account Executive — Gaming/Adtech Sales

Xsolla • Berlin

Vor Ort
EUR 76.000 - 111.000
Payment Systems Delivery Manager
Payment Systems Delivery Manager

Xsolla • Deutschland

Hybrid
EUR 90.000 - 120.000
Medical benefits
Flexible time off
Career roadmap