Senior Platform SRE

IG KnowHow

Kraków

On-site

PLN 240,000 - 420,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Growth opportunities
Mentoring programs
Networking clubs
Volunteer time off
Perks details

Job summary

IG Group is seeking a Senior Platform SRE to own and improve the reliability platform within a hybrid AWS/on-premises environment. The role focuses on observability, SLOs, and automated resilience across hundreds of services and engineers.

You will mentor others, drive incident prevention, and contribute to IG’s reliability standards. The position requires hands-on engineering, chaos testing, and strong collaboration with Platform Engineering and product teams in a fast-paced fintech setting.

Qualifications

  • Hands-on observability and instrumentation with OpenTelemetry and production use of monitoring tools.
  • Design and maintain SLOs, error budgets, and multi-window burn-rate alerts.
  • Build and operate CI/CD pipelines with blue/green/canary releases and automated rollback.
  • Experience with Kubernetes and Nomad in a hybrid cloud environment.
  • Production-quality coding in Java and/or Python.
  • Strong knowledge of distributed systems and failure modes.
  • Experience with incident management and blameless post-incident reviews.
  • Familiarity with chaos engineering practices and tooling.
  • Ability to write RFCs and contribute to engineering standards.

Responsibilities

  • Build and own the reliability platform and tooling.
  • Implement comprehensive monitoring and observability across services.
  • Maintain production readiness including automated deployments and zero-downtime patching.
  • Design self-healing capabilities and automated traffic rerouting.
  • Run chaos experiments to improve resilience and incident response.
  • Develop CI/CD pipelines embedding reliability practices.
  • Mentor engineers on reliability patterns and production engineering discipline.
  • Collaborate with platform teams to define SLOs on customer journeys.

Skills

Observability engineering
SLOs and error budgets
CI/CD and release engineering
Container orchestration
Software engineering (Java/Python)
Distributed systems
Incident management
Chaos engineering
Standards and governance

Tools

OpenTelemetry
Datadog
Dynatrace
Grafana
Kubernetes (EKS/AKS/GKE)
HashiCorp Nomad
Terraform
PagerDuty/ServiceNow

Job description

Job Title Senior Platform SRE Job Description So, who are we? IG is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions – built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated. What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses. The bar is high – bring a curious and forward-thinking mindset and we'll give you the platform to define what comes next. Join us at IG – the future gets built here.

The Platform SRE team is the engine of IG’s reliability programme. We sit within Infrastructure & Operations, working across IG’s hybrid estate of on-premises HashiCorp Nomad and AWS. We are not a reactive ops team. We build the platform, standards, and tooling that make reliability the default for every engineering team at IG. Through the SRE Guild, we connect with Domain SREs and Reliability Champions across the organisation, setting the bar and lifting it together.

Your role in the Team's Success You will be a hands-on technical contributor at the heart of the Platform SRE team, owning pieces of the reliability platform that hundreds of engineers depend on. You will work at the intersection of software engineering, observability, and systems reliability, turning reliability from a reactive concern into a proactive engineering discipline. You will partner with Platform Engineering, product teams, and Reliability Champions to define what good looks like in production and then make it the default. You will contribute to the SRE Guild, mentor engineers across the organisation, and when things go wrong, you will be on the call helping to mitigate, understand, and prevent a repeat.

What you'll do
  • Build and own the reliability platform
  • Implement comprehensive monitoring and observability using OpenTelemetry and distributed tracing.
  • Maintain SLO, error budgets and burn-rate tracking
  • Establish and maintain 24/7 operational readiness including automated deployments, blue/green releases, and zero-downtime patching strategies
  • Engineer self-healing capabilities: auto-remediation, error-budget-gated rollback, and automated traffic rerouting
  • Design and run chaos experiments across the AWS estate, turning severe-but-plausible failure scenarios into engineering improvements
  • Build automation tools and CI/CD pipelines that embed reliability practices, while applying software engineering discipline including version control, code reviews, and testing.
  • Contribute to the SRE AI agent, IG’s agentic tooling for incident investigation and reliability review, built on AWS frontier models
  • Mentor junior SREs and Reliability Champions on reliability patterns and production engineering discipline
  • Set and uphold standards
  • Author and evolve the SRE standards that underpin the Guild: SLO methodology, error budget policy, observability instrumentation guide, and Production Readiness Review (PRR) checklist
  • Mentor developers on reliability patterns including circuit breakers, retry logic, and fault tolerance
  • Work with development teams and Reliability Champions to design SLOs on customer journeys rather than per-service.
  • Assist and guide teams in system design, capacity planning, architectural reviews and closing observability gaps.
  • Own incident response and learning
  • Facilitate blameless post-incident reviews (PIRs) within five working days using contributing-factor methodology
  • Maintain the Lessons Register, track remediation actions to closure, and surface patterns across incidents quarterly
What you'll need for this role
  • Observability and instrumentation: hands-on OpenTelemetry experience (spans, metrics, traces, context propagation) and production use of Honeycomb, Datadog, Dynatrace, or Grafana; able to instrument Java or Python services directly.
  • SLOs and error budgets: proven track record designing customer-meaningful SLIs, setting error budgets, configuring multi-window burn-rate alerts, and working with development teams on reliability measurement
  • CI/CD and release engineering: experience building pipelines with safety mechanisms: blue/green and canary releases, automated rollback, and DORA metrics integration
  • Container orchestration: Kubernetes (EKS, AKS, or GKE) required; HashiCorp Nomad is a strong advantage on IG’s hybrid estate; solid understanding of cloud networking and IaC (Terraform preferred)
  • Software engineering: production-quality coding in Java and/or Python; comfortable contributing to application codebases to implement reliability patterns, not just configuring infrastructure around them
  • Distributed systems: strong understanding of how large-scale systems fail and how to make them fail safely; circuit breakers, bulkheads, idempotency, graceful degradation, and load-shedding; high-throughput, low-latency environments preferred
  • Incident management: on-call experience on production systems, blameless PIR facilitation, contributing-factor analysis, and driving action items to closure; PagerDuty and ServiceNow familiarity helpful
  • Chaos engineering: experience designing and executing hypothesis-driven experiments with blast-radius controls and gap-to-impact-tolerance analysis; AWS FIS, Gremlin, or equivalent
  • Community and standards: at ease in a guild or community-of-practice model; comfortable writing RFCs, presenting at engineering forums, and building standards that others will adopt
Experience Requirements

Track record in high-throughput, production environments (financial services, trading platforms, or similar mission-critical systems preferred)

  • Demonstrated ability to improve system reliability and performance at scale
  • Experience working collaboratively with development teams to implement observability and reliability improvements
  • Strong troubleshooting skills in distributed systems environments
Core Competencies
  • Systems thinking approach to problem-solving
  • Excellent communication skills for cross-functional collaboration and technical enablement
  • Ability to balance hands-on development work with operational responsibilities
  • Strong bias toward automation and eliminating manual toil
How we work

We try to take a thoughtful approach to our ways of working as a company. We follow a hybrid working model with 3 days in the office -- which we think balances the need to collaborate effectively and connect with each other. When it comes to how we deliver, there are 5 things we want everyone to do to drive high performance, better learning and career satisfaction: Lead and Inspire: Drives trust, alignment, and enthusiasm Think Big: Focus on the problems that most impact commercial outcomes Champion the client: Understand and prioritise client's needs Deliver at pace: Push for fast, sustainable growth; Raise the bar: Take ownership, be accountable and share feedback We believe that diversity is vital to success, it fuels creativity, drives innovation and sets us up for global success. We're committed to building teams with a variety of perspectives and skills to help us realise our vision and strategy, that's why we encourage applications from people with diverse backgrounds and experiences to join us on this journey. Learn more about our D & I approach here.

The Perks
  • Your growth fuels our success!
  • Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression.
  • Expand your network through committees, sports and social clubs.
  • Enjoy extra time off for volunteering and community work.
  • Learn more about the Perks here!
  • Join us for this exciting journey.

Join us for this exciting journey.

Number of openings 1 Looking for a career at a company that will support you, challenge you and help you grow? IG Group can provide that. IG Group is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions – built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated. What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses. The bar is high – bring a curious and forward-thinking mindset and we'll give you the platform to define what comes next. Join us at IG – the future gets built here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform SRE
Senior Platform SRE

IG Group • Kraków

On-site
PLN 250,000 - 420,000
Career development
Mentoring programs
Networking clubs
+3
Data Engineer
Data Engineer

IG Group • Kraków

On-site
PLN 180,000 - 240,000
Home office equipment reimbursement
Performance related bonus
Private medical cover (Medicover)
+2
Data Engineer
Data Engineer

IG KnowHow • Kraków

Hybrid
PLN 180,000 - 260,000
Private medical cover
Home office equipment reimbursement
Performance related bonus
+2
Senior Machine Learning Engineer
Senior Machine Learning Engineer

IG Group • Kraków

Hybrid
PLN 240,000 - 320,000
Competitive salary & benefits plan
Birthday day off
Volunteer days off
+3
Customer Experience Associate (Spanish)
Customer Experience Associate (Spanish)

IG Europe GmbH • Kraków

Hybrid
PLN 60,000 - 85,000
Hybrid working
Annual financial bonus
Private medical cover
+4
Customer Experience Associate (Spanish)
Customer Experience Associate (Spanish)

IG Group • Kraków

Hybrid
PLN 67,000 - 100,000
Hybrid working
Home office equipment reimbursement
Annual financial bonus
+5
Customer Experience Associate (English)
Customer Experience Associate (English)

IG Group • Kraków

On-site
PLN 81,000 - 97,000
Hybrid working
Home office reimbursement
Annual bonus
+6
Customer Experience Associate (Swedish)
Customer Experience Associate (Swedish)

IG Europe GmbH • Kraków

Hybrid
PLN 81,000 - 97,000
Hybrid working
Private medical cover
Life insurance
+2
Customer Experience Associate (Swedish)
Customer Experience Associate (Swedish)

IG Group • Kraków

On-site
PLN 81,000 - 97,000
Hybrid working
Private medical cover for you and your
Life insurance
+2
Customer Experience Associate (French)
Customer Experience Associate (French)

IG Europe GmbH • Kraków

Hybrid
PLN 81,000 - 97,000
Hybrid working
Home office equipment reimbursement
Performance related annual bonus
+2