Get more replies from employers
Send a job-specific resume in minutes.
Nium is seeking a Senior Manager of Site Reliability Engineering to lead global SRE teams ensuring high availability and performance of our payments platform. You will own incident management, observability, and capacity planning, partnering with product, security, and infrastructure to meet uptime and regulatory requirements across 100+ markets.
You will build a resilient, automated environment with blameless postmortems and disaster recovery practices to support a growing fintech platform.
Nium is looking for a Senior Manager, Site Reliability Engineering to lead the teams responsible for the availability, performance, scalability, and operational excellence of our global payments platform. This is a hands-on leadership role: you will build and grow a team of SREs, define the reliability roadmap, and partner closely with product engineering, security, and infrastructure teams to ensure Nium's systems meet the always-on expectations of a regulated financial platform operating across 100+ markets.
You will own incident management, observability, capacity planning, and production readiness practices, while championing a culture of blameless postmortems, proactive risk reduction, and engineering-driven automation. This role sits at the intersection of engineering leadership and operational rigor, and reports into Nium's engineering leadership.
Lead, mentor, and grow a team of SREs and reliability engineers across multiple time zones, setting clear goals, career paths, and performance expectations.
Own Nium's reliability strategy — defining and driving SLIs/SLOs/error budgets across critical payment, card issuance, and compliance services.
Drive incident management end-to-end: on-call structure, escalation paths, major incident response, and blameless postmortems that produce durable fixes, not just tickets.
Partner with product engineering leaders to embed reliability, scalability, and operational readiness into the software development lifecycle from design through launch.
Build and scale observability (metrics, logging, tracing, alerting) so that issues are detected and diagnosed before they impact customers or partner banks.
Lead capacity planning and performance engineering for systems processing high-volume, real-time financial transactions across a global, multi-region infrastructure.
Champion automation and self-healing systems to reduce toil, eliminate manual runbooks, and improve mean-time-to-detect and mean-time-to-resolve.
Own disaster recovery, business continuity, and chaos engineering practices, running regular game days to validate resilience assumptions.
Collaborate with Security and Compliance teams to ensure infrastructure practices meet regulatory and audit requirements (PCI-DSS, SOC 2, ISO 27001, and regional financial regulations).
Manage the reliability budget: tooling investments, cloud cost/performance trade-offs, and staffing plans, in partnership with finance and engineering leadership.
Represent SRE in executive reviews, translating technical risk and system health into business-relevant reporting for leadership and the board.
Establish and continuously refine production readiness reviews, runbooks, and operational standards across all engineering teams.
10+ years of experience in software engineering, infrastructure, or site reliability engineering, with 4+ years in a people-management or technical leadership role leading SRE/DevOps/Infrastructure teams.
Proven track record operating and scaling production systems for a high-availability, transaction-heavy platform — fintech, payments, banking, or e-commerce experience strongly preferred.
Deep hands-on expertise with cloud infrastructure (AWS) Kubernetes, container orchestration, and infrastructure-as-code (Terraform, CloudFormation, or similar).
Strong background in observability stacks (Prometheus, Grafana, Datadog, ELK/OpenSearch, or equivalent) and building alerting that reduces noise while catching real issues.
Demonstrated experience defining and operationalizing SLOs/error budgets, and using them to drive engineering prioritization.
Solid understanding of distributed systems, databases, caching, messaging queues, and API-driven microservice architectures at scale.
Experience leading major incident response for critical, customer-facing systems, including postmortem processes that drive real change.
Familiarity with security and compliance frameworks relevant to financial services (PCI-DSS, SOC 2, ISO 27001) and how they shape infrastructure and access practices.
Excellent communication skills — able to translate technical reliability concepts for engineering peers, product leaders, and executives alike.
A pragmatic, metrics-driven approach to engineering decisions: you measure before you optimize, and you test assumptions in production rather than in theory.
Experience with scripting/automation (Python, Go, or Bash) and CI/CD pipelines (Jenkins, ArgoCD, GitLab CI, or similar).
Experience operating systems subject to real-time payment rails, card networks, or banking core integrations.
Prior experience running SRE or infrastructure functions through hypergrowth or rapid international expansion.
Exposure to FinOps practices and cloud cost optimization at scale.
Experience building or scaling an SRE function from the ground up within an existing engineering organization.
We Value Performance:Through competitive salaries, performance bonuses, sales commissions, equity for specific roles and recognition programs, we ensure that all our employees are well rewarded and incentivized for their hard work.
We Care for Our Employees:The wellness of Nium’ers is our #1 priority. We offer medical coverage along with 24/7 employee assistance program, generous vacation programs including our year-end shut down. We also provide a flexible working hybrid working environment (3 days per week in the office).
We Upskill Ourselves:We are curious, and always want to learn more with a focus on upskilling ourselves. We provide role-specific training, internal workshops, and a learning stipend.
We Celebrate Together:We recognize that work is also about creating great relationships with each other. We celebrate together with company-wide social events, team bonding activities, happy hours, team offsites, and much more!
We Thrive with Diversity:Nium is truly a global company, with more than 33 nationalities, based in 18+ countries and more than 10 office locations. As an equal opportunity employer, we are committed to providing a safe and welcoming environment for everyone.