We’re working with a fast growing financial technology company developing an enterprise platform for banks and other regulated financial institutions. Its technology enables financial institutions to introduce and manage new digital financial products while retaining their existing core infrastructure. The platform operates across multiple international markets and supports both consumer-facing and enterprise environments. The company is now strengthening its production engineering capability as its platform, client base and transaction volumes continue to grow. The role
As a Senior Production Engineer, you’ll help maintain the reliability, resilience and performance of a business critical financial technology platform. This is a hands on engineering position combining production incident management, software development and continuous improvement. You’ll take technical ownership of complex production issues, investigate their underlying causes and implement permanent fixes rather than simply managing or escalating tickets. You’ll also deliver smaller customer facing enhancements, improve the maintainability of the codebase and help develop the team’s incident response practices. Working closely with software engineering, infrastructure and client-facing teams, you’ll play an important role in ensuring the platform continues to meet demanding availability and service commitments. The company operates a global support model, with teams collaborating across several time zones. Your normal working hours will remain aligned with the UK working week, alongside participation in a compensated on call rotation.
What you’ll be doing
- Leading the technical response to high priority production incidents, restoring services quickly and minimising customer impact.
- Conducting detailed root-cause analysis and working with engineering teams to implement lasting fixes.
- Developing and deploying smaller customer-facing features and platform improvements.
- Improving the performance, resilience and maintainability of production systems.
- Contributing to post-incident reviews and translating lessons learned into better tooling, runbooks and engineering practices.
- Working with infrastructure and platform teams to monitor system health and identify potential issues before they affect customers.
- Supporting and mentoring less experienced engineers in troubleshooting, incident response and coding standards.
- Improving handovers and knowledge sharing across geographically distributed teams.
- Helping mature incident, problem and change-management processes within a regulated environment.
- Providing technical input to client-facing and commercial teams when required.
What we’re looking for
- At least 5 years’ experience in software engineering, production engineering, sustaining engineering, site reliability engineering or a similar position.
- A strong software engineering background, with the ability to investigate, understand and modify production code.
- Commercial experience with Go, Java or another modern backend programming language.
- Advanced troubleshooting skills across complex, distributed production systems.
- Experience supporting event-driven or high-availability platforms.
- A track record of conducting root-cause analysis and implementing permanent corrective actions.
- Familiarity with observability, monitoring and incident-management tooling.
- Experience working to demanding service levels and improving measures such as availability and incident response times.
- Experience supporting systems operating within financial services or another regulated environment.
- The ability to communicate clearly during high-pressure incidents and collaborate with technical, commercial and client-facing stakeholders.
- Experience mentoring other engineers and improving team practices.
- Willingness to participate in a compensated out-of-hours on-call rota.
- A degree in computer science, engineering or a related discipline would be useful, although equivalent commercial experience will also be considered.
Helpful additional experience
- Experience within a fintech start up or scale up
- Knowledge of banking, payments, Open Banking, financial messaging or cardprocessing systems.
- Familiarity with event streaming, messaging and distributed data technologies.
- Experience with infrastructure as code, such as Terraform or CloudFormation.
- Familiarity with cloud platforms and modern monitoring or incident-response tools.
- Awareness of ITIL principles, particularly incident, problem and change management.
- Experience supporting sales, implementation or client-facing teams with technical input.