Site Reliability Engineer

Cboe

Greater London

On-site

GBP 90,000 - 130,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cboe Global Markets in London seeks a Site Reliability Engineer to keep our low-latency trading platforms highly available and resilient. The role combines software development, systems, network, and database management to support global markets from London.

You will work with cross-functional teams to implement automation, monitor performance, and ensure robust incident response across time zones. This position requires independent work and a focus on continuous improvement of platform

Qualifications

  • Extensive experience in site reliability engineering for financial markets.
  • Strong scripting in Python and automation in cloud and on-prem environments.
  • Experience with low-latency trading platforms and incident management.

Responsibilities

  • Provide production support across time zones and manage incidents.
  • Develop and maintain automation tools to improve platform reliability.
  • Handle capacity planning and performance tuning for trading platforms.
  • Collaborate with software, systems, and network teams.

Skills

Python
Linux
Docker
Kubernetes
Cloud platforms
SQL
Incident response
On-call readiness

Tools

Terraform
Solarflare/Onload
Git
Prometheus
Grafana

Job description

  • At Cboe Global Markets, we inspire our people to solve complex challenges together because what we do matters. We provide the financial infrastructure that powers the global economy. As a leading provider of market infrastructure and tradable products, Cboe delivers cutting-edge trading, clearing and investment solutions to market participants around the world
  • We’re building meaningful ways to support professional and personal development while strengthening the trust we’ve earned as a global market leader. Our teams are empowered to share ideas, actively pursue them and bring on a challenge. As champions of internal mobility and access to opportunity, we encourage our people to “go for it” and equip our managers with the training to coach their teams to the next level. We strive to provide employees a safe space to network, share ideas and create opportunities
  • The Site Reliability Engineer (London) is a role served by experienced technologists with a diverse set of skills ranging from software development to systems, network, application, and/or database management
  • The Cboe Site Reliability Engineering team is a highly skilled unit responsible for platform engineering, configuration management, implementation, capacity planning, performance tuning, analysis, troubleshooting, reporting, and process automation
  • This position is instrumental in support of both Cboe’s European markets and Cboe’sfollow-the-sun support modelfor its US Global Trading Hours (GTH) markets, providing critical overnight and early-session coverage from London that ensures continuous, high-availability operations across Cboe’s real-time low-latency trading platforms
  • The London-based SRE provides technical support to Cboe Trade Desk and Operations Support Center staff across time zones, and works closely with Software Engineering, Systems Engineering, and Network Engineering teams to troubleshoot complex issues and coordinate platform configuration updates. A Site Reliability Engineer must be able to work independently with little to no direct supervision in performing their duties
  • Platform Configuration Management:Provide configuration management of new and existing trading platforms and support implementation of new features and functionality based on new business requirements
  • Monitor development activities, change management tickets, and evaluate their impact on Cboe Operations. Execute daily change tickets assigned to Site Reliability Engineering in support of updates to production, disaster recovery, and certification systems. While the primary focus of this role involves support of bare-metalon-premises infrastructure, experience withcloud platforms(e.g., AWS, Azure, GCP) andcontainerization technologies(e.g., Docker, Kubernetes) is desirable
  • Incident Response & Technical Troubleshooting:Serve as a technical responder for production incidents occurring during US GTH market hours covered from the London time zone. Participate in incident triage, root cause analysis, and resolution in coordination with globally distributed engineering and operations teams. Provide timely, precise communication to stakeholders during active incidents and contribute to post-incident reviews and remediation tracking to drive long-term platform stability
  • System Availability & Technical Support:Provide technical support and operational oversight to sustain resiliency and high availability of critical business operations. Monitor production, disaster recovery, and certification systems for issues. Analyze and optimize performance of real-time trading platforms. Operate and maintain low-latency bare-metal infrastructure, including hardware health, Linux OS tuning, and kernel-bypass networking stacks such as Solarflare/Onload. Investigate software defects. Assist the build team to resolve build/deployment issues
  • Reporting & Data Analysis:Create and improve upon existing reports related to Operations management. Analyze technical data sets (e.g., order entry, market data, matching engine logs) to troubleshoot or explain perceived issues. Execute SQL queries against a database to perform data analysis for customers and associates. Service and support historical data product requests
  • Capacity Planning:Drive capacity planning decisions for Cboe Exchanges and support capacity planning needs of various Cboe business units. Participate in Capacity Planning meetings with engineering and technical operations management staff
  • Automation & Process Improvement:Support task and system health automation efforts through development, testing, and maintenance of Python tools. Leverage AI to maximize efficiency
  • On-Call & Weekend Testing:Participate in weekend testing (e.g., capacity testing, fail-over, etc.) and provide follow-the-sun on-call technical support as part of Cboe’s global Operations team
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineer - Site Reliability - cboe
Engineer - Site Reliability - cboe

OpenTalent • Greater London

On-site
GBP 70,000 - 110,000
Engineer - Site Reliability
Engineer - Site Reliability

United States Digital Space LLC • Greater London

On-site
GBP 60,000 - 85,000
Site Reliability Engineer — Trading Platform (London)
Site Reliability Engineer — Trading Platform (London)

OpenTalent • Greater London

On-site
GBP 70,000 - 110,000
Site Reliability Engineer – London: Low-Latency Trading
Site Reliability Engineer – London: Low-Latency Trading

Cboe • Greater London

On-site
GBP 90,000 - 130,000
Site Reliability Engineer (London) - Banking & Finance
Site Reliability Engineer (London) - Banking & Finance

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 81,000 - 99,000
Lucrative bonus scheme
SRE Manager (Low Latency Linux)
SRE Manager (Low Latency Linux)

Signify Technology • Greater London

Hybrid
GBP 120,000 - 180,000
Senior QA Manager
Senior QA Manager

Cboe Global Markets • Greater London

Hybrid
GBP 90,000 - 130,000
Medical insurance
Pension plan
Annual leave 25 days
+4
Site Reliability Engineer
Site Reliability Engineer

Selby Jennings • City Of London

On-site
GBP 90,000 - 140,000
Senior QA Manager
Senior QA Manager

Cboe • Greater London

Hybrid
GBP 90,000 - 125,000
Private medical insurance
25 days annual leave
Pension plan
+3
London SRE: Global Trading Platform
London SRE: Global Trading Platform

United States Digital Space LLC • Greater London

On-site
GBP 60,000 - 85,000