Software Engineering Manager, Site Reliability Engineering, Traffic Steering

United States Digital Space LLC

Greater London

On-site

GBP 90,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is seeking a Site Reliability Engineer to lead a seasoned team ensuring the availability and performance of mission-critical services. You will automate responses, oversee incident management, and drive reliability across globally distributed systems.

Join a culture of curiosity and collaboration, mentoring engineers, balancing on-call rotations, and delivering scalable infrastructure with a strong focus on DNS and load balancing.

Qualifications

  • Bachelor's degree in CS or related field.
  • 8+ years of experience in software development.
  • 3+ years of people management.
  • 3+ years leading projects.
  • 3+ years designing, analyzing, and troubleshooting distributed systems.

Responsibilities

  • Manage a mature team of engineers, on various projects, contributing to career development.
  • Own end-to-end availability and performance of key services and automate responses to non-exceptional conditions.
  • Lead by example, mentor the team and establish credibility through technical execution.
  • Serve as escalation point for launch planning, feature deployment, and incident management.
  • Take part in on-call rotations across continents using a follow-the-sun model.

Skills

Software development experience
Project leadership
People management
Distributed systems design
DNS familiarity

Education

Bachelor's degree in Computer Science or related field

Job description

Minimum qualifications:
  • Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
  • 8 years of experience with software development in one or more programming languages.
  • 3 years of experience managing people or teams.
  • 3 years of experience leading projects.
  • 3 years of experience designing, analyzing, and troubleshooting distributed systems.
Preferred qualifications:
  • Master's degree in Computer Science or Engineering.
  • Direct technical experience in production infrastructure and high availability systems.
  • Familiarity with Networking and Domain Name System (DNS).
About the job

Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault‑tolerant systems. SRE ensures that the company's services—both our internally critical and our externally‑visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE’s will keep an ever‑watchful eye on our systems capacity and performance.

Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale which are unique to the company, while using your expertise in coding, algorithms, complexity analysis and large‑scale system design.

SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame‑free environment. We promote self‑direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.

To learn more: check out our books on Site Reliability Engineering or read a career profile about why a Software Engineer chose to join SRE.

The Traffic Steering SRE team is responsible for ensuring users around the globe connect to the company's services and cloud platform reliably and efficiently including managing the company's domain name system (DNS) infrastructure and the advanced telemetry pipelines that optimize network performance and user experience.

In this role, you will load balancing and DNS, putting you at the centre of the infrastructure relied upon by every service at the company. Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next generation of the company platforms, we make the company's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.

Responsibilities
  • Manage a mature team of engineers, on a variety of interesting and transformational projects, contributing to their career development and success.
  • Own end‑to‑end availability and performance of key services and build automation to prevent problem recurrence. Automate response to all non‑exceptional service conditions.
  • Lead by example, mentor the team and establish credibility through quality technical execution.
  • Serve as an escalation point for launch planning, feature deployment, and incident management.
  • Take part in and manage on‑call rotations across continents, using a follow‑the‑sun model.

the company is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also the company's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer III, Site Reliability Engineering, Traffic Network Load Balancing
Software Engineer III, Site Reliability Engineering, Traffic Network Load Balancing

United States Digital Space LLC • Greater London

On-site
GBP 70,000 - 110,000
Software Engineer II, Site Reliability Engineering, Labs SRE
Software Engineer II, Site Reliability Engineering, Labs SRE

United States Digital Space LLC • Greater London

On-site
GBP 60,000 - 90,000
Software Engineer/ SRE (Linux)
Software Engineer/ SRE (Linux)

United States Digital Space LLC • Basingstoke

Hybrid
GBP 60,000 - 90,000
SRE Architect (68019)
SRE Architect (68019)

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Reward Gateway • Greater London

Hybrid
GBP 60,000 - 65,000
Hybrid work option
SRE Engineering Manager: Traffic Steering & DNS
SRE Engineering Manager: Traffic Steering & DNS

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 140,000
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hitachids • Greater London

On-site
GBP 90,000 - 140,000
SRE, London, UK
SRE, London, UK

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LSEG • Nottingham

On-site
GBP 70,000 - 90,000
Healthcare
Retirement planning
Paid volunteering days
+1
Head of Site Reliability Engineering (SRE)
Head of Site Reliability Engineering (SRE)

United States Digital Space LLC • Greater London

Hybrid
GBP 150,000 - 210,000
ClassPass
Unlimited vacation
Apple equipment
+3