Sr Network Engineer, Reliability and Observability

Blue Signal Search

Greater London

On-site

GBP 90,000 - 140,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity participation
Comprehensive benefits package

Job summary

Blue Signal Search is seeking a Sr Network Engineer focused on reliability and observability. The role blends production networking with software engineering to improve infrastructure resilience, telemetry, and automation across high-performance datacenter environments.

You will own incidents, design automated pipelines, and collaborate with hardware, software, and operations teams to elevate network health and performance.

Qualifications

  • 5+ years in network infrastructure with hands-on production ops.
  • Experience in data center networks, overlay tech, routing and fabric design.
  • Strong software engineering background with ITIL/Agile/xP and TDD practices.

Responsibilities

  • Drive network reliability programs using telemetry, incident history and infrastructure data.
  • Design automated data workflows, monitoring systems and tooling for continuous visibility.
  • Support high-performance data center networks by isolating and correcting traffic flow, configurations, and fiber connections.
  • Lead incident response, coordinate remediation and post-incident improvements.
  • Build reliability software in Golang with supporting Python or Rust tooling.

Skills

5+ years experience
Live production networks
High-performance computing networks
Software engineering mindset
Incident response leadership
Network protocols & fiber
Data center networking
Golang tooling

Tools

Golang
Python
Rust

Job description

Sr Network Engineer, Reliability and Observability
Location: Remote, United States or London, United Kingdom

US Location Preference: New York, San Francisco, Austin, Seattle, or Phoenix

Our client is scaling high-performance infrastructure that supports advanced artificial intelligence and accelerated computing workloads. They are seeking a Sr Network Engineer, Reliability and Observability who can combine deep production networking expertise with software engineering and a data-driven approach to infrastructure reliability. This engineer will help ensure complex datacenter fabrics remain resilient as the environment grows, while building the automation, telemetry, and engineering processes needed to identify problems earlier and prevent repeat failures.

At this time, our client is not open to C2C employment or sponsoring visas.

This Role Offers:
  • Exposure to sophisticated datacenter fabrics, high-performance networking, optics, and physical infrastructure.
  • Opportunity to build software and automation that improves network operations rather than relying solely on manual processes.
  • Competitive compensation with equity participation and a comprehensive benefits package.
Focus:
  • Drive network reliability programs by turning telemetry, operational trends, incident history, and infrastructure data into measurable improvements in availability and performance.
  • Design automated data workflows, monitoring systems, and operational tooling that provide continuous visibility into network health, service performance, and recurring failure patterns.
  • Support high-performance data center networking environments by isolating and correcting issues involving traffic flow, network design, system configurations, switching infrastructure, and physical connections.
  • Take technical ownership during complex production incidents, systematically isolate failure domains, coordinate response efforts, and carry issues through remediation and follow-up improvements.
  • Build scalable infrastructure and reliability software primarily in Golang, with Python or Rust used for supporting tools and automation, while applying disciplined development and testing practices.
  • Collaborate with deployment, datacenter operations, hardware, logistics, and software teams to improve infrastructure lifecycle processes, remediation workflows, and operational readiness.
  • Serve as a technical resource across several critical infrastructure domains, bringing specialized knowledge in areas such as network communications, high-speed connectivity, fiber systems, data transport, or power technologies.
  • Translate ambiguous infrastructure challenges into defined objectives, Jira initiatives, development pipelines, working code, documentation, and sustainable operational improvements.
Skill Set:
  • 5+ years of professional experience supporting network infrastructure, with at least 3 years focused on hands-on operational responsibilities within live production or high-performance computing environments.
  • Extensive experience supporting sophisticated data center networks, including routed and switched architectures, overlay technologies, dynamic routing, large-scale fabric designs, configuration troubleshooting, and connectivity issues across both logical and physical infrastructure.
  • Strong software engineering background, including experience with ITIL, Agile/xP, and TDD methodologies and development of hyperscale platforms using Golang with Python or Rust tooling.
  • Proven ability to lead incident response, troubleshoot complex infrastructure failures methodically, communicate effectively during outages, and maintain ownership through resolution.
  • Demonstrated technical depth across multiple infrastructure disciplines, with strong expertise in at least two areas such as network protocols, fiber technologies, transport systems, high-performance connectivity, or power infrastructure.
  • Proven ability to independently turn broad technical goals into structured plans and carry them through development, validation, implementation, and final delivery.
  • Strong operational judgment with the ability to prioritize effectively and make sound technical decisions when working with incomplete information or time-sensitive production issues.
  • Effective communication and documentation skills with the ability to convert technical findings and incident lessons into repeatable engineering practices.
  • Additional value placed on experience supporting high-performance computing infrastructure, advanced network technologies, monitoring and analytics platforms, data-driven troubleshooting, or hands-on data center operations.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Network Site Reliability Engineer (SREs)
Sr. Network Site Reliability Engineer (SREs)

Technopride Ltd • Greater London

On-site
GBP 70,000 - 90,000
Network SRE - DC – Network WAN
Network SRE - DC – Network WAN

Technopride Ltd • Greater London

On-site
GBP 111,000 - 185,000
Senior Network Engineer
Senior Network Engineer

IT WORLD LIMITED • England

On-site
GBP 70,000 - 78,000
Senior Network Engineer
Senior Network Engineer

UST • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Network Engineer – Data Center, Cloud
Senior Network Engineer – Data Center, Cloud

Jobtailor • Greater London

On-site
GBP 90,000 - 130,000
Senior Network Reliability Engineer Remote Observability
Senior Network Reliability Engineer Remote Observability

Blue Signal Search • Greater London

On-site
GBP 90,000 - 140,000
Equity participation
Comprehensive benefits package
Site Reliability Engineer
Site Reliability Engineer

ITAC Solutions • Birmingham

On-site
GBP 93,000 - 110,000
Work on cutting-edge distributed systems
Collaborate on cloud transformations
High-impact role with visibility across technology leadership
Lead Network Reliability Engineer
Lead Network Reliability Engineer

G-Research • Greater London

On-site
GBP 60,000 - 80,000
Highly competitive compensation plus annual discretionary bonus
Lunch provided
30 days' annual leave
+3
Infrastructure Tooling & Observability Engineer( UK)
Infrastructure Tooling & Observability Engineer( UK)

Radiant • Greater London

On-site
GBP 90,000 - 120,000
Senior Network Engineer_Cloud, Security & Infra (Architect II - Cloud Infrastructure Services)
Senior Network Engineer_Cloud, Security & Infra (Architect II - Cloud Infrastructure Services)

UST • Greater London

Hybrid
GBP 90,000 - 130,000