Senior Network Reliability Engineer

Gainbridge

United States

Remote

USD 135,000 - 190,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Dental insurance
401(k) plan with company matching
Employee assistance program

Job summary

Gainbridge is looking for a Sr. Network Reliability Engineer to lead its Site Reliability Engineering practice. This role focuses on building a robust network reliability framework across multi-cloud environments using SRE principles and automation tools like Terraform and Ansible.

The ideal candidate will have extensive experience in TCP/IP and SD-WAN architecture, along with a strong background in monitoring tools and automation. A commitment to continuous improvement and operational excellence is essential.

Qualifications

  • Proven experience with Terraform and Ansible in a production environment.
  • Strong proficiency in Python for automation and API interaction.
  • Hands-on experience with enterprise firewalls and security scales.

Responsibilities

  • Define SLOs and error budgets for the network platform.
  • Lead postmortems focusing on permanent remediation.
  • Move network state into code using Terraform and Ansible.

Skills

TCP/IP
BGP
OSPF
VPNs
SD-WAN architecture
Terraform
Ansible
Python
Cloudflare
Datadog

Tools

Grafana
Prometheus
Zscaler
eBPF

Job description

Group 1001 is a consumer‑centric, technology‑driven family of insurance companies on a mission to deliver outstanding value and operational performance by combining financial strength and stability with deep insurance expertise and a can‑do culture.

Why This Role Matters

The Platform Engineering Services team at Group 1001 is building a Site Reliability Engineering practice with a network scope. We’re hiring an Sr. Network Reliability Engineer who embodies Innovation and Excellence, and will apply SRE principles — code‑as‑source‑of‑truth, SLOs and error budgets, alerting on symptoms rather than causes, failure‑mode‑first design, and the elimination of toil — to the firm’s network platform from carrier edge through cloud fabric to Kubernetes pod boundary.

This is not a “keep the lights on” role. You will systematically engineer the lights‑on work out of existence, build the abstractions that let other engineering teams express network intent in code, and treat the network as a single engineered system rather than a collection of vendor consoles.

You will operate inside a DevSecOps practice spanning multi‑cloud, multi‑region environments, and you will partner closely with Cloud and Data Platforms, the NOC/SOC, and Cyber Security to extend reliability practice across the firm.

How You’ll Contribute
  • Define SLOs and error budgets for the network platform — DNS resolution, edge availability, mesh ingress success, cross‑region path health — and use them to gate changes, not just to color dashboards.
  • Lead postmortems with a focus on permanent remediation, not pattern‑recognition.
  • Alert on symptoms users feel, not on causes that may or may not produce impact.
  • Move network state into code using Terraform (or Pulumi), Ansible, and Python to replace CLI‑driven configuration with declarative, version‑controlled, peer‑reviewed change running through Infra CI/CD.
  • Build network policy as intent, not rule lists; express permitted flows, isolated segments, inspected egress, and DNS zone sharing; engineer compilers that translate intent into per‑vendor configuration.
  • Use Policy as Code (OPA/Rego, Sentinel, Cilium NetworkPolicy) to catch invariant violations at plan time, not apply time.
  • Design, deploy, and manage network infrastructure using Terraform or Ansible, moving the firm away from manual configuration to a code‑first approach.
  • Engineer the cloud network platform: operate and extend our multi‑account AWS Landing Zone, build platform abstractions for new accounts or services with declarative inputs.
  • Extend platform thinking into the container tier, covering Kubernetes networking, service mesh (Istio, Linkerd, Consul Connect), eBPF‑based observability and policy (Cilium, Hubble), and integration points where mesh‑level authz meets cloud‑tier identity.
  • Improve telemetry and observability with intent; build alerts as structured payloads with runbook links, suspected blast radius, and dependency‑aware suppression.
  • Author system‑health dashboards for operators and end‑user monitoring dashboards that reflect actual user experience; use Grafana, Elastic, Open Telemetry where each fits.
  • Mentor and grow the team, provide technical guidance to junior engineers, foster a culture of learning, and work out loud across Platform Engineering so the patterns you build cross‑pollinate to adjacent domains.
  • Handle hardware when required; provide maintenance and configuration support for routers, switches, and firewalls at data centers and offices; bring code‑first practices to physical hardware where possible.
  • Serve as an escalation point for network issues, perform troubleshooting with a focus on root cause analysis and permanent remediation, author runbooks and SOPs for the NOC.
What We’re Looking For
  • Deep understanding of TCP/IP, BGP, OSPF, VPNs, and SD‑WAN architecture.
  • Proven experience with Terraform (state management, modules) and Ansible (playbooks, roles) or similar in a production environment.
  • Proficiency in Python for automation and API interaction.
  • Hands‑on experience with Cloudflare, Zscaler, and/or enterprise firewalls.
  • Experience configuring monitoring tools (Datadog, Prometheus, Grafana) to create meaningful alerts and dashboards.
  • Nice to have: Service mesh experience (Istio, Linkerd, Consul Connect, Cilium), eBPF‑based observability (Hubble, Pixie), AWS multi‑account landing zone tooling (AFT, Control Tower, or equivalent).
  • Nice to have: Policy as Code experience (OPA/Rego, Sentinel, Cilium NetworkPolicy).
  • Strong belief that documentation is required for all work.
  • Toil‑reduction mindset, actively seeking to automate repetitive tasks.
  • Hybrid capability: willingness to handle physical hardware tasks when required while maintaining a software‑centric engineering mindset.
Compensation

The base pay for this position ranges from $135,000 per year in our lowest geographic market to $190,000 per year in our highest geographic market. Pay is based on factors such as market location, job‑related skills, and experience.

Benefits Highlights

Employees who work 30 hours or more weekly are eligible to enroll in Group 1001’s benefits package, which includes health, dental, eye, life, disability insurance, employee assistance program, wellness programs, and a 401(k) plan with company matching.

Diversity

Group 1001 is strongly committed to providing a supportive work environment where employee differences are valued. Diversity is an essential ingredient in building a welcoming place to work and a high‑performance team. All employees share the responsibility for maintaining a workplace culture of dignity, respect, and appreciation of individual and group differences.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Reliability Engineer
Senior Network Reliability Engineer

Group 1001 • Indianapolis (IN)

On-site
USD 135,000 - 190,000
Health insurance
Dental insurance
Vision insurance
+4
Senior Network Reliability Engineer
Senior Network Reliability Engineer

Gainbridge • Zionsville (IN), Northern (KY)

On-site
USD 135,000 - 190,000
Health Insurance
Dental Insurance
Vision Insurance
+4
Network Reliability Engineer
Network Reliability Engineer

Cloudflare • Austin (TX)

Hybrid
USD 90,000 - 120,000
Senior Customer Reliability Engineer
Senior Customer Reliability Engineer

Cloudflare • United States

On-site
USD 95,000 - 130,000
Medical/Rx Insurance
Dental Insurance
Vision Insurance
+2
Network Deployment Engineer
Network Deployment Engineer

Triwill Group • United States

Hybrid
USD 114,000 - 173,000
Equity plan
Comprehensive benefits package
Network Deployment Engineer
Network Deployment Engineer

CloudFlare • Austin (TX)

On-site
USD 110,000 - 170,000
Health insurance
401(k)
Principal Software Engineer: Distributed Systems (Config, Test, & Deployment)
Principal Software Engineer: Distributed Systems (Config, Test, & Deployment)

Cloudflare • Austin (TX)

On-site
USD 200,000 - 250,000
Medical/Rx Insurance
Dental Insurance
Vision Insurance
+4
Senior Systems Engineer
Senior Systems Engineer

Cloudflare • Seattle (WA)

On-site
USD 185,000 - 254,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Principal Software Engineer: Distributed Systems (Config, Test, & Deployment)
Principal Software Engineer: Distributed Systems (Config, Test, & Deployment)

Cloudflare • Washington

On-site
USD 220,000 - 275,000
Medical/Rx Insurance
Dental Insurance
Vision Insurance
+5
Network Deployment Engineer
Network Deployment Engineer

Cloudflare • Denver (CO)

On-site
USD 114,000 - 157,000
Equity participation
Health benefits
401(k)
+1