Customer Reliability Engineer (CRE)

Arista Networks

Santa Clara (CA)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Arista Networks is seeking a mid-level Customer Reliability Engineer to enhance our NDR platform’s performance and reliability. This pivotal role involves stabilizing existing systems and developing automated solutions to streamline operations.

With a strong background in Site Reliability Engineering and DevOps, you will work closely with engineering teams and customers to implement durable systems that prioritize ease and reliability. You must be ready to manage on-call responsibilities while also focusing on automation through coding and tooling enhancements.

Qualifications

  • 3+ years of experience in Site Reliability Engineering or DevOps.
  • Strong command of Linux systems and networking fundamentals.
  • Experience working directly with customers on technical issues.
  • Production experience with AWS and Terraform for automation.
  • Proficiency in Python or Go for automation tasks.

Responsibilities

  • Stabilize and map operational workloads with the team.
  • Automate operational tasks and build tooling.
  • Develop strategies for updating systems in isolated environments.
  • Monitor and debug complex production incidents.
  • Act as the on-call person for mission-critical systems.

Skills

Site Reliability Engineering
Linux systems administration
Networking fundamentals (TCP/IP, DNS, routing)
Customer engagement
AWS (VPC, EC2, IAM, S3)
Terraform
CI/CD pipelines
Python
Shell scripting (Bash)
Observability techniques

Job description

Job Description

Who You’ll Work With

Arista's Network Detection and Response (NDR) platform is a mission‑critical security tool for our customers. Its reliability is paramount. We are hiring a mid‑level, Customer Reliability Engineer (CRE) to join our team. This role is critical to the evolution of our customer‑facing infrastructure and operational posture.

What You’ll Do

This is not a traditional operations role. You will inherit a set of critical, manual, and hands‑on operational responsibilities essential to our customers' success. We need you to help with the effort to systematically dismantle this operational burden through automation, tooling, and systems. You will have a collaborative team of excellent engineers to work with.

The short‑term needs are: manual deployments, reactive troubleshooting, and on‑call escalations. But we need you to help us build a system where programmatic solutions have replaced human intervention. You must have the pragmatism to manage the current reality and the systematic impatience and technical skill to build its replacement.

Success in this role requires a dual mindset. You must be a skilled incident leader who can stabilize a crisis and a deliberate systems architect who can prevent the next one. You will work closely with our internal tools, platform, and product engineering teams to channel your direct operational knowledge into durable, long‑term solutions.

Your First Year and Beyond

Your work will follow a deliberate trajectory from reactive execution to proactive design.

Phase 1: Stabilize and Map - You will embed with the team, taking on the existing operational workload alongside the other customer SRE team members covering the USA and India time zones. This includes customer deployments, upgrades, and incident response. You will be expected to go on‑site for our air‑gapped customers, occasionally, to assist on‑prem deployments. Your initial goal is to achieve stability while mapping the landscape of our operational toil.

Phase 2: Automate and Influence - Armed with your map of toil, you will begin to automate. You will write code, build tooling, and deploy declarative infrastructure to eliminate the most critical operational burdens. For larger projects, you will act as a primary stakeholder, providing clear requirements to our internal tooling and platform teams and ensuring their solutions meet the operational need. Your success will be measured by a demonstrable reduction in the overall support effort, fewer pages, support escalations, and manual tasks.

Qualifications

DevOps and SRE Proficiency-You must have a strong 3+ years of background in Site Reliability Engineering or a closely related DevOps function. You also have a strong command of Linux systems administration and possess an understanding of networking fundamentals (TCP/IP, DNS, routing).

Customer‑Facing Experience-You must have experience working directly with external customers to solve difficult technical problems. Your communication must be clear, empathetic, and precise. You are comfortable developing and executing strategies for updating systems in isolated environments where traditional internet‑based tools are unavailable.

Cloud Infrastructure Expertise- You should have production experience with AWS (VPC, EC2, IAM, S3) and a proven track record of using Terraform and CI/CD pipelines to automate the delivery of infrastructure and software updates to remote or secure environments.

Monitoring and Observability- You will be responsible for both building and using our observability stack. This requires hands‑on experience instrumenting applications and managing the telemetry pipelines for metrics, logs, and traces. A core part of the role is then applying this data to debug complex production incidents, understand system behavior, and define SLOs.

Automation and Software Development- You must be proficient in writing code to automate operational tasks. Expertise in a high‑level language like Python or Go is required, as are strong shell scripting skills (e.g., Bash). We have a diverse tech stack including Python, Scala, C, C++, Haskell, Rust, PureScript, etc which requires experience with monitoring and debugging a complex system using system tools, command line utilities, networking debug tools, and filtering complex logs.

Operational Ownership- You must take pride as the On‑call/Directly responsible person for mission‑critical systems, prioritizing root‑cause analysis and durable, system‑level fixes and proactively collaborating with Product teams to directly build the automated solutions that resolve operational challenges permanently.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Customer Reliability Engineer
Senior Customer Reliability Engineer

United States Digital Space LLC • United States

Hybrid
USD 100,000 - 130,000
Health insurance
Flexible working hours
Opportunity for professional development
Customer Reliability Engineer - Automate Incidents
Customer Reliability Engineer - Automate Incidents

Arista Networks • Santa Clara (CA)

On-site
USD 100,000 - 130,000
Sr Software Engineer - Reliability Engineering
Sr Software Engineer - Reliability Engineering

Cox Enterprises • Village of North Hills (NY)

On-site
USD 150,000 - 185,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

Hybrid
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Manager, Site Reliability Engineering
Senior Manager, Site Reliability Engineering

PVH (Tommy Hilfiger/Calvin Klein) • United States

On-site
USD 190,000 - 240,000