Site Reliability Engineer (SRE) - Engineering Productivity

Arista Networks, Inc.

Hinoba-an

On-site

PHP 7,509,000 - 11,264,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Arista Networks is looking for a versatile EngProd engineer to maintain and support our rapidly expanding infrastructure and internal user base in a hybrid cloud environment. You will collaborate with software engineers to design, deploy and operate secure, scalable systems that power Arista's product development teams.

The role emphasizes automation, reliability and observability, using tools such as Ansible, Jenkins, Kubernetes, Grafana, Artifactory and more.

Qualifications

  • At least a 5-year track record in software or infrastructure engineering.
  • Experience with Go, Python or shell scripting to automate workflows.
  • Strong Linux administration and debugging skills.
  • Hands-on experience operating scalable software systems.
  • Proficiency with infrastructure-as-code and provisioning.

Responsibilities

  • Build, deploy and operate critical production systems focusing on scalability, reliability and security.
  • Monitor and enhance developer experience across services.
  • Develop automation to reduce toil and improve efficiency.
  • Respond to alerts and maintain incident runbooks.
  • Design and deploy scalable systems with observability as a primary goal.
  • Collaborate with product teams to resolve infrastructural bottlenecks.

Skills

Go
Python
Shell scripting
Linux administration
Infrastructure-as-code

Education

BSc Computer Science or Engineering
MS Computer Science or Engineering

Tools

Kubernetes
Ansible
Jenkins
Grafana
Artifactory
Gerrit
MySQL
ElasticSearch
Google Cloud
Perforce

Job description

  • Full-time
Company Description

Arista Networks is an industry leader in data-driven, client-to-cloud networking for large data center, campus and routing environments. Arista is a well-established and profitable company with over $7 billion in revenue. Arista's award-winning platforms, ranging in Ethernet speeds up to 800G bits per second, redefine scalability, agility, and resilience. Arista is a founding member of the Ultra Ethernet consortium. We have shipped over 20 million cloud networking ports worldwide with CloudVision and EOS, an advanced network operating system. Arista is committed to open standards, and its products are available worldwide directly and through partners.

At Arista, we value the diversity of thought and perspectives each employee brings. We believe fostering an inclusive environment where individuals from various backgrounds and experiences feel welcome is essential for driving creativity and innovation.

Our commitment to excellence has earned us several prestigious awards, such as the Great Place to Work Survey for Best Engineering Team and Best Company for Diversity, Compensation, and Work-Life Balance. At Arista, we take pride in our track record of success and strive to maintain the highest quality and performance standards in everything we do.

Job Description

Who You'll Work With

Arista Networks is looking for a skilled professional for our Engineering Productivity (EngProd) team to help maintain and support our rapidly expanding infrastructure and internal user base. The ideal candidate is someone who can wear many hats, is versatile and is enthusiastic about learning new technologies. As a part of the software engineering team, you will work with other team members to design, build and administer secure, scalable and fault-tolerant tools and infrastructure in a hybrid cloud environment.

Working in the EngProd group, you will collaborate and work with other engineers to design, build, scale, and operate the systems used by Arista's product development teams. Thes systems are based on industry-standards, including Ansible, Artifactory, Gerrit, Jenkins, Kubernetes, Grafana, Spinnaker, MySQL, ElasticSearch, Google Cloud, Varnish, Perforce, Gerrit etc, 3rd party storage appliances, as well as internal systems developed from the ground-up to automate CI/CD, testing, analysis, and visualization.

What You'll Do

  • Build, deploy safely and incrementally and operate critical production systems with focus on scalability, reliability, observability, performance and security.
  • Monitor, support and enhance developer experience across services.
  • Build automation to remove toil and efficiently operate production systems.
  • Proactively monitor, respond to, and enhance alerts and set up automated alert handling
  • Create and maintain the incident response runbooks.
  • Build and deploy new systems with scalability, reliability, and observability as primary requirements
  • Triage platform/infrastructural issues and help Arista software engineers in their triages. Engage with 3rd party vendor support.
  • Deploy new systems in a staged manner
  • Write postmortem documents and build solutions to avoid incidents from repeating.
  • Plan and communicate maintenance windows on production systems.
  • Work with Arista's product development teams to identify infrastructural issues that are causing bottlenecks and limitations in their workflows. Design and implement solutions to resolve them.
  • Survey and adopt best practices around infrastructure/platform to maintain secure, scalable and fault-tolerant systems.
  • Implement solutions to scale the systems
  • Implement fault-tolerance and performance to improve availability of the systems
  • Study the design and sufficient implementation details of OSS systems for better triage and fix resolution.
Qualifications

Essential to have all of the following skills

  • At least BSc Computer Science or Engineering + 5 years' experience, MS Computer Science or Engineering + 5 years' experience, or equivalent work experience.
  • Knowledge of one or more of Go, Python, shell scripting to be able to implement medium complexity automation workflows.
  • Knowledge of Linux (or UNIX) from administration and debugging perspective
  • Hands-on experience in operating software systems (infrastructure, complex applications etc) at scale
  • Experience in server provisioning (esp from storage and networking perspective).
  • Strong problem solving and software troubleshooting skills
  • Experience with infrastructure-as-code

Desirable to have one/more of the following skills

  • Experience with docker and virtualization technologies - kvm, qemu, kata-containers etc
  • Experience managing ElasticSearch clusters
  • Experience managing Artifactory, docker registry etc
  • Experience with infrastructure-as-code frameworks like Ansible
  • Experience managing large Java applications
  • Experience in storage infrastructure management eg: NAS, SAN, Ceph etc
Additional Information

Arista stands out as an engineering-centric company. Our leadership, including founders and engineering managers, are all engineers who understand sound software engineering principles and the importance of doing things right.

We hire globally into our diverse team. At Arista, engineers have complete ownership of their projects. Our management structure is flat and streamlined, and software engineering is led by those who understand it best. We prioritize the development and utilization of test automation tools.

Our engineers have access to every part of the company, providing opportunities to work across various domains. Arista is headquartered in Santa Clara, California, with development offices in Australia, Canada, India, Ireland, and the US. We consider all our R&D centers equal in stature.

Join us to shape the future of networking and be part of a culture that values invention, quality, respect, and fun.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - SRE (Platform Software Team)
Site Reliability Engineer - SRE (Platform Software Team)

Arista Networks, Inc. • Hinoba-an

On-site
PHP 900,000 - 1,800,000
Senior Systems Engineer (Pre Sales)
Senior Systems Engineer (Pre Sales)

Arista Networks, Inc. • Hinoba-an

On-site
PHP 1,653,000 - 2,975,000
Site Reliability Engineer, EngProd: Scale Secure Systems
Site Reliability Engineer, EngProd: Scale Secure Systems

Arista Networks, Inc. • Hinoba-an

On-site
PHP 7,509,000 - 11,264,000
Platform SRE: Scalable, Fault-Tolerant Infra
Platform SRE: Scalable, Fault-Tolerant Infra

Arista Networks, Inc. • Hinoba-an

On-site
PHP 900,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Philippines (PHC1) Avid Philippines • Philippines

Remote
PHP 1,200,000 - 1,800,000
Health & life insurance
Referral rewards
Generous leave policies
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Site Reliability Engineer
Site Reliability Engineer

Alsons/AWS Information Systems Inc. • Cebu City

Hybrid
PHP 600,000 - 1,000,000
Senior Site Reliability Engineer (AWS)
Senior Site Reliability Engineer (AWS)

Broadridge • Metro Manila

On-site
PHP 1,200,000 - 2,400,000
Staff Site Reliability Engineer – Cloud Efficiency
Staff Site Reliability Engineer – Cloud Efficiency

Super • España

On-site
PHP 1,200,000 - 1,600,000
Medical / Health Insurance
Employee Assistance Programme
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Vestas Services Philippines Inc. • Philippines

On-site
PHP 900,000 - 1,500,000
Global platform exposure
Engineering leadership
Fast-growing team