Site Reliability Engineer

Harvey Nash

United States

Remote

USD 120,000 - 150,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Harvey Nash in the United States is seeking a Site Reliability Engineer with a strong operational focus to ensure the reliability and performance of critical infrastructure and services.

You will deploy software for Cloud Prem and SAAS customers, respond to incidents, collaborate on root causes, and improve automation and run-books for efficient operations.

Qualifications

  • BS degree in Computer Science or related field.
  • 3+ years of Site Reliability Engineering experience.
  • Experience with cloud platforms, especially AWS.

Responsibilities

  • Deploy software for Cloud Prem and SAAS customers.
  • Respond to and diagnose incidents to minimize downtime.
  • Collaborate with engineers to establish root causes and resolutions.
  • Develop run-books and automation to streamline tasks.
  • Provide tier 2/3 support and on-call coverage.

Skills

Communication skills
Customer service orientation
Incident response

Education

BS degree in Computer Science or related field

Tools

Kubernetes
Helm
Linux
AWS (VPC, networking)
Terraform
GitOps
Prometheus
Grafana

Job description

Job Description: Overview:

We are seeking a highly motivated Site Reliability Engineer (SRE) with a strong operational focus to join our growing team. In this role, you will play a vital role in ensuring the smooth operation and performance of our critical infrastructure and services. You'll work cross-functionally to create alignment and deliver results alongside builders who have helped to shape the success of companies such as Google, Okta, AWS, Snowflake.

What you will do in this role:
  • Deploy software for Cloud Prem and SAAS customers.
  • Respond to and diagnose system incidents in a timely and efficient manner, minimizing downtime and impact on users.
  • Collaborate with other engineers to establish root causes and implement effective resolutions.
  • Continuously improve incident response processes and documentation for future occurrences.
  • Proactively monitor and maintain the health and performance of our infrastructure and services.
  • Perform routine administrative tasks such as system configuration, user management, and data backups.
  • Identify and implement operational improvements to ensure ongoing system reliability and efficiency.
  • Develop and implement scripts and automated solutions to streamline operational tasks and reduce manual workload.
  • Participate in the on-call rotation to address critical incidents outside of regular business hours.
  • Ensure effective handoff between on-call engineers and document post-incident information for future reference.
  • Document processes for support and create, maintain and execute run-books for identified situations
  • Provide tier 2/3 technical support to customers experiencing platform issues or requiring advanced troubleshooting
  • Work directly with customer technical teams to resolve complex deployment, configuration, and integration challenges
  • Conduct technical onboarding sessions and provide guidance on best practices for customer implementations
  • Collaborate with customer success teams to ensure smooth customer experiences and rapid issue resolution
  • Create and maintain customer-facing technical documentation, troubleshooting guides, and knowledge base articles
  • Escalate customer feedback and feature requests to product and engineering teams
  • Participate in customer calls and technical discussions to provide expert-level platform guidance
  • Track and analyze customer support metrics to identify trends and areas for improvement
What you will need to be successful in this role:
  • Education:BS degree in Computer Science or related field
  • Experience:3+ years of experience in Site Reliability Engineering
  • 2+ years experience working with cloud platform and cloud automation tools especially in AWS
  • Strong experience with Kubernetes, Helm, Linux, AWS networking(VPC) and Terraform
  • Experience with the GitOps model for deployment
  • Familiarity with distributed version control
  • Experience with monitoring and alerting tools (e.g., Prometheus, Grafana).
  • Excellent communication skills with ability to explain technical concepts to both technical and non-technical audiences
  • Strong customer service orientation with patience and empathy when working with frustrated customers
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Veritas Search Group • Tustin (CA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • New Jersey

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

On-site
USD 150,000 - 155,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000