Site Reliability Engineer - Production Support

Bayside Solutions

Cupertino (CA)

Remote

USD 83,000 - 96,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Bayside Solutions in Cupertino, CA is seeking a Site Reliability Engineer - Production Support for a W2 contract role. This remote position requires hands-on production support, incident management, and a focus on service reliability across cloud and data center deployments.

You will monitor health using Splunk, Datadog, and logs; manage incidents through the IM team; and work with Docker, Kubernetes, AWS, and Java in a fast-paced environment with strict SLAs.

Qualifications

  • Hands-on production support in SLA-bound environments with real incident history.
  • Experience with incident management, triage, escalation, RCA.
  • Proficient with Splunk and command-line troubleshooting.
  • Proficient in Linux/Unix and containerized workloads.
  • Cloud experience focusing on AWS; Azure/AKS acceptable.
  • Experience with payments/banking production environments is a plus.

Responsibilities

  • Provide production and application support to ensure service reliability.
  • Watch service health across production and drive issues to resolution.
  • Manage incidents and lifecycle with IM team, meet SLAs.
  • Monitor with Splunk and Datadog; investigate anomalies.
  • Support cloud and datacenter deployments; assist rollbacks.
  • Maintain in-house CD deployment process and scripts.
  • Keep Git repos up to date and debug deployment scripts.

Skills

Splunk
Linux/Unix
Docker
Kubernetes
AWS
Datadog
Java
Incident management
NOC
SLAs

Tools

Datadog
Jira
ServiceNow ITSM
Git

Job description

Site Reliability Engineer - Production Support

W2 Contract

Pay Rate: $60 - $70 per hour

Location: Cupertino, CA - Remote Role

Duties and Responsibilities:
  • Production support / application support / NOC / service reliability / incident management. Screen out: security analysts whose Splunk work is log-based threat detection.
  • Watch service health across production; catch, triage, and drive issues to resolution.
  • Run production support against SLAs and work the incident lifecycle directly with the incident management (IM) team.
  • Monitor and investigate using Splunk (and Datadog, where in play).
  • Support the application stack in cloud and datacenter environments.
  • Non-prod and prod deployments using our in-house CD tool; monitor deployments and roll back / debug when they fail.
  • Keep Git repos current; debug deployment scripts.
Requirements and Qualifications:
  • Hands-on production support/application support in an SLA-bound environment. Must be able to describe what they personally did during a real incident.
  • Incident management experience: triage, severity assessment, escalation, bridge participation, driving to resolution, RCA follow-up. Direct interaction with an IM team.
  • Splunk is required. They will be asked what commands/searches they run.
  • SOC/production monitoring
  • Troubleshooting depth demonstrable at the command line: log tracing, connection failures, latency, service crashes. Conceptual answers fail here.
  • Linux/Unix fluency.
  • Containerized workloads: Docker + Kubernetes (managing, not just using)
  • Cloud: AWS primary; Azure/AKS acceptable and precedented. Datadog - Naveen called it out as a plus point. New to this team's stack vs. prior reqs in Lighthouse is likely a growing surface.
  • Payments/banking production environment (card networks, Visa/Mastercard, wallets, issuer/acquirer processing).
  • Java application troubleshooting.
Preferred Qualifications
  • CI/CD tooling - Jenkins, specifically, can debug the scripts behind a deployment, not just trigger one.
  • Python and/or shell scripting for automation and runbook work.
  • Load balancers, SSL/TLS and certificate management, DNS, and general network troubleshooting.
  • Config management (Ansible/Chef/Puppet), autoscaling design in Kubernetes.
  • Terraform / infrastructure-as-code exposure.
  • ServiceNow / Jira or equivalent ITSM incident tooling; ITIL familiarity.

Bayside Solutions, Inc. is not able to sponsor any candidates at this time. Additionally, candidates for this position must qualify as a W2 candidate.

Bayside Solutions, Inc. may collect your personal information during the position application process. Please reference Bayside Solutions, Inc.'s CCPA Privacy Policy at

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Site Reliability Engineer - Production Support
Remote Site Reliability Engineer - Production Support

Bayside Solutions • Cupertino (CA)

Remote
USD 83,000 - 96,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Talentify • Woonsocket (RI)

Hybrid
USD 150,000 - 190,000
Health benefits
Referral program
Growth opportunities
Site Reliability Engineer
Site Reliability Engineer

Skill • Southlake (TX)

On-site
USD 66,000 - 73,000
Health insurance
Vision insurance
Dental insurance
+2
Graphic Production Artist
Graphic Production Artist

PRI Technology • Holmdel Township (NJ)

On-site
USD 110,208 - 117,096
Site Reliability / Production Engineer
Site Reliability / Production Engineer

GCS Recruitment Specialists • Cherry Hill Township (NJ)

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

On-site
USD 150,000 - 155,000
Site Reliability Engineer
Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

On-site
USD 120,000 - 155,000
Platform Engineer & Production Support
Platform Engineer & Production Support

Strategic Staffing Solutions • Charlotte (NC)

On-site
USD 100,000 - 150,000
Infrastructure/Cloud DevOps - SRE
Infrastructure/Cloud DevOps - SRE

Bayside Solutions • Cupertino (CA)

On-site
USD 150,000 - 230,000