Infrastructure/Cloud DevOps - SRE

Bayside Solutions

Cupertino (CA)

On-site

USD 150,000 - 230,000

Full time

11 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Bayside Solutions, Inc. is seeking a highly motivated DevOps / Site Reliability Engineer to support large-scale Kubernetes-based infrastructure and platform operations.

This role is focused on building, automating, and operating highly reliable systems that power critical engineering platforms and services. The role involves operating production environments at scale, developing automation, and improving CI/CD and deployment workflows.

Qualifications

  • Hands-on Kubernetes experience with EKS, GKE or AKS.
  • Experience running Kubernetes at scale in production.
  • Proficient in monitoring/observability with Grafana, Prometheus, or Splunk.
  • Strong scripting/automation using Python or Golang.
  • Experience with CI/CD pipelines and deployment automation.

Responsibilities

  • Design, build, automate, and support scalable Kubernetes platforms and services.
  • Operate and troubleshoot production environments at scale.
  • Develop automation to improve reliability and efficiency.
  • Monitor health, performance, and availability with observability tools.
  • Troubleshoot infrastructure, apps, and networking across distributed systems.
  • Collaborate with engineers to improve deployment and scalability practices.
  • Participate in incident response and root-cause analysis.
  • Improve CI/CD workflows and deployment automation.
  • Document processes and drive operational excellence.
  • Take ownership of projects and deliverables to completion.

Skills

Kubernetes
EKS
GKE
AKS
CI/CD
Python
Golang
Grafana
Prometheus
Splunk
Scripting
SRE basics

Tools

Grafana
Prometheus
Splunk

Job description

We are looking for a highly motivated DevOps / Site Reliability Engineer to support large-scale Kubernetes-based infrastructure and platform operations. This role is focused on building, automating, and operating highly reliable systems that power critical engineering platforms and services.

Duties and Responsibilities:
  • Design, build, automate, and support scalable Kubernetes-based platforms and services
  • Operate and troubleshoot production environments running at scale
  • Develop automation and tooling to improve operational efficiency and reliability
  • Monitor platform health, performance, and availability using observability tooling
  • Troubleshoot infrastructure, application, and networking issues across distributed systems
  • Work closely with engineering teams to improve deployment, reliability, and scalability practices
  • Participate in operational support, incident response, and root cause analysis
  • Improve CI/CD workflows and deployment automation
  • Drive operational excellence through documentation, automation, and process improvements
  • Take ownership of projects and independently drive deliverables to completion
Requirements and Qualifications:
  • Strong hands-on experience with Kubernetes platforms such as:
  • EKS
  • GKE
  • AKS or similar
  • Experience running and supporting applications on Kubernetes at scale
  • Strong understanding of containerized infrastructure and distributed systems
  • Experience with monitoring and observability tools, preferably:
  • Grafana
  • Experience with CI/CD pipelines and deployment automation
  • Experience with Splunk logging, log analysis, and troubleshooting
  • Strong scripting and automation experience using Python and/or Golang
  • Experience troubleshooting production systems under pressure
  • Strong communication and collaboration skills
  • Self-starter mentality with strong ownership and accountability
Preferred Qualifications
  • Experience operating Ray clusters/services
  • Strong networking and troubleshooting experience
  • Experience with cloud infrastructure and platform services
  • Experience with Infrastructure as Code and automation frameworks
  • Familiarity with SRE principles and operational best practices
Desired Skills and Experience

Kubernetes, Amazon EKS, Google Kubernetes Engine (GKE), Azure Kubernetes Service (AKS), DevOps, Site Reliability Engineering (SRE), Cloud Infrastructure, Containerized Infrastructure, Distributed Systems, Platform Engineering, Production Operations, CI/CD, Deployment Automation, Infrastructure as Code (IaC), Python, Golang, Grafana, Prometheus, Splunk, Observability, Log Analysis, Incident Response, Root Cause Analysis, Network Troubleshooting, Ray Clusters, Systems Automation, Performance Monitoring, Scalability, High Availability, Operational Excellence, Technical Documentation, Cross-Functional Collaboration, Project Ownership

Bayside Solutions, Inc. is not able to sponsor any candidates at this time. Additionally, candidates for this position must qualify as a W2 candidate.

Bayside Solutions, Inc. may collect your personal information during the position application process. Please reference Bayside Solutions, Inc.'s CCPA Privacy Policy at www.baysidesolutions.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Plano (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Koitecc Solutions • Chandler (AZ), Northern (KY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Seattle (WA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
SRE/Devops Engineer
SRE/Devops Engineer

INSPYR Solutions • Sunnyvale (CA)

Hybrid
USD 120,000 - 180,000
Work-life balance
No on-call requirements
Standard business hours
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000