SRE: Kubernetes, AWS & Observability

Veritas Search Group

Tustin (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Veritas Search Group is seeking an experienced Site Reliability Engineer to join the Product Platform team in an onsite capacity near Tustin, CA. You will bridge product development and platform engineering, translating infrastructure needs into scalable, reliable platform solutions.

The ideal candidate has deep Kubernetes and AWS experience, strong observability skills, and proficiency in Python scripting.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field with 4+ years of relevant experience or 6+ years of equivalent professional experience in lieu of a degree.
  • Deep hands-on experience with Kubernetes, including cluster creation, administration, deployments, networking, and troubleshooting.
  • Strong hands-on experience with AWS, including creating and maintaining cloud infrastructure and resources (S3, RDS).
  • Solid understanding of observability, monitoring, logging, metrics, and application performance concepts.
  • Experience with observability platforms such as Datadog, Splunk, Grafana, Prometheus, or similar tools.
  • Proficiency in Python for scripting, automation, or infra tasks.
  • Experience with CI/CD and GitOps workflows; strong root-cause analysis skills.
  • Excellent communication and collaboration across product, infra, cloud, networking, and security teams.

Responsibilities

  • Build, manage, and troubleshoot Kubernetes clusters and containerized environments.
  • Collaborate with product and application teams to translate infrastructure needs into platform solutions.
  • Serve as first point of contact for infrastructure issues; perform root‑cause analysis across layers.
  • Coordinate with platform, cloud, network, and security teams for cross-ownership issues.
  • Create, configure, and maintain AWS resources supporting platforms and apps.
  • Support and enhance an internal observability platform monitoring services and infra.
  • Onboard new apps to the observability platform with telemetry requirements.
  • Develop and maintain Python scripts for automation and troubleshooting.
  • Support CI/CD and GitOps deployment processes for Kubernetes environments.
  • Improve platform reliability, scalability, monitoring, and developer experience.
  • Engage in root-cause investigations and document platform standards and procedures.

Skills

Kubernetes
AWS
Observability
Troubleshooting
Cross‑functional collaboration
Python scripting
Communication

Education

Bachelor's degree in Computer Science, Engineering, IT

Tools

Datadog
Splunk
Grafana
Prometheus
S3
RDS
Argo CD
Helm
Kubernetes (EKS)
Terraform
Ansible

Job description

Veritas Search Group is seeking an experienced Site Reliability Engineer to join the Product Platform team in an onsite capacity near Tustin, CA. You will bridge product development and platform engineering, translating infrastructure needs into scalable, reliable platform solutions.

The ideal candidate has deep Kubernetes and AWS experience, strong observability skills, and proficiency in Python scripting.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Veritas Search Group • Tustin (CA)

On-site
USD 140,000 - 190,000
Kubernetes Platform SRE | AWS, Istio & Automation
Kubernetes Platform SRE | AWS, Istio & Automation

Okta • Washington

On-site
USD 194,000 - 267,000
Equity
Bonus
Health insurance
+5
Remote Site Reliability Engineer - Observability Expert
Remote Site Reliability Engineer - Observability Expert

Codevertex Innovations • Northern (KY)

Hybrid
USD 140,000 - 195,000
Kubernetes Platform SRE — AWS, Istio & Automation
Kubernetes Platform SRE — AWS, Istio & Automation

Okta • Chicago (IL)

On-site
USD 174,000 - 239,000
Kubernetes Platform SRE — AWS, Istio, Auto-Scaling
Kubernetes Platform SRE — AWS, Istio, Auto-Scaling

Okta • New York (NY)

On-site
USD 140,000 - 190,000
Kubernetes Platform SRE — AWS Cloud Automation Lead
Kubernetes Platform SRE — AWS Cloud Automation Lead

Triwill Group • Washington (IL)

On-site
USD 194,000 - 267,000
Equity
Bonus
Health, dental & vision insurance
+3
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Kubernetes SRE — Cloud Platform Reliability & Automation
Kubernetes SRE — Cloud Platform Reliability & Automation

Okta • Chicago (IL)

On-site
USD 174,000 - 214,000
Equity
Health insurance
Paid leave
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior Observability & SRE Engineer (GCP/Kubernetes)
Senior Observability & SRE Engineer (GCP/Kubernetes)

Ontrac Solutions • New York (NY)

On-site
USD 120,000 - 190,000