Site Reliability Engineer

Veritas Search Group

Tustin (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Veritas Search Group is seeking an experienced Site Reliability Engineer to join the Product Platform team in an onsite capacity near Tustin, CA. You will bridge product development and platform engineering, translating infrastructure needs into scalable, reliable platform solutions.

The ideal candidate has deep Kubernetes and AWS experience, strong observability skills, and proficiency in Python scripting.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field with 4+ years of relevant experience or 6+ years of equivalent professional experience in lieu of a degree.
  • Deep hands-on experience with Kubernetes, including cluster creation, administration, deployments, networking, and troubleshooting.
  • Strong hands-on experience with AWS, including creating and maintaining cloud infrastructure and resources (S3, RDS).
  • Solid understanding of observability, monitoring, logging, metrics, and application performance concepts.
  • Experience with observability platforms such as Datadog, Splunk, Grafana, Prometheus, or similar tools.
  • Proficiency in Python for scripting, automation, or infra tasks.
  • Experience with CI/CD and GitOps workflows; strong root-cause analysis skills.
  • Excellent communication and collaboration across product, infra, cloud, networking, and security teams.

Responsibilities

  • Build, manage, and troubleshoot Kubernetes clusters and containerized environments.
  • Collaborate with product and application teams to translate infrastructure needs into platform solutions.
  • Serve as first point of contact for infrastructure issues; perform root‑cause analysis across layers.
  • Coordinate with platform, cloud, network, and security teams for cross-ownership issues.
  • Create, configure, and maintain AWS resources supporting platforms and apps.
  • Support and enhance an internal observability platform monitoring services and infra.
  • Onboard new apps to the observability platform with telemetry requirements.
  • Develop and maintain Python scripts for automation and troubleshooting.
  • Support CI/CD and GitOps deployment processes for Kubernetes environments.
  • Improve platform reliability, scalability, monitoring, and developer experience.
  • Engage in root-cause investigations and document platform standards and procedures.

Skills

Kubernetes
AWS
Observability
Troubleshooting
Cross‑functional collaboration
Python scripting
Communication

Education

Bachelor's degree in Computer Science, Engineering, IT

Tools

Datadog
Splunk
Grafana
Prometheus
S3
RDS
Argo CD
Helm
Kubernetes (EKS)
Terraform
Ansible

Job description

This role requires candidates who are currently authorized to work in the U.S. without sponsorship, and C2C arrangements are not accepted. This role is onsite near Tustin, CA.
Job Description

We are seeking an experienced Site Reliability Engineer (SRE) to join our Product Platform team. This role will serve as a critical bridge between product development teams and the platform engineering organization, helping translate application infrastructure needs into scalable, reliable platform solutions.

The ideal candidate has deep, hands-on experience with Kubernetes, AWS, and observability, along with strong troubleshooting and cross-functional communication skills. This is an infrastructure-focused engineering role for someone who is comfortable working directly with product teams, diagnosing complex application and platform issues, and driving problems through resolution across multiple technical teams.

Responsibilities
  • Build, manage, maintain, and troubleshoot Kubernetes clusters and containerized application environments.
  • Partner closely with product and application development teams to understand infrastructure requirements and translate them into actionable platform engineering needs.
  • Serve as a first point of contact for infrastructure and reliability issues affecting product teams, performing root‑cause analysis across application, Kubernetes, cloud, networking, and security layers.
  • Resolve Kubernetes and platform‑related issues directly while coordinating with other infrastructure, cloud, network, and security teams when issues fall outside the platform team's ownership.
  • Create, configure, and maintain AWS resources supporting application and platform environments.
  • Support and enhance an internal observability platform used to monitor applications, services, and infrastructure.
  • Onboard new applications and use cases into the observability platform by partnering with technical teams to understand monitoring and telemetry requirements.
  • Develop and maintain Python scripts used for infrastructure automation, troubleshooting, platform operations, and observability.
  • Support CI/CD and GitOps‑based deployment processes for Kubernetes environments.
  • Improve platform reliability, scalability, monitoring, operational efficiency, and developer experience.
  • Participate in troubleshooting and root‑cause investigations involving multiple engineering teams and drive issues through successful resolution.
  • Document platform standards, troubleshooting procedures, operational processes, and technical solutions.
Required Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field with 4+ years of relevant experience, or 6+ years of equivalent professional experience in lieu of a degree.
  • Deep hands‑on experience with Kubernetes, including building clusters from scratch, cluster administration, deployments, networking, and troubleshooting.
  • Strong hands‑on experience with AWS, including creating and maintaining cloud infrastructure and resources.
  • Experience working with AWS services such as S3 and RDS.
  • Strong understanding of observability, monitoring, logging, metrics, and application performance concepts.
  • Experience with observability platforms such as Datadog, Splunk, Grafana, Prometheus, or similar technologies.
  • Experience using Python for scripting, automation, or infrastructure‑related tasks.
  • Strong troubleshooting and root‑cause analysis skills across complex application and infrastructure environments.
  • Excellent verbal and written communication skills with the ability to work effectively across product, application, infrastructure, cloud, networking, and security teams.
  • Ability and willingness to work in a highly collaborative position that combines hands‑on engineering with significant cross‑team coordination.
Preferred Qualifications
  • Experience with Helm for Kubernetes application packaging and deployment.
  • Experience with Argo CD and GitOps‑based deployment practices.
  • Experience building or maintaining CI/CD pipelines.
  • Experience with Prometheus, Grafana, and OpenTelemetry.
  • Experience with Amazon EKS and IAM.
  • Familiarity with Istio or other service mesh technologies.
  • Experience with Kafka or Amazon MSK.
  • Experience with infrastructure‑as‑code and configuration‑management technologies such as Terraform or Ansible.
  • Experience supporting internal developer platforms, infrastructure platforms, or observability products.
  • Experience onboarding development teams or applications onto centralized platform services.
  • Experience working in regulated or highly controlled technology environments.
  • Interest in learning and taking ownership of application code supporting internal platform tooling.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer (SRE) – Evening Shift
Site Reliability Engineer (SRE) – Evening Shift

Peraton • Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Java SRE Engineer
Java SRE Engineer

EITACIES Inc. • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000