Senior Platform Reliability & Automation Engineer

Cacheflow

Mountain View (CA)

Hybrid

USD 185,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Otter.ai in Mountain View seeks an Engineer with extensive system knowledge to build and operate large-scale systems enabling reliable deployment with effective monitoring and resilient operations.

You will own monitoring with Prometheus and Grafana, optimize Linux performance and security, manage infrastructure, and participate in on-call rotations while collaborating with IT, Security, and Engineering teams.

Qualifications

  • Bachelor's in Computer Science or Electrical Engineering (MS preferred).
  • 5+ years experience in SRE/Production Engineering, and IT Systems / Enterprise IT / Systems Engineering / DevOps-internal tooling, supporting internal customers.
  • Expert level experience architecting, developing, and troubleshooting large scale systems.
  • Advanced level proficiency with one or more programming languages (i.e. Python, Golang).
  • Extensive experience with CI/CD pipelines and infrastructure as code (Terraform, Ansible).
  • Strong familiarity with AWS services (i.e. ECS, S3, ALB, VPC).
  • Knowledge of containers and orchestration using Kubernetes.
  • Experience building production quality cloud infrastructure that enables reliable and rapid deployment of large-scale systems with effective monitoring and resilient operations.
  • Thrives in a fast paced startup environment.
  • Proven track record taking on projects from inception to launch.
  • Solid troubleshooting fundamentals across macOS/Windows/Linux, networking basics, and security hygiene.
  • Ability to work through ambiguous problems with IT, Security, and Engineering stakeholders.
  • Identity and access workflows (Okta/Entra, Google Workspace/M365), provisioning automation.

Responsibilities

  • Own, design and implement monitoring systems such as Prometheus and Grafana.
  • Manage and maintain infrastructure.
  • Own configuration management processes and build product features as appropriate.
  • Investigate and diagnose issues by digging into data and collaborating with engineers.
  • Participate in on-call rotation.
  • Help automate CI and testing processes to enable scale.

Skills

Python
Golang
CI/CD pipelines
SRE foundations
Cloud fundamentals

Education

Bachelor's in CS or EE
MS preferred

Tools

Terraform
Ansible

Job description

Otter.ai in Mountain View seeks an Engineer with extensive system knowledge to build and operate large-scale systems enabling reliable deployment with effective monitoring and resilient operations.

You will own monitoring with Prometheus and Grafana, optimize Linux performance and security, manage infrastructure, and participate in on-call rotations while collaborating with IT, Security, and Engineering teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Scale, Automate & Monitor Large Systems
Senior SRE: Scale, Automate & Monitor Large Systems

Otter.ai • Anchorage (AK)

On-site
USD 185,000 - 230,000
Senior Production Engineer
Senior Production Engineer

Cacheflow • Mountain View (CA)

Hybrid
USD 185,000 - 230,000
Senior Production Engineer
Senior Production Engineer

Otter.ai • Anchorage (AK)

On-site
USD 185,000 - 230,000
Staff Backend Engineer, AI Infrastructure
Staff Backend Engineer, AI Infrastructure

otterai • Seattle (WA)

On-site
USD 180,000 - 260,000
Senior Platform Reliability Engineer (Kubernetes & CI/CD)
Senior Platform Reliability Engineer (Kubernetes & CI/CD)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Senior Data Engineer: Build Scalable Data Platforms
Senior Data Engineer: Build Scalable Data Platforms

Otter.ai • Mountain View (CA)

On-site
USD 185,000 - 230,000
Senior Data Engineer - Scale Data Pipelines (Hybrid)
Senior Data Engineer - Scale Data Pipelines (Hybrid)

otterai • Mountain View (CA)

Hybrid
USD 185,000 - 230,000
Staff Backend Engineer: AI Infra & Scalable Systems
Staff Backend Engineer: AI Infra & Scalable Systems

Otter.ai • Seattle (WA)

On-site
USD 210,000 - 275,000
Senior Backend Architect for AI Infrastructure
Senior Backend Architect for AI Infrastructure

Otter.ai • Seattle (WA)

On-site
USD 185,000 - 230,000
Salary top-tier compensation
Senior Platform Engineer - Reliability & AI Observability
Senior Platform Engineer - Reliability & AI Observability

Next Ventures • New York (NY)

On-site
USD 150,000 - 190,000