Staff+ Production Engineer

Sanas.AI Inc.

Palo Alto, Northern (CA, KY)

Hybrid

USD 140,000 - 170,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Sanas is seeking a Production Engineer to own the infra that builds, deploys, and operates our real-time speech AI platform globally. You will architect deployment infrastructure across public clouds and drive the production lifecycle with strong observability, on-call discipline, and robust tooling.

You will design scalable, fault-tolerant systems, champion operational excellence, and collaborate across teams to push forward the next generation of speech + AI capabilities.

Qualifications

  • 5+ years of hands-on Production/Platform engineering experience on AWS.
  • Expertise in infrastructure-as-code (Terraform) and containerization (Docker).
  • Advanced observability experience with Datadog, NewRelic, Prometheus, and Grafana.
  • Strong Python or Rust for internal tooling and automation.
  • Excellent cross-functional communication and ability to document complex distributed systems clearly.
  • Proven track record in cost optimization and resource management in large cloud environments.

Responsibilities

  • Design and implement mission-critical, multi-region deployment infrastructure on AWS.
  • Lead architecture for platform deployment into corporate private clouds (Azure/GCP as applicable).
  • Improve developer velocity with rapid, high-confidence deployment workflows (CI/CD).
  • Build and maintain a cohesive telemetry, monitoring, and alerting ecosystem for high availability.
  • Standardize on-call docs, incident response, and post-mortems.
  • Create self-service tooling and comprehensive documentation for engineering teams.

Skills

AWS
Terraform
Docker
Observability tooling
Python or Rust
Cross-functional communication
Cost optimization
Multi-region deployment

Tools

Terraform
Docker

Job description

Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real-time speech AI platform capable of accent translation, noise cancellation, speech enhancement, cross-language communication, and more.

Sanas makes conversations clearer, more inclusive, and more effective, removing barriers that prevent people from being understood, regardless of accent, background noise, or native language.

Sanas is currently one of the fastest growing startups in Silicon Valley, growing from $16M to $50M ARR in 2025. The company's core business is profitable and is on track to end 2026 with >$77M ARR. Our team combines deep expertise in model innovation and systems engineering with a design-minded product engineering culture to build and ship cutting-edge AI models and experiences — entirely in-house.

Sanas is a 130 person team, established in 2020. In this short span, we've successfully secured over $100 million in funding. Our innovation has been supported by the industry's leading investors, including Insight Partners, Google Ventures, Quadrille Capital, General Catalyst, Quiet Capital, and other influential investors. Our reputation is further solidified by collaborations with numerous Fortune 100 companies. With Sanas, you're not just adopting a product; you're investing in the future of communication.

If you're looking to have a significant role in roadmapping and driving technical directions, if you're looking to deploy challenging and big ideas without much overhead or slowness, if you're looking to leave your mark on an ambitious, generational mission to change how the worlds thinks about speech + AI, then Sanas is a well-suited place for you.

About the Role

We are seeking a Production Engineer to take complete, end-to-end ownership of the infrastructure that builds, deploys, and operates our high-scale, real-time speech AI platform globally. This role involves architecting and delivering deployment infrastructure across public and corporate private clouds (AWS, Azure, and GCP). You will champion the entire production lifecycle, ensuring operational excellence across internal development workflows, deployment pipelines, world-class observability (telemetry, monitoring, alerting), and robust on-call systems for high availability and globally-distributed production environments.

Job Description
  • Design and implement mission-critical, multi-region deployment infrastructure on AWS to ensure exceptional scalability and fault tolerance.
  • Lead the architecture for deploying the platform into corporate private cloud environments across diverse architectures (AWS, Azure, GCP).
  • Drive developer velocity by engineering and maintaining rapid, high-confidence development, testing, and staging deployment workflows.
  • Establish and maintain a cohesive, state-of-the-art telemetry, monitoring, and alerting ecosystem to achieve deep operational observability and industry-leading system uptime.
  • Standardize and orchestrate world-class on-call documentation, incident response processes, and post-mortem procedures.
  • Proactively build robust self-service tools, systems, and comprehensive documentation that empowers engineering teams to manage and scale their own services.
Qualifications
  • 5+ years of deep, hands-on expertise as a Platform/Production engineer building, scaling, and operating critical, high-availability production environments on AWS.
  • Proven mastery of infrastructure-as-code and containerization technologies, specifically Terraform and Docker.
  • Expert-level experience designing and implementing advanced cloud monitoring and observability systems (e.g., Datadog, NewRelic, Prometheus/Grafana).
  • Advanced capability in building and maintaining internal tooling and automation using Python or Rust.
  • Possess strong "engineering taste" and the ability to define, champion, and enforce a high-quality bar for operational excellence and site reliability across the entire organization.
  • Excellent cross-functional communication and consensus-building skills, with a focus on clearly documenting and describing complex distributed systems to technical and non-technical audiences.
  • Demonstrated track record in cost optimization and resource management in a high-scale cloud environment, balancing efficiency with performance needs.
Bonus
  • Significant deployment and operational experience across multiple major cloud environments (Azure or GCP).
  • Deep experience with GPU performance tuning, resource orchestration, and scaling for AI/ML workloads.
  • Familiarity with the unique challenges of real-time, low-latency systems, particularly conversational speech or streaming services, aligning with Sanas's core voice pipelines.
  • Experience designing and operating robust self-serve developer platforms, including features like usage-based billing, quota management, or tiered access controls.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff+ Production Engineer
Staff+ Production Engineer

Tensec • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Member of Technical Staff, ML Inference Engineering
Member of Technical Staff, ML Inference Engineering

Sanas • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior Production Engineer — Real-Time Speech AI Platform
Senior Production Engineer — Real-Time Speech AI Platform

Sanas.AI Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 140,000 - 170,000
Member of Technical Staff, LLM Post-Training, Applied
Member of Technical Staff, LLM Post-Training, Applied

Sanas • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Field Marketing Specialist
Field Marketing Specialist

Sanas • Palo Alto (CA)

On-site
USD 100,000 - 150,000
Strategic Account Growth Manager
Strategic Account Growth Manager

Sanas • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Growth Account Executive
Growth Account Executive

Sanas.AI Inc. • Palo Alto (CA), Northern (KY)

On-site
USD 120,000 - 180,000
Enterprise Account Executive
Enterprise Account Executive

Sanas • Boston (MA)

On-site
USD 120,000 - 210,000
Senior ML Inference Engineer — Real-Time Systems
Senior ML Inference Engineer — Real-Time Systems

Sanas • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior Member of Technical Staff
Senior Member of Technical Staff

Thomas Talent Network • San Francisco (CA)

On-site
USD 225,000 - 300,000