Principal Cloud Platform Engineer

SambaNova

Austin (TX)

On-site

USD 144,000 - 189,000

Full time

7 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
HSA with employer contribution
Dental, Vision
Headspace subscription
Gympass+ membership

Job summary

SambaNova Systems is seeking a Principal Cloud Platform Engineer to own the AI inference service reliability, latency, and scalability. You will bridge software and operations, ensuring uptime and efficient resource use, with on-call rotation for 24/7 support.

You will automate infrastructure in new regions, build monitoring, and drive CI/CD for model updates while collaborating with finance on cloud spend. Strong SRE/DevOps and cloud experience are essential.

Qualifications

  • B.S. in Computer Science, Computer Engineering, or related field.
  • 3+ years of Site Reliability Engineering or DevOps experience.
  • Experience supporting a large-scale customer-facing service in a public cloud environment (AWS, GCP, Azure).
  • Strong programming in Python, Go, Rust, or Java.
  • Proven experience with containerization and orchestration (Docker and Kubernetes).
  • Deep understanding of monitoring/observability tools (Prometheus, Grafana, Datadog, ELK).
  • Experience with Infrastructure as Code (Terraform, CloudFormation).
  • Experience with CI/CD (Jenkins, GitHub Actions, ArgoCD).
  • Strong Linux/Unix system administration fundamentals.

Responsibilities

  • Own production inferencing service reliability across regions including availability, latency, and performance.
  • Automate AI infrastructure provisioning in new regions.
  • Participate in on-call rotation and lead incident response.
  • Build monitoring dashboards (Prometheus, Grafana, Datadog) for health, latency, throughput, and accelerator usage.
  • Identify and eliminate performance bottlenecks; design auto-scaling policies.
  • Manage infrastructure as code (Terraform, Ansible) and CI/CD pipelines for model deployments.
  • Forecast infra needs, align with product roadmap, and collaborate with finance on cloud spend.
  • Define and report SLOs/SLIs for the inferencing platform.

Skills

SRE/DevOps experience
Public cloud (AWS/GCP/Azure)
Python/Go/Rust/Java
Docker & Kubernetes
Terraform/CloudFormation
CI/CD (Jenkins/GitHub Actions/ArgoCD)
Linux/Unix administration

Education

B.S. in Computer Science/Computer Engineering or related field

Tools

Prometheus
Grafana
Datadog
ELK Stack
Terraform
CloudFormation
Jenkins
GitHub Actions
ArgoCD
Docker
Kubernetes

Job description

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale.

SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.

About The Team

The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.

About The Role

As a Principal Cloud Platform Engineer, you will be specializing in our AI Inferencing Service and will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability.

Responsibilities
  • Shared ownership of the production inferencing service across regions, covering availability, latency, performance, change management, and capacity planning
  • Standing-up and automating AI infrastructure in new regions
  • Participating in a shared primary/secondary on-call rotation, and leading incident response
  • Building monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog for service health, model latency and throughput, and accelerator utilization
  • Finding and eliminating performance bottlenecks
  • Designing auto-scaling policies that handle variable inference loads
  • Managing cloud and on-prem infrastructure as code in Terraform and Ansible
  • Building CI/CD pipelines that safely deploy new model versions and service updates
  • Forecasting infrastructure needs against the product roadmap and usage trends, and working with finance to manage cloud spend
  • Defining and reporting on SLOs and SLIs for the inferencing platform, using that data to prioritize reliability work
Required Qualifications
  • B.S. in Computer Science, Computer Engineering, or related field
  • 3+ years of experience in a Site Reliability Engineering, DevOps
  • Experience supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)
  • Strong programming and scripting skills in languages like Python, Go, Rust, or Java
  • Proven experience with containerization and orchestration technologies (Docker and Kubernetes)
  • Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog)
  • Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)
  • Experience with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD)
  • Strong Linux/Unix system administration fundamentals
Preferred Qualifications
  • Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure.
  • Direct experience supporting ML/AI inferencing services in production.
  • Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs.
  • Knowledge of model serving frameworks like vLLM, SGLang or Ray.
  • Understanding of MLOps principles and practices.
  • Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached).
Base Salary Range

Base Pay Range $144,000—$189,000 USD

EEO Policy

SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary For US-Based, Full-Time Employment Positions

SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution.

  • We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans.
  • Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care.
  • Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Platform Engineer
Senior Cloud Platform Engineer

Front Door Defense • San Jose (CA)

On-site
USD 140,000 - 210,000
Competitive compensation
Equity and benefits
Cloud Platform Engineer
Cloud Platform Engineer

SambaNova • San Jose (CA)

On-site
USD 150,000 - 210,000
Equity
Flexible work environment
Comprehensive benefits
Senior Software Engineer - ML Infrastructure
Senior Software Engineer - ML Infrastructure

SambaNovaSystems • United States

On-site
USD 200,000 - 275,000
Health insurance
Health Savings Account (HSA)
Headspace subscription
+2
Software Engineering Director - Inference Platform
Software Engineering Director - Inference Platform

SambaNovaSystems • San Jose (CA)

On-site
USD 245,000 - 325,000
Health benefits (premium covered)
Headspace
Gympass+ membership
Director, Software Engineering San Jose, California, United States
Director, Software Engineering San Jose, California, United States

SambaNova • San Jose (CA)

On-site
USD 245,000 - 325,000
95% premium coverage for employee medical insurance
77% premium coverage for dependents
Health Savings Account (HSA)
+2
Cloud Platform Architect
Cloud Platform Architect

SambaNova Systems • San Jose (CA)

On-site
USD 245,000 - 325,000
Health insurance coverage
HSA with employer contribution
Dental & Vision insurance
+1
Software Architect
Software Architect

SambaNova • San Jose (CA)

On-site
USD 245,000 - 325,000
95% premium coverage for employee medical insurance
77% premium coverage for dependents
Health Savings Account (HSA)
+1
Forward Deployment Engineer
Forward Deployment Engineer

External SambaNova Systems • United States

On-site
USD 138,000 - 170,000
Equity
Health insurance
Health Savings Account (HSA)
Forward Deployment Engineer SambaNova Remote - US
Forward Deployment Engineer SambaNova Remote - US

Neura Market • Northern (KY)

Hybrid
USD 138,000 - 170,000
Equity
Health insurance
Well-being benefits
Senior Principal Machine Learning Engineer
Senior Principal Machine Learning Engineer

SambaNovaSystems • San Jose (CA)

On-site
USD 220,000 - 300,000
Equity
Health insurance
HSA
+1