Cloud Native Engineer

SproutsAI

Palo Alto (CA)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity
Hybrid work model

Job summary

A cutting-edge AI company in Palo Alto is seeking a Cloud Native Engineer to build and optimize the microservices architecture of their AI cloud platform. The role involves designing resilient systems using cloud native technologies while ensuring the platform can handle massive workloads efficiently. Applicants should have over 7 years of software engineering experience, strong proficiency in Go, and hands-on Kubernetes expertise. This is a hybrid role requiring in-office work three days a week.

Qualifications

  • 7+ years of software engineering experience with focus on distributed systems and cloud native technologies.
  • Strong proficiency in Go with deep understanding of concurrency patterns and standard libraries.
  • Hands-on experience with Kubernetes including development, deployment, and maintenance.

Responsibilities

  • Design and develop microservices using Go for AI cloud platform.
  • Build and maintain Kubernetes-based infrastructure for workload management.
  • Implement and optimize cloud native solutions for scalability and reliability.

Skills

Distributed systems expertise
Proficiency in Go
Kubernetes experience
Containerization technologies
Problem-solving mindset
Collaboration skills

Tools

Docker
Prometheus
Grafana
Service mesh technologies

Job description

Zettabyte delivers high-performance AI computing infrastructure to enterprises and entrepreneurs globally. We specialize in offering NVIDIA GPUs—such as H100, A100, and RTX series—through a proprietary software platform called Zsuite, which enables intelligent scheduling, resource optimization, and efficient management of distributed AI workloads. Drawing expertise from hyperscale cloud computing and renowned academic institutions, we address gaps in AI orchestration and reliability, helping customers efficiently train and deploy AI models. The Zsuite platform provides robust orchestration and management capabilities for AI workloads, aiming to democratize access to advanced computing power. Zettabyte collaborates with leading technology companies and is a UN Global Marketplace member, supporting innovation with ethical standards and a global perspective. Our infrastructure supports on-demand cloud GPU instances, performance management, and tailored AI data center deployments, contributing toward advancing the AI ecosystem, notably in Asia. We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

About the Role

We`re looking for a Cloud Native Engineer to build and optimize the microservices architecture powering our AI cloud platform. You`ll design resilient, scalable systems using cutting‑edge cloud native technologies, ensuring our platform can handle massive AI workloads with reliability and performance.

As part of our cloud native engineering team, you`ll work on Kubernetes‑based infrastructure, Go microservices, and container orchestration systems that serve as the backbone of our AI computing platform. You`ll architect the systems that make AI compute seamless for thousands of developers and enterprises.

This is a unique opportunity for someone who`s excited to work with the latest cloud native technologies, solve complex distributed systems challenges, and build infrastructure at scale in the rapidly growing AI space.

What You'll Do
  • Design and develop microservices using Go that power our AI cloud platform`s core functionality
  • Build and maintain Kubernetes‑based infrastructure for container orchestration and workload management
  • Implement and optimize cloud native solutions for scalability, reliability, and performance
  • Contribute to code reviews, technical documentation, and knowledge sharing within the engineering team
  • Explore and integrate emerging cloud native technologies like Volcano, Prometheus, and service mesh solutions
  • Design distributed systems architecture for high‑availability AI workload processing
  • Collaborate with DevOps and SRE teams to ensure production reliability and monitoring
  • Leverage AI‑assisted coding tools (GitHub Copilot, ChatGPT, Cursor IDE, etc.) to boost productivity and code quality
You'll Thrive Here If You
  • 7+ years of software engineering experience with focus on distributed systems and cloud native technologies
  • Strong proficiency in Go with deep understanding of concurrency patterns and standard libraries
  • Familiarity with Python or other backend languages for polyglot development environments
  • Hands‑on Kubernetes experience including development, deployment, and maintenance of production clusters
  • Solid understanding of microservices architecture design patterns and implementation best practices
  • Experience with containerization technologies (Docker, containerd) and container runtime optimization
  • Problem‑solving mindset with ability to independently design and implement complex system components
  • Strong collaboration and communication skills for working in cross‑functional teams
  • Experience using AI‑assisted coding tools and willingness to integrate them into development workflow
  • Familiarity with cloud native ecosystem tools such as Prometheus, Grafana, Volcano, or service mesh technologies
  • Open source contributions to cloud native projects (Kubernetes, CNCF ecosystem)
  • Experience with large‑scale Kubernetes cluster operations and troubleshooting in production environments
  • Knowledge of microservices architecture patterns including circuit breakers, service discovery, and distributed tracing
Compensation

We provide Competitive salary and equity based on your experience and skillset.

This is a Hybrid role - 3 days in office, 2 days WFH; Must locate in Palo Alto

Additional Information

Applicants must be authorized to work in the United States without need for visa sponsorship.

Must locate in Palo Alto and be available for 3 days per week in office per company policy.

Skills
  • Programming Languages & AI‑Assisted Coding
  • Containerization & Orchestration
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior, Staff Backend Engineer - Distributed System
Senior, Staff Backend Engineer - Distributed System

SproutsAI • Palo Alto (CA)

Hybrid
USD 190,000 - 270,000
Competitive salary
Equity based on experience
Senior/Staff Backend Engineer - Distributed System
Senior/Staff Backend Engineer - Distributed System

Zettabyte Inc • Palo Alto (CA)

Hybrid
USD 180,000 - 240,000
Competitive salary
Equity based on experience
Hybrid work model
Virtualization & Orchestration Engineer
Virtualization & Orchestration Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 140,000 - 210,000
Competitive pay
Bonuses & incentives
Benefits: medical/dental/vision/401k
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000
Go Cloud Native Engineer | Kubernetes & AI Platform
Go Cloud Native Engineer | Kubernetes & AI Platform

SproutsAI • Palo Alto (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary and equity
Hybrid work model
GPU Cloud Platform Engineer
GPU Cloud Platform Engineer

Yotta Labs • United States

Remote
USD 120,000 - 160,000
Flexible remote work environment
Innovative team collaboration
Cutting-edge technology challenges
DGX Cloud Automation Engineer
DGX Cloud Automation Engineer

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
DGX Cloud Automation Engineer
DGX Cloud Automation Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Platform Engineer - AI Infrastructure
Platform Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 180,000 - 250,000
Significant freedom and ownership in project development
Work on challenging problems related to ultra-low latency
Join a high-growth environment