Senior/Staff Backend Engineer - Distributed System

Zettabyte Inc

Palo Alto (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Equity based on experience
Hybrid work model

Job summary

A tech company focusing on AI infrastructure is seeking a Senior/Staff Backend Engineer to create innovative systems for managing GPU clusters. Ideal candidates will possess extensive experience in backend engineering with distributed systems, proficiency in languages like Go or Python, and a passion for solving complex AI challenges. This hybrid role requires in-office presence in Palo Alto three days a week, offering a competitive salary and equity based on experience.

Qualifications

  • 5+ years backend engineering experience with distributed systems.
  • Strong proficiency in Go, Python, or similar backend languages.
  • Experience with resource scheduling and orchestration.
  • Understanding of hardware constraints and system optimization.
  • Linux systems knowledge and containerization experience (Docker, Kubernetes).
  • Comfortable working with expensive resources where efficiency directly impacts costs.
  • Excited about solving novel problems in AI infrastructure (not just another CRUD app).
  • Startup mindset—comfortable with ambiguity and rapid iteration.

Responsibilities

  • Design APIs for GPU operations.
  • Build scheduling algorithms for GPU utilization.
  • Develop resource management systems for GPU lifecycle.
  • Create usage tracking and billing systems for GPU-hours, memory usage, and compute utilization.
  • Implement monitoring for GPU-specific metrics, health checks, and automatic failure recovery.
  • Build multi-tenancy systems with resource isolation, quota management, and fair scheduling.
  • Optimize cold starts for model serving and implement efficient model loading strategies.
  • Collaborate with frontend engineers to expose complex infrastructure through intuitive interfaces.
  • Leverage AI-assisted coding tools to boost productivity and code quality.

Skills

Backend engineering experience with distributed systems
Proficiency in Go, Python, or similar
Resource scheduling and orchestration
API design (REST, GraphQL, gRPC)
Linux systems knowledge
Containerization (Docker, Kubernetes)
Problem-solving in AI infrastructure
Comfort with high-cost efficiency
Startup mindset
Kubernetes
Multi-tenancy
Scheduling
Billing/meters

Tools

Docker
Kubernetes
GitHub Copilot
Claude Code

Job description

Senior/Staff Backend Engineer - Distributed System

Zettabyte Inc

About Us

At Zettabyte, we’re on a mission to make AI compute ubiquitous, seamless, and limitless. We’re building a cloud where AI just works—anywhere, anytime. “AI Power. Everywhere.” Be part of the team designing the infrastructure for the AI-first world.

Why this role exists

We need a Backend Engineer to build the systems that orchestrate GPU clusters for AI workloads. You'll create APIs that handle GPU allocation, memory management, compute scheduling, and multi-tenant isolation—challenges unique to AI infrastructure beyond typical backend engineering. You will solve questions: How do we efficiently share expensive GPU resources across users? How do we handle GPU memory constraints for large AI models? How do we ensure quality of service when workloads compete for compute? This is an opportunity to build infrastructure where every API call could allocate thousands of dollars worth of compute per hour, where your optimizations directly impact whether AI startups can afford to train their models.

What you’ll do
  • Design APIs that abstract complex GPU operations into simple developer experiences
  • Build scheduling algorithms that maximize GPU utilization while ensuring SLA compliance
  • Develop resource management systems for GPU lifecycle—provisioning, allocation, scheduling, and release
  • Create usage tracking and billing systems for GPU-hours, memory usage, and compute utilization
  • Implement monitoring for GPU-specific metrics, health checks, and automatic failure recovery
  • Build multi-tenancy systems with resource isolation, quota management, and fair scheduling
  • Optimize cold starts for model serving and implement efficient model loading strategies
  • Collaborate with frontend engineers to expose complex infrastructure through intuitive interfaces
  • Leverage AI-assisted coding tools (GitHub Copilot, Claude Code, Cursor IDE, etc.) to boost productivity and code quality.
Qualifications
  • 5+ years backend engineering experience with distributed systems
  • Strong proficiency in Go, Python, or similar backend languages
  • Experience with resource scheduling, orchestration, and API design (REST, GraphQL, gRPC)
  • Understanding of hardware constraints and system optimization
  • Linux systems knowledge and containerization experience (Docker, Kubernetes)
  • Comfortable working with expensive resources where efficiency directly impacts costs
  • Excited about solving novel problems in AI infrastructure (not just another CRUD app)
  • Startup mindset—comfortable with ambiguity and rapid iteration
Bonus Qualifications
  • GPU or HPC cluster management experience
  • Understanding of ML/AI workload patterns and requirements
  • Experience with high-value resource allocation systems
  • Background in performance optimization for compute-intensive workloads
  • Familiarity with GPU virtualization and sharing technologies
  • Experience building billing or metering systems
Details
  • Competitive salary and equity based on your experience and skillset
  • Hybrid role: 3 days in office, 2 days WFH; must locate in Palo Alto
  • Authorized to work in the United States without visa sponsorship

Referrals increase your chances of interviewing at Zettabyte Inc by 2x.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior, Staff Backend Engineer - Distributed System
Senior, Staff Backend Engineer - Distributed System

SproutsAI • Palo Alto (CA)

Hybrid
USD 190,000 - 270,000
Competitive salary
Equity based on experience
Cloud Native Engineer
Cloud Native Engineer

SproutsAI • Palo Alto (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary and equity
Hybrid work model
Backend Software Engineer (Distributed Systems / Python)
Backend Software Engineer (Distributed Systems / Python)

Glint Tech Solutions • San Francisco (CA)

Hybrid
USD 170,000 - 230,000
Competitive equity package
Comprehensive benefits
Senior GPU Infrastructure Engineer
Senior GPU Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Senior GPU Infrastructure Engineer - AI Infrastructure
Senior GPU Infrastructure Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours