Software Engineer - AI Infrastructure

Andromeda Cluster

San Francisco (CA)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Ownership and autonomy in projects
Engagement with customers and providers
Opportunity to shape systems

Job summary

A cutting-edge AI infrastructure firm is seeking a Software Engineer specializing in Infrastructure to design core platform components. You will develop APIs and services, translate customer needs into actionable requirements, and enhance system performance. The ideal candidate has over 5 years of relevant experience, strong expertise in Kubernetes and automation tools, and a solid foundation in backend engineering. This role offers the opportunity to shape the future of AI infrastructure in a rapidly growing sector.

Qualifications

  • 5+ years of experience in Infrastructure, Platform, or Backend Engineering roles.
  • Strong systems fundamentals with deep understanding of Linux, networking, storage, and distributed systems.
  • Proven expertise with Kubernetes, VMs, or bare-metal environments.

Responsibilities

  • Design and develop core platform components, including infrastructure orchestration.
  • Build robust APIs and services that abstract over diverse infrastructure types.
  • Translate customer usage patterns into product requirements.

Skills

Infrastructure Engineering
Backend Engineering
Systems Fundamentals
Kubernetes expertise
Linux
Networking
Storage
Distributed Systems
APIs development
Automation tools

Tools

Terraform
Ansible
Helm

Job description

Software Engineer - AI Infrastructure

Location: North America Remote / San Francisco · Full-Time

About Andromeda

Andromeda Cluster, founded by Nat Friedman and Daniel Gross, is on a mission to democratize access to cutting-edge AI infrastructure previously reserved for hyperscalers. What began with a single managed cluster has quickly evolved into a global platform, connecting leading AI labs, data centers, and cloud providers. Our orchestration layer seamlessly routes training and inference jobs across the world, unlocking flexibility and efficiency in one of the fastest-growing sectors on earth. Our long‑term vision is to establish a global marketplace for AI compute—powering AGI with the same fluidity as world financial markets.

We are scaling rapidly and seeking exceptional talent in AI infrastructure, research, and engineering.

The Role

As an Infrastructure Product Engineer, you will play a pivotal role in building the backbone of Andromeda’s platform. You'll transform complex, real‑world infrastructure challenges into scalable product capabilities that benefit our customers.

Positioned at the intersection of infrastructure and product engineering, this role is deeply technical and systems‑oriented, yet laser‑focused on building solutions with broad leverage.

What You’ll Do
  • Design and develop core platform components, including infrastructure orchestration, provisioning, and lifecycle management solutions.
  • Build robust APIs, services, and control planes that abstract over diverse infrastructure types (VMs, Kubernetes, bare metal, schedulers).
  • Translate customer usage patterns into product requirements, delivering impactful features and improvements.
  • Create automation and internal tooling to eliminate manual or ad‑hoc operational work.
  • Enhance reliability, performance, and observability at the platform level, emphasizing durable improvements over quick fixes.
  • Collaborate with peer teams to define clear ownership boundaries between platform capabilities and customer‑specific solutions.
  • Write clean, maintainable, and well‑documented code with a focus on long‑term sustainability.
  • Participate in technical design discussions and contribute to the architectural evolution of our platform.
What We’re Looking For
  • 5+ years of experience in Infrastructure, Platform, or Backend Engineering roles.
  • Strong systems fundamentals: deep understanding of Linux, networking, storage, and distributed systems.
  • Proven expertise with Kubernetes, VMs, or bare‑metal environments.
  • Advanced software engineering skills; capable of building production‑grade APIs and services (Python, Go, or similar).
  • Extensive experience with infrastructure as code and automation tools (Terraform, Ansible, Helm, etc.).
  • Demonstrated ability to navigate ambiguity and distill complex problems into clear, maintainable abstractions.
  • Product‑focused mindset: care about interfaces, defaults, reliability, and sustainable operations.
  • Excellent written and verbal communication skills; effective collaborator across engineering and product functions.

Nice to Have:

  • Hands‑on experience with GPU or AI infrastructure.
  • Experience with control‑plane or orchestration systems.
  • Background spanning both infrastructure and application/backend engineering.
  • Experience architecting multi‑tenant systems.
  • Strong skills in technical writing and design documentation.
  • Early‑stage startup experience.
Why You’ll Love It Here

This is a true builder’s opportunity: you’ll have ownership and autonomy to shape our systems, engage directly with customers and providers, and lay the foundations for scalable, reliable AI infrastructure. Join us at Andromeda and help power the future of AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Andromeda • San Francisco (CA)

Remote
USD 120,000 - 160,000
Customer Reliability Engineer
Customer Reliability Engineer

Andromeda Cluster • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda • San Francisco (CA)

On-site
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Infrastructure Manager
Infrastructure Manager

The Resume Database • San Francisco (CA)

On-site
USD 100,000 - 130,000
Competitive compensation
Meaningful equity
Comprehensive benefits including healthcare
+1
Member of the Business Staff - Compute Markets
Member of the Business Staff - Compute Markets

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 90,000 - 120,000
Competitive compensation
Meaningful equity
Comprehensive healthcare benefits
+1
Member of the Technical Staff - Systems
Member of the Technical Staff - Systems

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Health insurance
Equity
Unlimited PTO
+1
Solutions Architect
Solutions Architect

Andromeda • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Equity
Healthcare, dental, and vision
401(k)
+1
Compute Trader
Compute Trader

Andromeda • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive compensation
Meaningful equity
Comprehensive healthcare benefits
+1