Software Engineer - Platform Infrastructure (Rust, C++)

Pantera Capital

Palo Alto (CA)

On-site

USD 180,000 - 440,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

equity
comprehensive medical
vision
dental coverage
access to a 401(k) retirement plan
short & long-term disability insurance
life insurance
various other discounts and perks

Job summary

SpaceXAI seeks engineers to design, build, and optimize a large-scale distributed system powering one of the world’s largest supercomputing clusters. You will dive into the low-level stack to profile and optimize across GPUs, Linux kernel, networking, and filesystems for peak efficiency.

You will collaborate on hardware-software co-design and maintain a scalable, reliable codebase while developing tools to boost team productivity and workflows.

Qualifications

  • Systems programming experience in C, C++, or Rust.
  • Hands-on expertise with Kubernetes (K8s) including cluster architecture, networking, storage and production-grade operations.
  • Strong fundamentals of computer systems and how code runs from hardware to software.

Responsibilities

  • Design, build, and implement a large-scale distributed system powering a major supercomputing cluster.
  • Profile, debug, and optimize performance across GPUs, Linux kernel, networking, and filesystems.
  • Collaborate on hardware, software, and algorithm co-design to advance AI training.
  • Maintain and improve the codebase for scalability and reliability.
  • Develop tools to boost team productivity and streamline workflows.

Skills

Systems programming
Kubernetes
OS fundamentals

Tools

Docker
containerd
crio

Job description

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands‑on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

RESPONSIBILITIES:
  • Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
  • Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
  • Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
  • Maintain and innovate on our codebase to ensure scalability and reliability.
  • Develop tools to enhance team productivity and streamline workflows.
BASIC QUALIFICATIONS:
  • Systems programming experience in C, C++, or Rust
  • Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
  • Hands‑on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production‑grade operations
PREFERRED SKILLS AND EXPERIENCE:
  • Collaborate in a fast‑paced, open environment to design and foundational systems.
  • Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers
  • Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives)
  • Proficiency in performance analysis, profiling, and low‑level optimization techniques
  • Solid understanding of computer networks and the TCP/IP stack
  • Experience working with Linux kernel concepts or systems‑level debugging tools (e.g., perf, gdb, strace, Wireshark)
  • Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows
  • Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel
  • Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar)
COMPENSATION AND BENEFITS:

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long‑term disability insurance, life insurance, and various other discounts and perks.

  • equity
  • comprehensive medical
  • vision
  • dental coverage
  • access to a 401(k) retirement plan
  • short & long‑term disability insurance
  • life insurance
  • various other discounts and perks

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Platform Infrastructure (Rust, C++)
Software Engineer - Platform Infrastructure (Rust, C++)

SpaceXAI • Bellevue (WA)

On-site
USD 180,000 - 440,000
Software Engineer - Platform Infrastructure (Rust, C++)
Software Engineer - Platform Infrastructure (Rust, C++)

Socket.dev • Palo Alto (CA)

On-site
USD 440,000
Equity
Medical coverage
Vision coverage
+5
Software Engineer - Platform Infrastructure (Rust, C++)
Software Engineer - Platform Infrastructure (Rust, C++)

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
401(k) plan
+1
Software Engineer - Platform Core (C++, C)
Software Engineer - Platform Core (C++, C)

Pantera Capital • United States

On-site
USD 180,000 - 440,000
Equity
Medical coverage
401(k)
+3
Software Engineer - Platform Core (C++, C)
Software Engineer - Platform Core (C++, C)

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Software Engineer - Linux Kernel (C++, C)
Software Engineer - Linux Kernel (C++, C)

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical
Vision
+5
Software Engineer - Linux Kernel (C++, C)
Software Engineer - Linux Kernel (C++, C)

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical
Vision
+5
Software Engineer - Platform Core (C++, C)
Software Engineer - Platform Core (C++, C)

SpaceXAI • Seattle (WA)

On-site
USD 180,000 - 440,000
Equity
Medical, vision & dental coverage
401(k) retirement plan
+2
Software Engineer - Platform Security
Software Engineer - Platform Security

SpaceXAI • Palo Alto (CA)

On-site
USD 100,000 - 258,000
Software Engineer - X Data Engineering
Software Engineer - X Data Engineering

Pantera Capital • Palo Alto (CA)

On-site
USD 125,000 - 400,000
Equity
Medical coverage
401(k) retirement plan
+2