Founding Machine Learning Infrastructure Engineer

Peano AI

Palo Alto (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Peano AI is building the Agent Cloud infrastructure and seeks a Founding ML Infrastructure Engineer to design and optimize high‑performance model serving for large‑scale open‑source models in Palo Alto.

You will own end‑to‑end ML systems, work with accelerators like TPUs/GPUs, improve latency and throughput, and collaborate with research and product teams to ship production‑grade infrastructure.

Qualifications

  • Strong experience in ML systems, distributed systems, or high-performance computing.
  • Experience optimizing inference or training workloads for large models.
  • Familiarity with TPUs, GPUs, or other accelerators.
  • Experience with CUDA, Triton, NCCL, JAX/XLA, PyTorch internals, vLLM, SGLang, or TensorRT‑LLM.
  • Strong systems debugging skills.
  • Comfort working across model code, runtime, infrastructure, and product requirements.
  • High ownership in an early-stage startup environment.

Responsibilities

  • Optimize large-scale LLM inference and serving systems.
  • Improve tokens per second, latency, throughput, and cost efficiency.
  • Work on serving infrastructure for open-source models across accelerators.
  • Improve batching, scheduling, KV cache management, memory usage, and accelerator utilization.
  • Support long-context inference up to 1M context.
  • Debug performance bottlenecks across model execution, runtime, networking, and infrastructure.
  • Work with frameworks such as JAX/XLA, PyTorch, vLLM, SGLang, TensorRT-LLM, or related systems.
  • Collaborate with the application team to optimize infrastructure for agentic workloads.
  • Help turn research prototypes into reliable production systems.

Skills

ML systems
Distributed systems
High-performance computing
Inference optimization
Training optimization
TPUs
GPUs

Tools

CUDA
Triton
NCCL
JAX/XLA
PyTorch internals
vLLM
SGLang
TensorRT-LLM

Job description

Founding Machine Learning Infrastructure Engineer

Location: Onsite in Palo Alto

Compensation: Competitive Salary + Equity

About Peano AI

Peano AI is building the infrastructure and application stack for the next generation of agentic AI systems .

We believe token usage will grow exponentially over the coming years, but routing all inference through closed model providers will remain too expensive for many users and enterprises. Our thesis is that agentic applications require a vertically integrated stack: high-throughput, cost-efficient serving infrastructure paired with an application layer designed for long-running, agentic workloads.

Peano AI is building the Agent Cloud, a serving and training infrastructure platform purpose-built for agentic workloads, long-context inference, and large-scale open-source model deployment. By combining infrastructure and application design, we aim to make open-source models significantly more performant, practical, and competitive.

About This Role

We are looking for an ML Systems Engineer to help build and optimize the core serving infrastructure behind Agent Cloud. This role focuses on high-performance inference across different accelerators.

You will work on model serving performance, accelerator utilization, long-context inference, batching, scheduling, KV cache management, runtime efficiency, and cost reduction. This is a deeply technical role at the intersection of ML systems, infrastructure, and product.

Direct TPU experience is a strong plus, but not required. We care most about strong ML systems fundamentals, performance intuition, and the ability to ship reliable systems quickly.

What You'll Do

  • Optimize large-scale LLM inference and serving systems.

  • Improve total tokens per second, decode tokens per second, latency, throughput, and cost efficiency.

  • Work on serving infrastructure for open-source models across different types of accelerators.

  • Improve batching, scheduling, KV cache management, memory usage, and accelerator utilization.

  • Support long-context inference, including workloads targeting up to 1M context.

  • Debug performance bottlenecks across model execution, runtime, networking, and infrastructure.

  • Work with frameworks such as JAX/XLA, PyTorch, vLLM, SGLang, TensorRT-LLM, or related systems.

  • Collaborate closely with the application team to ensure infrastructure is optimized for agentic workloads, not just generic chatbot inference.

  • Help turn research prototypes into reliable, high-performance production systems.

Qualifications

  • Strong experience in ML systems, distributed systems, or high-performance computing.

  • Experience optimizing inference or training workloads for large models.

  • Familiarity with TPUs, GPUs, or other accelerators.

  • Experience with one or more of CUDA, Triton, NCCL, JAX/XLA, PyTorch internals, vLLM, SGLang, TensorRT-LLM, distributed inference, or distributed training.

  • Strong systems debugging skills.

  • Comfort working across model code, runtime, infrastructure, and product requirements.

  • High ownership and the ability to operate effectively in an early-stage startup environment.

Cultural Fit

  • Hands-on technical excellence and strong engineering judgment.

  • End-to-end ownership, from design to implementation to production outcomes.

  • Bias for action: ship quickly, learn from failures, and iterate.

  • High intensity during critical milestones, with a focus on real customer impact.

  • Ability to do deep, focused work and sustain execution.

  • Clear communication with teammates, customers, and stakeholders.

  • Comfort with ambiguity, rapid change, and wearing multiple hats.

  • Low ego, high integrity, high accountability, and strong collaboration.

  • Continuous learning and a belief that judgment, intelligence, and capability compound over time.

If you are excited to build the infrastructure and agent systems behind the next generation of AI applications, push open-source models to production-grade performance, and turn ambitious research ideas into real-world impact, Peano AI is the place for you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive compensation & equity
Health benefits
Cutting-edge AI work
+1
Founding Full-Stack Engineer: AI/ML Infra Leader
Founding Full-Stack Engineer: AI/ML Infra Leader

Octavia Technologies • San Francisco (CA)

On-site
USD 160,000 - 210,000
Staff / Principal Machine Learning Engineer, Serving - USA
Staff / Principal Machine Learning Engineer, Serving - USA

Inworld • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits
Senior Software Engineer, ML Infrastructure
Senior Software Engineer, ML Infrastructure

Arena Intelligence • Bay (AR)

On-site
USD 120,000 - 180,000
Equity
Health benefits
Work on cutting-edge AI
Founding Engineer - Applied AI
Founding Engineer - Applied AI

Product Pulse • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Full health insurance
Free lunch and dinner
+2
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits package
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure

Apple Inc. • San Francisco (CA)

On-site
USD 210,000 - 300,000
Founding Forward Deployed Machine Learning Engineer
Founding Forward Deployed Machine Learning Engineer

adaption • San Francisco (CA)

On-site
USD 100,000 - 140,000
Flexible work
Annual travel stipend
Weekly meal allowance
+2
Staff Engineer, Machine Learning
Staff Engineer, Machine Learning

ActAI • Palo Alto (CA)

On-site
USD 190,000 - 270,000