ML Systems Engineer

ChipAgents

San Jose (CA)

On-site

USD 150,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Unlimited PTO
Full benefits (medical, vision, dental, 401k)
Free parking and private gym

Job summary

ChipAgents is looking for a skilled ML Systems Engineer in San Jose, California. In this technical role, you will optimize large language model inference for our agentic AI platform, impacting chip design efficiency.

Your responsibilities include implementing performance optimizations, designing multi-node clusters, and collaborating with top semiconductor firms. With a competitive salary range of $150K–$350K, we also offer substantial benefits and equity options.

Qualifications

  • B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field.
  • Experience with large‑scale ML systems, GPU computing, or high‑performance inference optimization.
  • Strong proficiency in Python and C++/CUDA.

Responsibilities

  • Design, deploy, and optimize LLM inference systems across multi‑node clusters.
  • Implement and benchmark inference optimizations.
  • Profile and analyze inference bottlenecks at systems level.

Skills

Large-scale ML systems
GPU computing
High-performance inference optimization
Python
C++/CUDA
SGLang
vLLM
PyTorch
Debugging and profiling

Education

B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field

Tools

Distributed computing frameworks (Ray)

Job description

About ChipAgents

ChipAgents is redefining the future of chip design and verification with agentic AI workflows. Our platform leverages cutting‑edge generative AI to assist engineers in RTL design, simulation, and verification, dramatically accelerating chip development. Founded by experts in AI and semiconductor engineering, we partner with top semiconductor firms, cloud providers, and innovative startups to build intelligent AI agents. The company is a Series A company backed by tier‑1 VC firms. ChipAgents is deployed in production to companies that have shipped 16B chips.

Position Overview

We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform. This is a technical role focused on low‑level systems optimization. You will implement performance optimizations, build evaluation harnesses, and architect multi‑node clusters for training and inference that push the limits of LLM throughput and latency. Your work will directly impact the responsiveness and cost‑efficiency of AI agents used by leading semiconductor companies to design chips.

Key Responsibilities
  • Design, deploy, and optimize LLM inference systems across multi‑node clusters, maximizing throughput and minimizing latency for production workloads.
  • Implement and benchmark concrete inference optimizations.
  • Profile and analyze inference bottlenecks at the systems level—from GPU kernel execution to memory bandwidth constraints.
  • Build robust evaluation harnesses and benchmarking frameworks that measure accuracy, throughput, latency, and resource utilization across various parallelism strategies.
  • Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.
  • Investigate and apply emerging techniques from research papers and open‑source projects to continuously improve inference performance.
Qualifications
  • B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience).
  • Experience with large‑scale ML systems, GPU computing, or high‑performance inference optimization.
  • Strong proficiency in Python and C++/CUDA; hands‑on experience with SGLang, vLLM, PyTorch, or similar inference frameworks.
  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.
  • Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization.
  • Strong systems‑level debugging and profiling skills; comfort working at multiple layers of the stack from CUDA kernels to application logic.
  • Familiarity with distributed computing frameworks (Ray, multi‑node training/inference) is a plus.
  • Self‑directed problem solver who is interested in working on ambitious optimization challenges.
Why Join Us
  • Work on cutting‑edge LLM inference optimization problems with real‑world production impact.
  • Access to substantial GPU compute resources for experimentation and benchmarking.
  • Collaborate with a world‑class team spanning AI research, systems engineering, and EDA.
  • Shape the performance characteristics of AI systems used by leading semiconductor companies.
What we offer
  • $150K/yr – $350K/yr + Offers Equity. We are open to discuss above‑scale compensation with exceptional candidates on a case‑by‑case basis.
  • Unlimited PTO and full benefits (medical, vision, dental, 401k).
  • Two engineering‑centric offices with free parking, private gym, and free lunch, drinks and snacks.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer — Production-Scale LLM Inference
ML Systems Engineer — Production-Scale LLM Inference

ChipAgents • San Jose (CA)

On-site
USD 150,000 - 350,000
Unlimited PTO
Full benefits (medical, vision, dental, 401k)
Free parking and private gym
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

Netpreme • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Relocation assistance
Visa sponsorship
Lunch stipend
+1
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

Netpreme • Cambridge (MA)

On-site
USD 190,000 - 230,000
Relocation assistance
Visa sponsorship
Daily lunch stipend
+2
Research Scientist
Research Scientist

ChipAgents • San Jose (CA)

On-site
USD 150,000 - 350,000
Unlimited PTO
Full benefits (medical, vision, dental, 401k)
Free parking and private gym
+1
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

adaption • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Flexible work
Annual travel stipend
Weekly meal allowance
+1
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

adaption • Redwood City (CA)

Hybrid
USD 120,000 - 160,000
Flexible work
Annual travel stipend
Weekly meal allowance
+1
ML Systems Engineer
ML Systems Engineer

EM DASH LABS • Town of Texas (WI)

On-site
USD 120,000 - 180,000
NVIDIA hardware access
Cloud credits
Ownership of production systems
Full-Stack AI Engineer
Full-Stack AI Engineer

ChipAgents • San Jose (CA)

On-site
USD 150,000 - 350,000
Unlimited PTO
Full medical benefits (medical, vision, dental)
401k
+3
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Research Scientist - AI for Electronic Design Automation
Research Scientist - AI for Electronic Design Automation

ScOp Venture Capital • Santa Clara (CA)

On-site
USD 150,000 - 230,000