ML Systems Engineer

ScOp Venture Capital LLC.

Santa Barbara, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Equity
Professional growth opportunities
Access to GPU compute resources

Job summary

ChipAgents seeks an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering its agentic AI platform. This role centers on low-level systems optimization and building scalable multi-node clusters for training and inference that push LLM throughput and latency limits.

You will implement performance optimizations, build evaluation harnesses, and collaborate with researchers to integrate new model architectures into production infrastructure.

Qualifications

  • 3+ years of experience with large-scale ML systems and GPU inference optimization.
  • Strong Python and C++/CUDA skills; hands-on experience with SGLang, vLLM, PyTorch or similar frameworks.
  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.
  • Experience deploying and optimizing LLMs in production: model serving, batching, distributed inference, or quantization.
  • Strong systems-level debugging and profiling skills; comfortable across CUDA kernels to application logic.
  • Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus.
  • Self-directed problem solver with a passion for ambitious optimization challenges.

Responsibilities

  • Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads.
  • Implement and benchmark concrete inference optimizations.
  • Profile and analyze inference bottlenecks from GPU kernel execution to memory bandwidth constraints.
  • Build robust evaluation harnesses and benchmarking frameworks measuring accuracy, throughput, latency, and resource utilization across parallelism strategies.
  • Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.
  • Investigate and apply emerging techniques from research papers and open-source projects to continuously improve inference performance.

Skills

Python
C++/CUDA
GPU architecture
ML systems
Profiling
Distributed computing
Model serving
SGLang
vLLM
PyTorch
CUDA kernels
Ray

Education

Bachelor/Master/PhD in CS/EE or related

Tools

SGLang
vLLM
PyTorch

Job description

ChipAgents is redefining the future of chip design and verification with agentic AI workflows. Our platform leverages cutting-edge generative AI to assist engineers in RTL design, simulation, and verification, dramatically accelerating chip development. Founded by experts in AI and semiconductor engineering, we partner with top semiconductor firms, cloud providers, and innovative startups to build intelligent AI agents. The company is a Series A company backed by tier-1 VC firms. ChipAgents is deployed in production to companies that have shipped 16B chips.

Position Overview

We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform. This is a technical role focused on low-level systems optimization. You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency. Your work will directly impact the responsiveness and cost-efficiency of AI agents used by leading semiconductor companies to design chips.

Key Responsibilities
  • Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads.
  • Implement and benchmark concrete inference optimizations.
  • Profile and analyze inference bottlenecks at the systems level—from GPU kernel execution to memory bandwidth constraints.
  • Build robust evaluation harnesses and benchmarking frameworks that measure accuracy, throughput, latency, and resource utilization across various parallelism strategies.
  • Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.
  • Investigate and apply emerging techniques from research papers and open-source projects to continuously improve inference performance.
Qualifications
  • B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience).
  • 3+ years of experience with large-scale ML systems, GPU computing, or high-performance inference optimization.
  • Strong proficiency in Python and C++/CUDA; hands‑on experience with SGLang, vLLM, PyTorch, or similar inference frameworks.
  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.
  • Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization.
  • Strong systems‑level debugging and profiling skills; comfort working at multiple layers of the stack from CUDA kernels to application logic.
  • Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus.
  • Self‑directed problem solver who is interested in working on ambitious optimization challenges.
Why Join Us
  • Work on cutting‑edge LLM inference optimization problems with real‑world production impact.
  • Access to substantial GPU compute resources for experimentation and benchmarking.
  • Collaborate with a world‑class team spanning AI research, systems engineering, and EDA.
  • Shape the performance characteristics of AI systems used by leading semiconductor companies.
  • Competitive compensation, benefits, meaningful equity, and professional growth opportunities.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Systems Engineer — LLM Inference & Multi-Node Performance
ML Systems Engineer — LLM Inference & Multi-Node Performance

ScOp Venture Capital LLC. • Santa Barbara (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity
Professional growth opportunities
+1
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs • Irvine (CA)

Hybrid
USD 140,000 - 190,000
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Sr. Staff AI Engineer, Silicon Design
Sr. Staff AI Engineer, Silicon Design

Cognichip • Redwood City (CA)

On-site
USD 150,000 - 200,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Research Scientist - AI for Electronic Design Automation
Research Scientist - AI for Electronic Design Automation

ScOp Venture Capital LLC. • Santa Barbara (CA), Northern (KY)

Hybrid
USD 140,000 - 180,000
Equity options
Competitive compensation
Professional growth opportunities