Software Engineer, ML Infrastructure

Realmlabs

Sunnyvale (CA)

On-site

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Market aligned compensation
Founding engineer equity
Medical, Dental, Vision, and Life insurance
401-K
In-office lunch

Job summary

An innovative AI startup is seeking a Founding ML Infrastructure Engineer to take charge of deploying and optimizing production-grade LLM systems. In this core role, you will be responsible for building and managing a full ML serving stack, working closely with product teams to ensure system performance and reliability. The ideal candidate will have extensive experience in ML infrastructure, particularly with LLMs, and will be proficient in relevant technologies such as PyTorch, TensorFlow, and Kubernetes. This position offers a unique opportunity to shape the future of AI within the company.

Qualifications

  • Extensive experience with GPU inference technologies.
  • Proficient in building production-grade ML infrastructure.
  • Keen awareness of latency and throughput optimization.

Responsibilities

  • Own the end-to-end LLM inference stack.
  • Design high-performance LLM serving systems.
  • Collaborate with product teams for deployment.

Skills

Deep understanding of LLM internals
GPU inference optimization
Software engineering fundamentals
Collaboration with cross-functional teams

Education

5+ years of professional experience in ML infrastructure

Tools

PyTorch
TensorFlow
TensorRT
Triton Inference Server
Kubernetes

Job description

Role Overview

We are hiring a Founding ML Infrastructure Engineer to own the end-to-end deployment, optimization, and operation of our suits of models in production.

This is a core founding role focused on building and operating production-grade LLM systems. You will apply deep knowledge of model internals to deploy, optimize, and run modern LLMs at scale, owning performance end-to-end across latency, throughput, and reliability.

You will design and operate the full ML serving stack from model artifacts to GPU execution, and work closely with Product and ML teams to ensure our models can support high QPS, strict SLAs, and production correctness.

This role is ideal for someone who deeply understands how LLMs work internally but chooses to specialize in making them fast, stable, and production-ready.

About Realm Labs

Realm Labs is an AI trust and security startup. We help enterprises detect, debug, and prevent AI’s misbehaviors in production. We are backed by top VCs and serve some of the most iconic global enterprises.

Key Responsibilities
  • Own the end-to-end LLM inference stack, including:
    • Model loading and execution
    • GPU utilization and memory efficiency
    • Runtime performance tuning
    • Production deployment and scaling
  • Design and operate high-performance LLM serving systems using technologies such as:
    • vLLM, TensorRT / TensorRT-LLM, Triton Inference Server, SGLang
  • Optimize inference across:
    • Latency
    • Throughput (QPS)
    • GPU memory footprint
    • Cost efficiency
  • Work hands-on with PyTorch and TensorFlow models, including:
    • Model graph understanding
    • Attention mechanisms, KV cache behavior, batching strategies
    • Precision tradeoffs (FP16, BF16, INT8, etc.)
  • Build and maintain production-grade GPU services:
  • Multi-model serving
  • Autoscaling strategies
  • Fault isolation and graceful degradation
  • Collaborate with application and platform teams to:
    • Define serving APIs
    • Ensure correctness and safety of outputs
    • Debug production issues end-to-end
  • Build a reproducible model training and versioning system for customer deployments
  • Establish best practices for:
    • Model versioning
    • Rollouts and rollbacks
    • Performance benchmarking
    • Production validation
  • Expected Qualifications
    • 5+ years of professional experience in ML infrastructure, systems engineering, or production ML roles.
    • Strong software engineering fundamentals; ability to write robust, maintainable production code.
    • Deep hands-on experience with LLM inference infrastructure, including:
      • PyTorch (required)
      • TensorFlow (working knowledge)
    • Proven experience with GPU inference optimization, including:
      • TensorRT / TensorRT-LLM
      • vLLM
      • Triton Inference Server
      • SGLang or similar serving runtimes
    • Strong understanding of LLM internals, such as:
      • Transformer architectures
      • Attention and KV caching
      • Batching, streaming, and token-level generation
    • Experience running ML systems in production with high traffic and SLAs.
    • Comfortable working in Linux-based, cloud production environments.
    Preferred Qualifications
    • Experience deploying LLMs on Kubernetes and GPU clusters.
    • Familiarity with CUDA, NCCL, or low-level GPU performance concepts.
    • Experience with:
      • Model sharding and parallelism strategies
      • Multi-GPU inference
      • Streaming inference systems
    • Knowledge of observability for ML systems (metrics, latency breakdowns, GPU monitoring).
    • Experience working at startups or owning systems with minimal abstraction layers.
    Additional Information
    • This is a founding, high-ownership role with direct impact on core product capabilities.
    • You will be expected to build, run, and own systems end-to-end.
    • The role may include limited on-call responsibilities aligned with production ownership.
    Compensation & Benefits
    • Market aligned compensation and benefits.
    • Founding engineer equity (Equity is a significant component of this role and will be discussed).
    • Medical, Dental, Vision, Life insurance, 401-K, In-office lunch etc.
    Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and candidate. But if we make you an offer, we will make all reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
    Get your free, confidential resume review.
    or drag and drop your file here.
    Similar jobs

    Similar jobs worth comparing

    Member of Technical Staff, ML Systems
    Member of Technical Staff, ML Systems

    Netpreme • Santa Clara (CA)

    On-site
    USD 180,000 - 240,000
    Relocation assistance
    Visa sponsorship
    Lunch stipend
    +1
    ML Ops Engineer — Agentic AI Lab (Founding Team)
    ML Ops Engineer — Agentic AI Lab (Founding Team)

    Fabrion • San Francisco (CA)

    On-site
    USD 120,000 - 150,000
    Competitive salary
    Meaningful equity
    Member of Technical Staff - Inference
    Member of Technical Staff - Inference

    Prime Intellect • San Francisco (CA)

    Hybrid
    USD 150,000 - 300,000
    Cash compensation range of $150-300k
    Flexible work arrangement (remote or San Francisco office)
    Full visa sponsorship and relocation support
    +3
    Member of Technical Staff, ML Systems
    Member of Technical Staff, ML Systems

    Netpreme • Cambridge (MA)

    On-site
    USD 190,000 - 230,000
    Relocation assistance
    Visa sponsorship
    Daily lunch stipend
    +2
    Senior Machine Learning Engineer (LLMs)
    Senior Machine Learning Engineer (LLMs)

    Albiware Inc. • Chicago (IL)

    On-site
    Competitive salary
    Generous PTO
    Medical, dental, and vision coverage
    +2
    Member of Technical Staff - Inference
    Member of Technical Staff - Inference

    Prime Intellect • United States

    Hybrid
    USD 120,000 - 150,000
    Competitive compensation
    Flexible work arrangement
    Full visa sponsorship
    +2
    Senior Machine Learning Engineer (LLMs)
    Senior Machine Learning Engineer (LLMs)

    Albi • Chicago (IL)

    On-site
    Competitive salary
    Generous PTO
    Medical, dental, and vision coverage
    +2
    Member of Technical Staff — Model Optimization and Inference (New Grad)
    Member of Technical Staff — Model Optimization and Inference (New Grad)

    Nuance Labs • Seattle (WA)

    On-site
    USD 200,000 - 300,000
    Health Savings Account with $2,000 annual contributions
    15 days of PTO plus public holidays
    Lunch, drinks, and snacks provided daily
    Distributed Systems Engineer, Real-Time Inference at Scale
    Distributed Systems Engineer, Real-Time Inference at Scale

    adaption • San Francisco (CA)

    On-site
    Software Engineer, Inference Runtime
    Software Engineer, Inference Runtime

    LM Studio • New York (NY)

    Hybrid
    USD 150,000 - 350,000
    Competitive salary and equity grants
    Great medical, vision, dental plans
    Catered team lunch / expensed dinners
    +3