Founding Systems Engineer: Low-Latency LLM & Edge Infra

Berkley Hunt

New York (NY)

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Berkley Hunt is partnering with a well-funded startup building next-gen infrastructure for high-performance, low-latency LLM inference at scale. This founding-level role focuses on low-level systems, distributed environments, and mission-critical performance challenges.

The engineer will own global routing, backend infra, and security controls, shaping edge-based inference and scalable production systems across global locations.

Qualifications

  • Experience building distributed, concurrent systems focused on speed and reliability.
  • Deep expertise in systems programming (Rust, C++, or Zig).
  • Strong Python skills for automation and control-plane development.
  • Experience designing infrastructure for LLM inference or high-scale CDN environments.
  • Proficiency with AWS, Postgres, Redis, and Kafka.
  • Familiarity with Linux internals, performance tuning, and kernel-level debugging.
  • Observability tools such as Zipkin or Jaeger.
  • Clear, concise communicator with a strong documentation and operational mindset.

Responsibilities

  • Architect and implement global routing engines with microsecond decision times.
  • Build and manage high-performance backend infrastructure from scratch, focusing on concurrency, throughput, and resilience.
  • Own the control plane, monitoring systems, and production documentation for all deployments.
  • Lead efforts in internal security and access management systems.
  • Design and optimize infrastructure for edge-based inference with maximum hardware utilization and minimal latency.
  • Collaborate cross-functionally to drive infrastructure decisions from prototype to production.

Skills

Distributed systems
Systems programming
Python scripting
Cloud services
Observability

Tools

Kafka
PostgreSQL
Redis
AWS
Zipkin/Jaeger

Job description

Berkley Hunt is partnering with a well-funded startup building next-gen infrastructure for high-performance, low-latency LLM inference at scale. This founding-level role focuses on low-level systems, distributed environments, and mission-critical performance challenges.

The engineer will own global routing, backend infra, and security controls, shaping edge-based inference and scalable production systems across global locations.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Infrastructure System Engineer
Infrastructure System Engineer

Berkley Hunt • New York (NY)

On-site
USD 150,000 - 210,000
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
LLM Inference Engineer — High-Performance Rust Systems
LLM Inference Engineer — High-Performance Rust Systems

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Head of LLM Inference Platform & Architecture
Head of LLM Inference Platform & Architecture

United States Digital Space LLC • San Francisco (CA)

On-site
USD 240,000 - 360,000
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Distributed LLM Inference Engineer - Scale & Resilience
Distributed LLM Inference Engineer - Scale & Resilience

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior LLM Platform Engineer – Enterprise AI Systems
Senior LLM Platform Engineer – Enterprise AI Systems

Parallel Wireless • United States

On-site
USD 180,000 - 280,000
Staff Research Engineer - LLM Inference & Serving
Staff Research Engineer - LLM Inference & Serving

Modal Labs • New York (NY)

On-site
USD 180,000 - 240,000
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4