Member of Technical Staff – AI Inference platform, robustness

Lyceum

Zürich

Vor Ort

CHF 120.000 - 200.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Lyceum is seeking a Member of Technical Staff to own and operate its AI inference platform. You will architect a reliable, secure, and scalable serving stack that can handle thousands of concurrent users while evolving the platform as we grow.

Your focus includes scalable routing, autoscaling across GPU clusters, robust observability, and performance optimization from ingestion to model execution. You will join a world-class team and help maintain security, multi-tenancy, and high availability.

Qualifikationen

  • 3+ years of backend, infrastructure, or systems engineering experience.
  • Strong proficiency in Go and Python.
  • Experience building or operating a model serving platform or ML platform.
  • Solid understanding of systems performance - profiling, benchmarking, and optimising latency and throughput.
  • Familiarity with observability tooling (Prometheus, Grafana, OpenTelemetry, or similar).
  • Understanding of security fundamentals - network isolation, authentication, encryption, multi-tenancy.

Aufgaben

  • Design and implement scalable routing, load balancing, and autoscaling across GPU clusters.
  • Build robust monitoring, alerting, and incident response tooling to reduce downtime.
  • Profile and optimise the full inference path from request to response to reduce latency and improve throughput.
  • Evaluate open-source inference frameworks and tooling to improve throughput and stability.
  • Ensure security and multi-tenancy considerations are integrated into the serving stack.

Kenntnisse

Go
Python
Model serving platform
Performance profiling
Observability tooling
Security fundamentals

Tools

Prometheus
Grafana
OpenTelemetry
Kubernetes
NVIDIA Dynamo
vLLM
Triton

Jobbeschreibung

Member of Technical Staff – AI Inference platform

Full-time

Your mission

You will make Lyceum's AI inference platform reliable, secure, and scalable - ensuring it performs under pressure as we grow to thousands of concurrent users. While others on the team expand what the platform can do, your job is to make sure it keeps working, fails gracefully, and gets faster over time.

Your focus
  • Scalability: Architect and implement the systems that allow our inference platform to scale to thousands of concurrent users. This includes request routing, load balancing, autoscaling, and resource scheduling across GPU clusters.
  • Reliability and observability: Build robust monitoring, alerting, and incident response tooling. Design for graceful degradation, automatic recovery, and minimal downtime.
  • Performance engineering: Profile and optimise the full inference path from request ingestion through model execution to response delivery. Identify and eliminate bottlenecks at every layer.
  • Infrastructure evolution: Evaluate and integrate open-source inference frameworks and tooling (Dynamo, vLLM, Triton, etc.) where they improve throughput, latency, or stability of the serving stack.
Your KPIs
  • Platform uptime and availability (SLA adherence)
  • P50/P95/P99 latency and throughput under load
  • Time-to-detection and time-to-resolution for incidents
  • Scalability milestones (concurrent users, requests per second, GPU utilisation)
Your profile

We consider candidates from diverse backgrounds, with a deep love for technical challenges and the desire to take on ownership beyond what's reasonably expected.

Requirements
  • 3+ years of experience in backend, infrastructure, or systems engineering
  • Strong proficiency in Go and Python
  • Experience building or operating a model serving platform or ML platform
  • Solid understanding of systems performance - profiling, benchmarking, and optimising latency and throughput
  • Familiarity with observability tooling (Prometheus, Grafana, OpenTelemetry, or similar)
  • Understanding of security fundamentals - network isolation, authentication, encryption, multi-tenancy
Nice to have
  • Experience with NVIDIA Dynamo or similar inference orchestration/routing frameworks
  • Hands‑on experience with GPU serving infrastructure (vLLM, Triton, TensorRT-LLM)
  • Experience with Kubernetes in a production environment (deployment, networking, resource management)
  • Experience operating at scale (10k+ RPS, multi-region, multi-cluster)
  • Comfortable working on‑call or in incident response when things break
Why us?
  • Outstanding team: Work with some of the best engineers in the world, coming from hedge funds, big tech, AI startups and top universities.
  • Once in a lifetime opportunity: Early‑stage company in the fastest-growing market in the world
  • Ownership: Shape how European AI companies access GPU compute
  • European mission: Build sovereign, GDPR‑compliant AI infrastructure for the next generation of deep‑tech
About us

Lyceum is building AI-native GPU infrastructure for the next generation of deep‑tech companies. We’re an early‑stage team moving fast, obsessed with customer value, and focused on building a category-defining product.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Member of Technical Staff – AI Inference platform
Member of Technical Staff – AI Inference platform

Lyceum • Zürich

Vor Ort
CHF 130.000 - 170.000
Member of Technical Staff – AI Inference platform, features
Member of Technical Staff – AI Inference platform, features

Lyceum • Zürich

Vor Ort
CHF 120.000 - 180.000
GPU Platform Engineer
GPU Platform Engineer

Lyceum • Zürich

Vor Ort
CHF 120.000 - 180.000
Junior Member of Technical Staff – Platform Infrastructure, GPUs
Junior Member of Technical Staff – Platform Infrastructure, GPUs

Lyceum • Zürich

Vor Ort
CHF 90.000 - 120.000
Head of Engineering
Head of Engineering

Lyceum • Zürich

Vor Ort
CHF 180.000 - 260.000
Staff Engineer, Scalable AI Inference Platform
Staff Engineer, Scalable AI Inference Platform

Lyceum • Zürich

Vor Ort
CHF 130.000 - 170.000
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Lyceum • Zürich

Vor Ort
CHF 120.000 - 200.000
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Schweiz

Vor Ort
CHF 100.000 - 140.000
Staff / Principal Machine Learning Engineer, Serving - Switzerland
Staff / Principal Machine Learning Engineer, Serving - Switzerland

Inworld • Schweiz

Remote
CHF 90.000 - 120.000
Senior Full Stack Developer
Senior Full Stack Developer

Delta Labs AG • Zürich

Vor Ort
CHF 110.000 - 165.000
Competitive salary
Equity opportunity
Early-stage startup environment