Senior AI Inference Platform Engineer

Lyceum

Zürich

Vor Ort

CHF 120.000 - 200.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Lyceum is seeking a Member of Technical Staff to own and operate its AI inference platform. You will architect a reliable, secure, and scalable serving stack that can handle thousands of concurrent users while evolving the platform as we grow.

Your focus includes scalable routing, autoscaling across GPU clusters, robust observability, and performance optimization from ingestion to model execution. You will join a world-class team and help maintain security, multi-tenancy, and high availability.

Qualifikationen

  • 3+ years of backend, infrastructure, or systems engineering experience.
  • Strong proficiency in Go and Python.
  • Experience building or operating a model serving platform or ML platform.
  • Solid understanding of systems performance - profiling, benchmarking, and optimising latency and throughput.
  • Familiarity with observability tooling (Prometheus, Grafana, OpenTelemetry, or similar).
  • Understanding of security fundamentals - network isolation, authentication, encryption, multi-tenancy.

Aufgaben

  • Design and implement scalable routing, load balancing, and autoscaling across GPU clusters.
  • Build robust monitoring, alerting, and incident response tooling to reduce downtime.
  • Profile and optimise the full inference path from request to response to reduce latency and improve throughput.
  • Evaluate open-source inference frameworks and tooling to improve throughput and stability.
  • Ensure security and multi-tenancy considerations are integrated into the serving stack.

Kenntnisse

Go
Python
Model serving platform
Performance profiling
Observability tooling
Security fundamentals

Tools

Prometheus
Grafana
OpenTelemetry
Kubernetes
NVIDIA Dynamo
vLLM
Triton

Jobbeschreibung

Lyceum is seeking a Member of Technical Staff to own and operate its AI inference platform. You will architect a reliable, secure, and scalable serving stack that can handle thousands of concurrent users while evolving the platform as we grow.

Your focus includes scalable routing, autoscaling across GPU clusters, robust observability, and performance optimization from ingestion to model execution. You will join a world-class team and help maintain security, multi-tenancy, and high availability.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Staff Engineer, Scalable AI Inference Platform
Staff Engineer, Scalable AI Inference Platform

Lyceum • Zürich

Vor Ort
CHF 130.000 - 170.000
Member of Technical Staff – AI Inference platform, robustness
Member of Technical Staff – AI Inference platform, robustness

Lyceum • Zürich

Vor Ort
CHF 120.000 - 200.000
Member of Technical Staff – AI Inference platform
Member of Technical Staff – AI Inference platform

Lyceum • Zürich

Vor Ort
CHF 130.000 - 170.000
Senior AI Inference Systems Engineer - GPU HPC
Senior AI Inference Systems Engineer - GPU HPC

NVIDIA • Schweiz

Vor Ort
CHF 140.000 - 210.000
Senior AI Inference & HPC Systems Engineer
Senior AI Inference & HPC Systems Engineer

NVIDIA • Schweiz

Vor Ort
CHF 120.000 - 180.000
Low-Latency AI Backend Engineer
Low-Latency AI Backend Engineer

AI Chopping Block • Zürich

Vor Ort
CHF 120.000 - 170.000
GPU Platform Engineer - Build Sovereign AI Compute
GPU Platform Engineer - Build Sovereign AI Compute

Lyceum • Zürich

Vor Ort
CHF 120.000 - 180.000
GPU Platform Engineer
GPU Platform Engineer

Lyceum • Zürich

Vor Ort
CHF 120.000 - 180.000
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Schweiz

Vor Ort
CHF 100.000 - 140.000
Senior AI-Focused Full-Stack Engineer (LLM/Agents)
Senior AI-Focused Full-Stack Engineer (LLM/Agents)

Embodied AI • Schweiz

Hybrid
CHF 94.000 - 169.000
Stock options
Hybrid/Remote CET±4
Mental health app access
+1