Senior Scientist, Efficient LLM Inference & Optimization

United States Digital Space LLC

Zürich

Vor Ort

CHF 180.000 - 250.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Competitive compensation
Career growth and learning
Flexibility and ownership
Collaborative culture
Impactful AI projects

Zusammenfassung

Token Factory is seeking a Senior Applied Scientist to turn frontier inference bottlenecks into production-ready solutions. You will design experiments, write code, and collaborate with engineers to ship results as production components.

You will lead research programs on efficient LLM/VLM inference, publish credible work, and mentor teams while building prototypes and open artifacts for external credibility.

Qualifikationen

  • PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field.
  • Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.
  • Strong hands-on coding ability in Python and PyTorch; ability to move from idea to experiment to prototype quickly.
  • Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs.
  • Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis.
  • Excellent written and verbal communication.

Aufgaben

  • Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff.
  • Prepare internal reports, technical blogs, or papers when the work is externally credible.
  • Partner directly with MLEs to ensure research prototypes become usable production components.
  • Define and execute research programs in efficient LLM and VLM inference with measurable production impact.
  • Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse, and model/runtime co-optimization.
  • Build high-quality prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks, then productionize them.
  • Design rigorous evaluation methodology covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.
  • Publish papers, technical reports, blog posts, and open-source artifacts that build external credibility for the company Token Factory.
  • Collaborate with MLE, GPU kernel, backend infrastructure, product, and customer teams to choose high-leverage research bets.
  • Mentor engineers and scientists on experimental design, scientific rigor, and model/system tradeoffs.

Kenntnisse

Python
PyTorch
Experiment design
Research publications
ML systems

Ausbildung

PhD in CS/ML/related field

Tools

CUDA
Triton

Jobbeschreibung

Token Factory is seeking a Senior Applied Scientist to turn frontier inference bottlenecks into production-ready solutions. You will design experiments, write code, and collaborate with engineers to ship results as production components.

You will lead research programs on efficient LLM/VLM inference, publish credible work, and mentor teams while building prototypes and open artifacts for external credibility.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Senior Applied Scientist, Efficient LLM Inference & Model Optimization

United States Digital Space LLC • Zürich

Vor Ort
CHF 180.000 - 250.000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
High-Performance Software Engineer — Verified ML Inference Engine
High-Performance Software Engineer — Verified ML Inference Engine

École polytechnique fédérale de Lausanne, EPFL • Lausanne

Vor Ort
CHF 110.000 - 150.000
Frontier AI models access
Travel for collaboration
Professional development
+1
Production-Grade Verifiable ML Inference Engineer
Production-Grade Verifiable ML Inference Engineer

EPFL • Lausanne

Vor Ort
CHF 120.000 - 180.000
Dynamic team on high-profile project
Multicultural academic environment
Continuing education and professional
+2
Hybrid Remote AI Research Scientist: LLMs & Multimodal
Hybrid Remote AI Research Scientist: LLMs & Multimodal

Embodied AI • Lausanne

Hybrid
CHF 120.000 - 170.000
AI Scientist - LLM Systems
AI Scientist - LLM Systems

Artificialy • Lugano

Vor Ort
CHF 120.000 - 180.000
Competitive compensation
Growth opportunities
Scientific environment
+1
Research Scientist - LLM Efficiency
Research Scientist - LLM Efficiency

Apple • Zürich

Vor Ort
CHF 180.000 - 240.000
Senior ML Engineer, Inference & Optimization
Senior ML Engineer, Inference & Optimization

Nebius Group • Zürich

Vor Ort
CHF 150.000 - 220.000
Competitive compensation
Career growth and learning机会
Flexibility and ownership
+3
AI Scientist: Frontier Models & Real-World Impact
AI Scientist: Frontier Models & Real-World Impact

Mistral Ai • Zürich

Vor Ort
CHF 120.000 - 180.000
AI Scientist - LLM Systems
AI Scientist - LLM Systems

Artificialy • Zürich

Vor Ort
CHF 130.000 - 180.000
AI Research Engineer - Scalable ML & NLP Innovator
AI Research Engineer - Scalable ML & NLP Innovator

Thomson Reuters • Zug

Hybrid
CHF 120.000 - 190.000
Hybrid work model
Flex My Way policies
Career development