Senior ML Engineer (Token Factory)

Jobgether

Germany (OH)

On-site

USD 150,000 - 210,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
International environment
High ownership

Job summary

Jobgether is seeking an experienced AI infrastructure engineer to work on the cutting edge of foundation models and ML systems. You will help build inference and fine-tuning technologies for large-scale AI platforms, focusing on quality, efficiency, and scalable deployment.

You will collaborate with researchers and engineers in a fast-moving, international environment, applying strong software engineering practices and contributing to the evolution of AI infrastructure.

Qualifications

  • Deep understanding of ML foundations and reinforcement learning.
  • Strong expertise in modern DL techniques for language processing and generation.
  • Experience training large ML models across multiple compute nodes.
  • Solid understanding of performance optimization for large neural networks, including sharding and custom kernels.
  • Strong Python programming and software engineering skills.
  • Extensive experience with modern DL frameworks, especially JAX.
  • CI/CD, version control, unit testing, production-grade practices.

Responsibilities

  • Develop and improve fine-tuning methodologies for foundation models (LoRA-based and full-parameter).
  • Optimize model quality and training efficiency across large-scale ML workloads.
  • Address bottlenecks in LLM inference to improve production performance and resource use.
  • Build training and evaluation pipelines using JAX with advanced inference techniques.
  • Experiment with model architectures including dense and mixture-of-experts models.
  • Develop scaling laws to guide model development and resource allocation.
  • Investigate low-precision training/inference approaches (FP8, MXFP4, etc.).
  • Work with distributed training environments across multiple GPU clusters.
  • Analyze performance factors like sharding, custom kernels, hardware capabilities.
  • Translate research into robust, scalable, production ML systems.

Skills

LoRA fine-tuning
Full-parameter tuning
Distributed training
Python
JAX
Model optimization
Inference optimization
Hardware utilization
Speculative decoding
Multi-node training

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps.

This role offers the opportunity to work at the forefront of large-scale AI infrastructure and machine learning systems.
You will help build inference and fine-tuning technologies for foundation models spanning language, vision, audio, and multimodal architectures.
Your work will focus on improving model quality, training efficiency, inference performance, and hardware utilization at massive scale.
You will tackle technically challenging problems involving distributed training, low-precision computation, optimization, and reinforcement learning.
Working primarily with Python and JAX, you will turn advanced research ideas into reliable, production-ready systems.
The role combines deep technical ownership with opportunities to influence engineering practices and contribute to the evolution of AI platforms.
You will collaborate with highly experienced engineers and researchers in a fast-moving, international environment where your work can have significant impact.

Accountabilities
  • Develop and improve advanced fine-tuning methodologies, including LoRA-based and full-parameter approaches, for cutting-edge foundation models.
  • Optimize model quality and training efficiency across large-scale machine learning workloads.
  • Identify and address bottlenecks in large language model inference to improve production performance and resource efficiency.
  • Build training and evaluation pipelines using JAX for techniques such as speculative decoding and advanced inference optimization.
  • Experiment with different model architectures, including dense and mixture-of-experts models as well as autoregressive and parallel approaches.
  • Develop and evaluate scaling laws to inform model development, performance optimization, and resource allocation.
  • Investigate low-precision training and inference approaches, including FP8, NVFP4, and MXFP4, for supervised fine-tuning and reinforcement learning.
  • Work with distributed training environments spanning multiple computational nodes and large GPU clusters.
  • Analyze performance considerations such as sharding strategies, custom kernels, and modern hardware capabilities.
  • Translate research concepts and experimental results into robust, scalable, production-quality machine learning systems.
  • Apply strong software engineering practices, including CI/CD, version control, unit testing, and maintainable code design.
  • Collaborate across engineering and research teams while communicating technical concepts clearly and contributing to technical direction.
Requirements
  • Deep understanding of the theoretical foundations of machine learning and reinforcement learning.
  • Strong expertise in modern deep learning techniques for language processing and generation.
  • Demonstrated experience training large machine learning models across multiple computational nodes.
  • Solid understanding of performance optimization for large neural network training, including sharding strategies, custom kernels, and hardware-specific capabilities.
  • Strong software engineering skills, particularly with Python.
  • Extensive experience with modern deep learning frameworks, particularly JAX.
  • Proficiency in contemporary software development practices, including CI/CD, version control, unit testing, and production-quality engineering.
  • Strong communication, collaboration, and technical leadership abilities.
  • Experience working with language models or related NLP technologies is highly valued.
  • Familiarity with concepts such as multi-head attention, RoPE, ZeRO/FSDP, Flash Attention, and quantization is advantageous.
  • Experience building and delivering products in dynamic, startup-like environments is a plus.
  • Strong engineering background in distributed systems or high-load web services is beneficial.
  • Open-source projects demonstrating advanced engineering capabilities are valued.
  • Excellent English communication skills, including strong technical writing and articulation.
Benefits
  • Competitive compensation.
  • Career growth and continuous learning opportunities.
  • Flexible working environment with a high degree of ownership.
  • Opportunity to work on impactful, large-scale AI and machine learning projects.
  • Collaborative culture with experienced engineers and researchers.
  • International environment with diverse and highly skilled teams.
  • Opportunity to contribute to advanced foundation model training, fine-tuning, inference optimization, and AI infrastructure.
  • Exposure to cutting-edge GPU computing, distributed systems, and modern machine learning technologies.
  • Inclusive workplace committed to equal employment opportunities.
  • Support and reasonable accommodations throughout the hiring process when required.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist - Distributed Machine Learning
Research Scientist - Distributed Machine Learning

Ifm Us • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Health benefits
401K Plan
Paid time off
+2
Applied AI Engineer
Applied AI Engineer

Norbert Health • New York (NY)

On-site
USD 100,000 - 150,000
Equity participation
Competitive salary
High autonomy and technical ownership
Technical Lead - Machine Learning
Technical Lead - Machine Learning

USA Tech Recruit • San Francisco (CA)

On-site
USD 180,000 - 230,000
AI Research Engineer
AI Research Engineer

Fuel Talent • Seattle (WA)

On-site
USD 147,000 - 220,000
Visa support
Open-source collaboration
Seattle-based opportunity
Senior Technical Program Manager (Engineering) - AI Tooling & Systems
Senior Technical Program Manager (Engineering) - AI Tooling & Systems

AI Chopping Block • United States

On-site
USD 180,000 - 240,000
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Research Engineer, Applied AI
Research Engineer, Applied AI

Akoncagua AI • Lakeland (FL)

On-site
USD 120,000 - 180,000
Senior Technical Program Manager (Engineering) - AI Tooling & Systems
Senior Technical Program Manager (Engineering) - AI Tooling & Systems

Madrona Venture Labs • United States

On-site
USD 150,000 - 230,000
Senior Technical Program Manager (Engineering) - AI Tooling & Systems
Senior Technical Program Manager (Engineering) - AI Tooling & Systems

Deepgram • United States

On-site
USD 180,000 - 240,000
AIML Engineer
AIML Engineer

Qubeaxis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Performance bonus (up to 20% of base)
Equity participation
Health, dental, and vision insurance
+3