Research Engineer (Agentic Models)

United States Digital Space LLC

Berlin, München

Vor Ort

EUR 90.000 - 120.000

Vollzeit

Vor 11 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

United States Digital Space LLC in Berlin is seeking a Research Engineer to design and optimize SFT and RL post-training pipelines for multi-step coding agents integrated into developer IDEs.

You will work with PyTorch-based stacks, large-scale distributed training on GPU clusters, and evaluation loops to shape models that power our agentic workflows. The role emphasizes end-to-end ownership and collaboration with product and infra teams.

Qualifikationen

  • Extensive hands-on experience training LLMs (pre-training, fine-tuning, or post-training) in a research or production setting.
  • Strong practical knowledge of PyTorch and large-scale ML tooling.
  • Experience with distributed training on GPU clusters and data pipelines.

Aufgaben

  • Design, implement, and maintain SFT and RL post-training pipelines for multi-step coding agents.
  • Train and adapt LLMs for agent workflows including planning and tool use.
  • Build evaluation environments to measure agent behavior and iterate on training data and rewards.
  • Collaborate with research, product, and infra teams to ship features.

Kenntnisse

Python
PyTorch
LLM training
Distributed training
Data pipelines

Tools

Megatron
NeMo
verl

Jobbeschreibung

At the company, code is our passion. Ever since we started, back in 2000, we’ve been striving to make the strongest, most effective developer tools on earth. Today, AI-powered assistance and agents are becoming a core part of how developers work in our IDEs.

We’re building multi-step coding agents that can understand large codebases, plan changes, call tools, and iterate with the user. As a Research Engineer in the Agentic Models team, you’ll be responsible for the models, training loops, and evaluation pipelines that power these agents.

You’ll work at the intersection of SFT and RL-style post-training, and product-driven evaluation, using our distributed GPU and MapReduce clusters to ship models into the company products.

As part of our team, you will:
  • Design, implement, and maintain SFT and RL post-training pipelines for multi-step coding agents.
  • Train and adapt LLMs for agent workflows, including planning, tool use, and multi-step interactions inside the company IDEs.
  • Build and develop evaluation and simulation environments where coding agents can act, be measured, and compared on realistic developer tasks.
  • Design evaluation frameworks and metrics for agent behavior, analyze traces and logs, and close the loop from evaluation back into training, data, and reward design.
  • Analyze training and evaluation results to propose and implement improvements to model architectures, training recipes, and datasets.
  • Work with large-scale infrastructure, including distributed training on GPU clusters and large MapReduce-style data processing for pre-training and fine-tuning datasets.
  • Collaborate closely with research, product, and infrastructure teams to turn high-level product visions into concrete models, experiments, and shipped features.
We’ll be happy to bring you on board if you have:
  • Extensive hands-on experience training LLMs (pre-training, fine-tuning, or post-training) in a research or production setting.
  • Deep expertise in modern deep learning frameworks such as PyTorch, and specialized LLM training stacks (e.g. Megatron, NeMo, verl, or similar).
  • Strong theoretical and practical understanding of LLM fundamentals: architectures, tokenization, data pipelines, batching, mixed precision, distributed training, and debugging unstable runs.
  • The ability to own projects end to end, starting from a high-level problem or product pain point and overseeing it through the design, experimentation, implementation, and iteration phases.
  • A product-aware mindset – you care about how developers actually use agents and can translate product needs and failure modes into modeling and evaluation work.
  • At least 3 years of Python experience writing clean, maintainable code in modern ML codebases.
Our ideal candidate would have experience with:
  • ML orchestrators and workflow tools such as Kubeflow, Dagster, Airflow, ZenML, and/or job schedulers like Kubernetes or SLURM.
  • Large-scale data and training pipelines, e.g. MapReduce-style clusters, multi-node GPU training, or workloads on the order of 1M+ CPU/GPU hours.
  • Designing and maintaining evaluation pipelines for LLMs or agents, including metrics, dashboards, experiment tracking, and automated regression checks.
  • AI agent development, such as tool-using agents, planners, or multi-step coding workflows, and familiarity with agentic frameworks or patterns.
  • Experiment tracking and observability using tools like Weights & Biases, MLflow, Langfuse, or similar.
  • Inference optimization and serving optimized models in production.

#LI-KP1

We are an equal opportunity employer

We know great ideas can come from anyone, anywhere. That’s why we do our best to create an open and inclusive workplace – one that welcomes everyone regardless of their background, identity, religion, age, accessibility needs, or orientation.

We process the data provided in your job application in accordance with the Recruitment Privacy Policy.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Research Engineer (Agentic Models)
Research Engineer (Agentic Models)

JetBrains • München

Vor Ort
EUR 70.000 - 95.000
Staff Research Engineer (LLM Pre-Training)
Staff Research Engineer (LLM Pre-Training)

United States Digital Space LLC • Berlin, München

Hybrid
EUR 90.000 - 140.000
Staff Research Engineer (LLM Pre-Training)
Staff Research Engineer (LLM Pre-Training)

JetBrains • Berlin

Vor Ort
EUR 70.000 - 90.000
Staff Research Engineer (LLM Pre-Training)
Staff Research Engineer (LLM Pre-Training)

JetBrains • München

Vor Ort
EUR 60.000 - 85.000
Senior ML Engineer (JetBrains Research)
Senior ML Engineer (JetBrains Research)

United States Digital Space LLC • Berlin, München

Hybrid
EUR 70.000 - 110.000
Senior AI Engineer (Core Engine)
Senior AI Engineer (Core Engine)

United States Digital Space LLC • Berlin, München

Hybrid
EUR 80.000 - 110.000
Strong base salary
Flexible work location
Remote work up to 30 days abroad
+9
Research Engineer (Kineto) (m/f/d)
Research Engineer (Kineto) (m/f/d)

JetBrains GmbH • München

Vor Ort
EUR 85.000 - 120.000
AI Engineer (Germany)
AI Engineer (Germany)

United States Digital Space LLC • Berlin

Hybrid
EUR 70.000 - 120.000
Flexible PTO
Health insurance
Employee assistance programs
+4
Senior AI Engineer - Agentic AI Evaluation Brain Team · Munich, Singapore ·
Senior AI Engineer - Agentic AI Evaluation Brain Team · Munich, Singapore ·

Resaro • München

Hybrid
EUR 90.000 - 130.000
ML Research Engineer - PhD
ML Research Engineer - PhD

Obsidian • Berlin

Vor Ort
EUR 60.000 - 90.000