Eine zielgenaue Bewerbung für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.
Oak Tree Software is seeking an experienced ML/ML Ops Engineer to build and operate production-grade ML infrastructure in a Germany-based role. You will implement end-to-end pipelines with Kubeflow, manage GPU resources in Kubernetes, and fine-tune LLMs while tracking experiments in MLflow.
You’ll work with SQL Server and DuckDB for fast in-cluster data processing, coding in Python, and adhering to zero-trust security policies in enterprise cloud environments.
Build and orchestrate ML pipelines using Kubeflow Pipelines (KFP v2)
Train models on GPUs, including GPU resource management within Kubernetes
Fine-tune transformers/LLMs, track experiments and models via MLflow
Build classic ML models (XGBoost, CatBoost)
Work with data using SQL Server and DuckDB as a lightweight OLAP solution for efficient in-cluster processing of large datasets
Develop in Python (pipelines, integrations, tooling based on uv)
Work within a zero-trust / secure-by-default environment (network policies, restrictive container rights)
Location: Germany
German language at B2 level is a must have
An experienced ML/ML Ops Engineer with a strong engineering background, combining:
Hands-on, production-grade ML infrastructure experience (not just notebook-level ML)
Deep proficiency in the Python ecosystem and modern engineering practices
Experience working in regulated/enterprise cloud-native environments
Willingness to quickly ramp up on the client's domain (healthcare/billing data), even without prior background in it