Data Platform Engineer – Data Operations (all genders)

Jackalope Digital LLC

München

Hybrid

EUR 70.000 - 110.000

Vollzeit

Vor 3 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

STARK is a defence technology company building autonomous systems and a scalable data platform. You will own the data backbone, ETL pipelines, and metadata systems to turn raw field data into ready-to-use datasets for perception and autonomy.

You will collaborate with data collectors, annotation vendors, and ML engineers to design automated workflows, data tooling, and reliable, high-quality data at scale in a startup environment.

Qualifikationen

  • Strong Python software engineering with typed, tested, maintainable code.
  • Experience designing, building, and operating robust ETL/data ingestion pipelines.
  • Solid SQL (PostgreSQL preferred) with data modeling and metadata systems.
  • Hands-on experience with object storage (GCP/GCS, S3) and basics of containerization (Docker) and CI/CD.
  • Unix/Linux proficiency; comfortable with shell scripting and Linux environments.
  • Ability to balance quick fixes with long-term architectural solutions in a startup.
  • Strong cross-functional communication, coordinating with vendors and non-technical stakeholders.
  • Automation mindset to reduce manual tasks and improve data workflows.

Aufgaben

  • Design, write, and maintain Python ETL pipelines that ingest and merge complex field recordings into GCP storage.
  • Build programmatic tools to sub-sample video and extract frames, packaging datasets.
  • Design and maintain metadata databases and data catalogs with solid SQL and schema design.
  • Own end-to-end data curation and labeling workflows with external subcontractors.
  • Develop backend APIs, dashboards, and data-access tools for the AI/engineering org.
  • Enforce software hygiene: testing, typing, documentation, logging, and CI/CD.
  • Thrive in a startup: tackle ambiguity, migrate legacy data, and automate repetitive tasks.

Kenntnisse

Python
ETL
SQL
Cloud
Unix/Linux
CI/CD
Documentation

Tools

Docker
GCS/S3
DVC
FiftyOne

Jobbeschreibung

About Us

STARK is a new kind of defence technology company revolutionizing the way autonomous systems are deployed across multiple domains. We design, develop and manufacture high-performance unmanned systems that are software-defined, mass-scalable, and cost-effective. This provides our operators with a decisive edge in highly contested environments.

We're focused on delivering deployable, high-performance systems — not future promises. In a time of rising threats, STARK is bolstering the technological edge of NATO Allies and their Partners to deter aggression and defend Europe — today.

About the team

The Data Operations team owns the entire data lifecycle behind STARK's AI stack: collection, acquisition, generation, curation, and management. We run our own data-collection campaigns across Europe, evaluate new sensors and platforms, and build the internal data platform that turns raw recordings into ready-to-use datasets. Everything we produce feeds directly into the perception and autonomy systems deployed on STARK's platforms — a real data advantage is built, not bought. The team is scaling up right now: real scope, direct impact, no legacy.

Your mission

Data is the fuel of STARK’s AI stack — you build the engine that makes it usable. You own the software backbone of our data platform: the metadata systems, ETL pipelines, data contracts, catalogs, databases, and internal tools that let engineers find, understand, validate, and reuse terabytes of multi-sensor field data in minutes, not days. You treat data context as a product: structured, searchable, version-aware, documented, and traceable from raw recording to processed asset, annotation delivery, dataset, and downstream ML workflow. Success in this role requires seamless cross-functional collaboration—acting as the central hub between the teams collecting data in the field, our external annotation vendors, and the ML engineers training the models. Your job is to design the automated workflows and data tooling that unify these groups, transforming raw operational data into reliable, high-quality systems at scale.

Responsibilities
  • Build the Data Backbone: Design, write, and maintain robust Python software and ETL pipelines that ingest, process, and merge complex field recordings (flight data, rosbags, video) and external deliveries into our GCP cloud storage.
  • Create Data Tooling: Build the programmatic tools that bridge raw data and downstream usage. This includes writing services to automatically sub-sample video feeds, extract valuable frames, package datasets, and build self-serve data access tools.
  • Database & Metadata Engineering: Design, implement, and maintain our metadata databases and data catalogs (tracking datasets, recordings, sensors, labels, and lineage) using solid SQL and schema design principles.
  • Data Operations & Labeling Workflows: Own the end-to-end technical workflows for data curation and labeling. You will build the operational tooling and coordinate with external labeling subcontractors to ensure high-quality data deliveries, track progress, and run automated QA.
  • Internal Tooling & APIs: Develop backend APIs, self-service dashboards, and data-access tools that allow the entire AI and engineering organization to quickly search, filter, and understand terabytes of multi-sensor data.
  • Drive Engineering Excellence: Establish and enforce good software engineering hygiene in a young codebase: rigorous testing, typing, documentation, logging, and CI/CD pipelines.
  • Generalist Problem Solving: Thrive in an evolving startup environment. Take on ambiguous problems, migrate legacy data, handle access management, and aggressively automate away manual support tasks.
Qualifications
  • Strong Software Engineering in Python: You write clean, typed, tested, and maintainable Python code. You approach data problems with a software developer's mindset.
  • Data Engineering & ETL: Proven experience designing, building, and operating robust ETL/data ingestion pipelines.
  • Database Mastery: Solid SQL (PostgreSQL preferred) with a strong grasp of schema design, data modeling, and metadata systems.
  • Cloud & Infrastructure: Hands-on experience with object storage (GCP/GCS, S3) and the basics of containerization (Docker) and CI/CD.
  • Unix/Linux Environments: Strong familiarity with Unix-based systems, shell scripting, and standard Unix tools. You are comfortable working natively in a Linux environment.
  • Pragmatic & Adaptable: You know how to balance a quick, scrappy fix with a long-term architectural solution, and you are highly comfortable with the changing priorities of a startup environment.
  • Communicator & Coordinator: You are comfortable working cross-functionally and coordinating with external vendors and non-technical stakeholders to drive data labeling and curation efforts.
  • Automation Mindset: You aren't allergic to jumping in to do support tasks, but you have the technical chops to automate them away so you never have to do them twice.
Nice to have
  • GenAI / Synthetic Data: Experience with synthetic data generation or GenAI-assisted workflows (auto-labeling, data augmentation, foundation-model-based curation).
  • Data Ecosystems: Familiarity with data versioning (DVC, LakeFS, FiftyOne), computer vision annotation formats (e.g., COCO), or large-scale data curation workflows.
  • Advanced GCP: Experience with GCP services beyond basic storage (BigQuery, Cloud Run, IAM).
  • Robotics / Complex Data: Exposure to robotics data formats (ROS bags, MCAP, PX4 logs) or handling heavy, multi-modal data streams (video, lidar).

A note on our process: we value critical thinking, grit, and the ability to learn over a perfect checklist. If you don't hit every bullet point but you love building data systems that real engineers depend on every day, we still want to hear from you.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Data Platform Engineer – Data Operations (all genders)
Data Platform Engineer – Data Operations (all genders)

United States Digital Space LLC • München

Vor Ort
Data Platform Engineer – Data Operations (all genders)
Data Platform Engineer – Data Operations (all genders)

Meyandy LLC • München

Hybrid
EUR 70.000 - 110.000
Data Platform Engineer – Data Operations (all genders)
Data Platform Engineer – Data Operations (all genders)

STARK • München

Vor Ort
EUR 85.000 - 125.000
Data Operations & Labeling Specialist (all genders)
Data Operations & Labeling Specialist (all genders)

Jackalope Digital LLC • München

Hybrid
EUR 40.000 - 56.000
Field Robotics Engineer – Data Operations (all genders)
Field Robotics Engineer – Data Operations (all genders)

Stark Defence • München

Vor Ort
EUR 70.000 - 100.000
Data Operations & Labeling Specialist (all genders)
Data Operations & Labeling Specialist (all genders)

Meyandy LLC • München

Hybrid
EUR 32.000 - 48.000
Data Operations & Labeling Specialist (all genders)
Data Operations & Labeling Specialist (all genders)

United States Digital Space LLC • Berlin

Vor Ort
EUR 42.000 - 65.000
Field Robotics Engineer – Data Operations (all genders)
Field Robotics Engineer – Data Operations (all genders)

STARK • München

Vor Ort
EUR 70.000 - 110.000
Field Robotics Engineer – Data Operations (all genders)
Field Robotics Engineer – Data Operations (all genders)

Meyandy LLC • München

Hybrid
EUR 90.000 - 120.000
Field Robotics Engineer – Data Operations (all genders)
Field Robotics Engineer – Data Operations (all genders)

Jackalope Digital LLC • München

Hybrid
EUR 70.000 - 110.000