Data Engineer (Detect & Track Distillation)

Blackshark

Zürich

Vor Ort

CHF 90.000 - 150.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Mach aus dieser Rolle ein Vorstellungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Harmattan AI is seeking a Data Engineer to own the data layer for autonomous defense ML models, working from Paris, Lausanne, or Zurich. You will turn raw field logs and public datasets into clean, versioned datasets for modelers.

You’ll shape data processing for deep-learning pipelines and enable efficient data delivery to training workloads. Ideal candidates have a STEM background, strong Python and data-engineering skills, and experience with unstructured data pipelines.

Qualifikationen

  • Degree in a STEM field or equivalent practical experience.
  • Experience building pipelines for unstructured data at scale (video, images, or sensor data).
  • Strong Python skills and data-loader optimization for training frameworks (e.g. PyTorch).
  • Bonus: multimodal sensor data, dataset versioning tooling, and distributed data processing tooling.

Aufgaben

  • Ingest, decode, and store raw, unstructured field data (video and other sensor streams) from field logs into efficient, controlled formats.
  • Multimodal alignment to make paired data usable for training.
  • Curation: convert ingested and public data into high-value datasets with parsing, filtering, de-duplication and revision.
  • Define data gathering and labeling requirements and own the dataset-construction workflow and tooling.
  • Build task-specific datasets for training and evaluation in collaboration with acquisition, annotation and product teams.
  • Version datasets and maintain lineage for reproducible training runs.
  • Store data in training-ready formats and manage storage tiering to balance cost and latency.
  • Deliver clean, documented datasets and tools for loading to keep training from being I/O-bound.

Kenntnisse

Python
Data engineering
Pipelines at scale
PyTorch

Ausbildung

STEM degree or equivalent

Jobbeschreibung

About Us

Harmattan AI is a next-generation defense prime building autonomous and scalable defense systems. Following the close of a $200M Series B, valuing the company at $1.4 billion, we are expanding our teams and capabilities to deliver mission-critical systems to allied forces.

Our work is guided by clear values: building technologies with real-world impact, pursuing excellence in everything we do, setting ambitious goals, and taking on the hardest technical challenges. We operate in a demanding environment where rigor, ownership, and execution are expected.

ABOUT THE ROLE

Our ML teams train models on datasets derived from large volumes of raw, unstructured data. Model quality depends directly on data quality, and today that data is handled largely by hand.

As a Data Engineer, operating out of Paris, Lausanne, or Zurich, you will own the data layer that feeds the team's models, from raw field logs or public datasets through curated, versioned, training-ready datasets. You will manage terabytes of raw, unstructured data and turn it into clean, documented, versioned datasets, so that the modelers spend their time designing and training models, not waiting on data loaders or wrangling corrupted files. You join at an early stage with real influence over how field and public data gets processed for deep-learning pipelines.

RESPONSIBILITIES
  • Ingestion Pipeline: Ingest, decode, and store raw, unstructured field data (video and other sensor streams) from field logs into efficient, controlled formats.

  • Multimodal Alignment: Align multiple data streams temporally and spatially so paired data is usable for training.

  • Curation: Transfer both ingested data and public datasets into high-value data, including parsing, filtering, de-duplication and revision.

  • Data & Labeling Requirements: Define which data to gather and what and how to label it, and own the dataset-construction workflow and labeling tooling. Coordinate with the teams responsible for data gathering and labeling.

  • Dataset Construction: Build task-specific datasets for the team’s training and evaluation needs, in collaboration with acquisition, annotation and product teams.

  • Versioning & Lineage: Version datasets and maintain lineage so training runs stay reproducible.

  • Storage & Formats: Store data in efficient, training-ready formats (such as columnar or sharded formats) and manage storage tiering to balance cost and latency as datasets grow.

  • Efficient Delivery: Deliver clean, documented datasets and the corresponding tools for loading to keep training from being I/O-bound, shaping their structure with the modelers.

CANDIDATE REQUIREMENTS
  • Educational Background: A degree in a STEM field, or equivalent practical experience. Practical data engineering experience matters more than the specific degree.

  • Data Pipelines at Scale: Built and maintained pipelines for unstructured data at scale (video, images, or sensor data), covering ingestion, decode, storage, curation, and versioning.

  • Engineering: Strong in Python and data engineering, and comfortable optimizing data loaders for common training frameworks (for example PyTorch).

  • Bonus: Multimodal sensor data, labeling or dataset construction for ML, dataset versioning tooling, and distributed data processing tooling.

  • Attributes: Systematic, quality-minded, pragmatic, and service-oriented so the modelers are enabled, with a knack for taming messy data via automation.

  • Commitment: 100% dedication to Harmattan AI's mission of providing a defensive edge to allied nations through ethical, high-impact technology.

We look forward to hearing how you can help shape the future of autonomous defense systems at Harmattan AI.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Data Engineer (Detect & Track Distillation)
Data Engineer (Detect & Track Distillation)

Harmattan AI • Lausanne

Hybrid
CHF 110.000 - 170.000
Data Engineer (Detect & Track Distillation)
Data Engineer (Detect & Track Distillation)

Harmattan AI • Zürich

Hybrid
CHF 120.000 - 170.000
Machine Learning Engineer (Detect & Track Distillation)
Machine Learning Engineer (Detect & Track Distillation)

Blackshark • Zürich

Vor Ort
CHF 120.000 - 180.000
Machine Learning Engineer (Detect & Track Distillation)
Machine Learning Engineer (Detect & Track Distillation)

Harmattan AI • Zürich

Vor Ort
CHF 120.000 - 180.000
Machine Learning Engineer - Foundational
Machine Learning Engineer - Foundational

Harmattan AI • Zürich

Vor Ort
CHF 110.000 - 150.000
Data Engineer, ML Data Pipelines & Versioned Datasets
Data Engineer, ML Data Pipelines & Versioned Datasets

Blackshark • Zürich

Vor Ort
CHF 90.000 - 150.000
Autonomous Defense Data Engineer
Autonomous Defense Data Engineer

Harmattan AI • Zürich

Hybrid
CHF 120.000 - 170.000
Computer Vision Engineer
Computer Vision Engineer

Harmattan AI • Zürich

Vor Ort
CHF 100.000 - 130.000
Director of Engineering - Mission Intelligence
Director of Engineering - Mission Intelligence

Harmattan AI • Zürich

Vor Ort
CHF 180.000 - 300.000
ML Ops Engineer
ML Ops Engineer

Harmattan AI • Zürich

Vor Ort
CHF 90.000 - 150.000