Senior Data Scientist

Intellias

Polska

On-site

PLN 220,000 - 320,000

Full time

13 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Intellias is building a modern data and AI platform to empower real-time decision making and enterprise-scale analytics. You will lead platform setup on Databricks, manage secure data ingestion, and implement scalable feature engineering for pseudonymised entity resolution and image processing at scale.

You’ll collaborate with customers’ platform teams to ensure robust, reproducible workflows and governance.

Qualifications

  • Hands-on experience with PySpark and Spark SQL on large-scale data.
  • Proficient in Databricks including Unity Catalog and Delta Lake.
  • Experience with AWS and data security best practices.
  • Strong knowledge of MLOps and model monitoring.

Responsibilities

  • Platform setup and configuration of Databricks environment.
  • Secure data ingestion with pseudonymisation and data quality checks.
  • Develop scalable feature engineering pipelines.
  • Support modelling experiments and comparative analyses.
  • Automate end-to-end workflows and monitoring.
  • Document runbooks and provide handover.

Skills

PySpark
Spark SQL
Databricks
AWS
Pseudonymisation
Entity resolution
Image processing
PyTorch
MLOps
Governance
Secrets Manager

Tools

Unity Catalog
Delta Lake
Auto Loader
S3
KMS

Job description

Let's breathe life into great tech ideas! With 3,000 people globally, Intellias is a company where benchmark technological solutions are born. Join in and take your part in digitalizing the world.

Project Overview:

Join a transformative data and AI platform initiative aimed at modernizing enterprise-scale capabilities and enabling real-time decision-making. This project delivers a comprehensive roadmap covering AI, MLOps, data governance, and platform scalability, supporting a shift towards data-first operations and intelligent automation.

Requirements:
  • PySpark & Spark SQL: Nested-JSON flattening (explode, structs),`mapInPandas` / pandas UDFs, window functions for velocity and distinct-count aggregates, partitioning and performance tuning
  • Databricks: Unity Catalog (catalogs, schemas, grants, column masks, tags, lineage), external locations on S3, cluster / compute policies incl. GPU, Auto Loader, Delta (MERGE, time travel, VACUUM), Workflows, Repos, secret scopes
  • AWS: S3, IAM roles / instance profiles, KMS decryption in jobs, Secrets Manager, basic cost awareness
  • Pseudonymisation in pipelines: HMAC tokenization with per-key-type keys, normalisation (Unicode NFKD, case, token sort), phonetic keys, never logging raw values, scripted deletion (DROP / VACUUM / secret destruction)
  • Entity resolution (implementation): Blocking / candidate generation with exact and approximate keys, edit-distance tolerance, match scores, surrogate IDs, conflict counting rather than merging
  • Image processing at scale: Running a GPU batch job that decrypts, crops and embeds images in memory (PyTorch / timm), PCA, LSH bucketing
  • Mixed-data clustering: Hands-on: `kmodes` (k-prototypes), `gower` + `kmedoids` with sampling, R `kamila` on Databricks, `StepMix` (latent class), bootstrap ARI
  • Feature selection: Mutual information, correlation pruning, PCA. Can run a BPSO / GA search with`mealpy`
  • MLOps & monitoring: M*Lflow experiments and registry (aliases),Lakehouse Monitoring* (profile and drift metrics on Delta tables), SQL alerts, PSI / JS divergence
  • Handover: Writes runbooks and README-level documentation. Leaves a single end-to-end Workflow that someone else can run
  • Nice to have
  • Experience with AssureID / identity-verification JSON outputs.
  • PyOD and basic PyTorch training (to support A on the VAEs in week 2).
Responsibilities:
  • Platform setup. With customer's platform team, provision and configure the Databricks environment in customer's cloud account: catalog and permissions, compute (including GPU), secrets, experiment tracking.
  • Secure data ingestion. Build repeatable ingestion of raw verification outputs and document images: parsing, pseudonymisation, data-quality checks, schema handling.
  • Feature engineering at scale. Implement the feature tables:
    • batch image-embedding jobs;
    • exact and approximate linkage keys;
    • cross-transaction and velocity aggregates;
    • pseudonymous entity resolution.
  • Modelling support. Implement feature-selection and mixed-data clustering methods. Run model comparisons and stability tests, and support density-model training runs.
  • Automation. Orchestrate the end-to-end pipeline as a scheduled, idempotent, re-runnable workflow.
  • Monitoring. Set up the data-quality and drift-monitoring prototype, with alerting, on feature and score tables.
  • Privacy engineering. Implement access controls, provenance fields, lineage and the scripted deletion procedure. Confirm that no raw personal data reaches the analytic layer.
  • Documentation and handover. Write runbooks and give a live walkthrough, so engineers can operate everything independently
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer \\ Data Scientist — Unstructured Data & AI Pipelines
Data Engineer \\ Data Scientist — Unstructured Data & AI Pipelines

Neurons Lab LTD. • Warszawa

On-site
PLN 83,000 - 152,000
Data Engineer (Databricks)
Data Engineer (Databricks)

Addepto • Poland

On-site
PLN 80,000 - 100,000
Flexible work arrangements
20 fully paid days off
Medical and sports packages
+1
Senior Data Engineer
Senior Data Engineer

Luxoft Poland • Poland

On-site
PLN 180,000 - 280,000
Internal Mobility program
Private medical & dental care & life保险
Data Engineer
Data Engineer

Neurons Lab • Poland

On-site
PLN 78,000 - 112,000
Senior Data Engineer (Investment Data)
Senior Data Engineer (Investment Data)

Luxoft Germany • Polska

On-site
PLN 307,000 - 482,000
Azure Data Engineer
Azure Data Engineer

Aon plc • Kraków

Hybrid
PLN 150,000 - 230,000
Senior Data Engineer (Databricks Migration)
Senior Data Engineer (Databricks Migration)

Sigma Software • Kraków

On-site
PLN 240,000 - 360,000
Senior Data Engineer (Databricks on Azure)
Senior Data Engineer (Databricks on Azure)

Luxoft Poland • Poland

On-site
PLN 180,000 - 320,000
Databricks Data Architect (Healthcare & Life Sciences)
Databricks Data Architect (Healthcare & Life Sciences)

Entrada • Polska

Remote
PLN 180,000 - 360,000
Remote work from Poland
MacBook provided
Referral bonus
+2
Senior AI/Data Engineer
Senior AI/Data Engineer

Xebiacee • Poland

On-site
PLN 180,000 - 300,000