Machine Learning Engineer

Articul8

Dublin

On-site

EUR 90,000 - 130,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Articul8 AI in Ireland is seeking machine learning engineers to join our team full-time. You will help build pipelines for data collection, extraction, filtering, synthetic data generation, and analysis to power domain-specific models end to end.

You will own data acquisition projects, collaborate with researchers and engineers, and ship new systems quickly, with a focus on rapid prototyping, data quality, and scalability across large-scale datasets and diverse sources.

Qualifications

  • BS/MS/PhD in Computer Science or a related field.
  • Proficiency in at least one deep learning framework, such as PyTorch.
  • Experience in ML projects in text or vision, training models.
  • Strong expertise in large stateful distributed systems and data processing.
  • Strong proficiency in building large-scale data processing pipelines with distributed workloads (multiprocessing, Ray, Docker, Kubernetes).
  • Proficiency in Python and writing clean, maintainable code.
  • Excellent problem-solving and attention to data quality, handling anomalies and biases.

Responsibilities

  • Design and develop data processing pipelines, including data extraction, data filtering, data labeling, etc.
  • Implement machine learning models to improve the quality and diversity of data (e.g., quality classifier, document layout model, code verification model).
  • Own and lead engineering projects in the area of data acquisition, including web crawling, data ingestion, and processing.
  • Collaborate with our Applied Research, Technology, and Architecture teams to ensure smooth data flow and system operability.
  • Develop and deploy highly scalable distributed systems capable of handling terabytes of data.
  • Architect and implement algorithms for data indexing and search capabilities.
  • Build and maintain backend services for data storage, including work with key-value databases and synchronization.
  • Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks.

Skills

PyTorch
Text/vision ML
Distributed systems
Data pipelines
Python
Data quality
GitHub contributions
Large datasets
Web crawling/tools
Custom data libs
Multilingual

Education

BS/MS/PhD in Computer Science or related field

Tools

Hadoop
Datasketch
Scrapy
Selenium
Ray
Docker
Kubernetes

Job description

About us:

At Articul8 AI, we relentlessly pursue excellence and create exceptional AI products that exceed customer expectations. We are a team of dedicated individuals who take pride in our work and strive for greatness in every aspect of our business. We believe in using our advantages to make a positive impact on the world and inspiring others to do the same.

Job Description:

We are seeking machine learning engineers to join our team full-time. As part of your role, you will help us build pipelines of data collection, data extraction, data filtering/synthetic data generation and data analysis. You will own all work related to acquiring high-quality data to power the training of our domain-specific models end to end. You will work closely with other researchers and engineers to empower our next generation of domain-specific models. We value rapid prototyping, iterating, and shipping new systems quickly.

Required Qualifications:
  • BS/MS/PhD in Computer Science or a related field.

  • Proficiency in at least one deep learning framework, such as PyTorch.

  • Experience in machine learning projects in text or vision, e.g., has trained machine learning models to tackle a specific problem.

  • Strong expertise in large stateful distributed systems and data processing.

  • Strong proficiency in building large-scale data processing pipelines, familiar with distributed workload (e.g., multiprocessing, Ray, Docker, Kubernetes).

  • Proficiency in at least one programming language commonly used in machine learning, such as Python and ability to write clean, maintainable code.

  • Excellent problem-solving skills and attention to detail, especially when handling data anomalies and biases to further improve data quality.

Key Competencies
  • Active Github contributions are a big plus.

  • Experience in building large-scale datasets.

  • Familiar with at least one of the following tools for data crawling (e.g. Scrapy), data collection (e.g., VPNs, Selenium), data processing (e.g., Hadoop, Datasketch).

  • Building bespoke data processing libraries from scratch.

  • Keeping up with state-of-the-art techniques for preparing AI training data.

  • Organizing and meticulously bookkeeping data across multiple clouds, of multiple modalities, and from many sources.

  • Multilingual which contributes to enriching the language diversity crucial for robust model training.

Responsibilities:
  • Design and develop data processing pipelines, including data extraction, data filtering, data labeling, etc.

  • Implement machine learning models to improve the quality and diversity of data (especially in the data extraction stage), e.g., quality classifier, document layout model, code verification model, etc.

  • Own and lead engineering projects in the area of data acquisition, including web crawling, data ingestion, and processing.

  • Collaborate with our Applied Research, Technology, and Architecture teams to ensure smooth data flow and system operability.

  • Develop and deploy highly scalable distributed systems capable of handling terrabytes of data.

  • Architect and implement algorithms for data indexing and search capabilities.

  • Build and maintain backend services for data storage, including work with key-value databases and synchronization.

  • Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks.

By joining our team, you become part of a community that embraces diversity, inclusiveness, and lifelong learning. We nurture curiosity and creativity, encouraging exploration beyond conventional wisdom. Through mentorship, knowledge exchange, and constructive feedback, we cultivate an environment that supports both personal and professional development.

Your future experience at Articul8 will include continuous learning and growth opportunities as we embark on an exciting journey to disrupt the status quo. If you're excited about joining a team that's passionate about making a difference, we want to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Researcher (Dublin, CA)
Applied AI Researcher (Dublin, CA)

Articul8 • Dublin

On-site
EUR 70,000 - 90,000
AI/ML Computational Scientist Manager
AI/ML Computational Scientist Manager

Accenture • Dublin

On-site
EUR 90,000 - 130,000
Machine Learning Engineer - Data Pipelines & Training Data
Machine Learning Engineer - Data Pipelines & Training Data

Articul8 • Dublin

On-site
EUR 90,000 - 130,000
Senior AI/ML Computational Scientist
Senior AI/ML Computational Scientist

Accenture • Dublin

On-site
EUR 110,000 - 180,000
Senior AI/ML Computational Scientist
Senior AI/ML Computational Scientist

Accenture UK & Ireland • Dublin

On-site
EUR 140,000 - 210,000
Principal Applied AI Researcher - Domain- Specific Models (Dublin, CA)
Principal Applied AI Researcher - Domain- Specific Models (Dublin, CA)

Articul8 • Dublin

On-site
EUR 120,000 - 150,000
Member of Technical Staff, Machine Learning
Member of Technical Staff, Machine Learning

ActAI • Ireland

On-site
EUR 90,000 - 130,000
Senior Applied AI Researcher (Dublin, CA)
Senior Applied AI Researcher (Dublin, CA)

Articul8 • Dublin

On-site
EUR 120,000 - 180,000
Lead Data Scientist (AI Engineering)
Lead Data Scientist (AI Engineering)

Mastercard • Ireland

On-site
EUR 130,000 - 170,000
Senior Applied AI Researcher - End-to-End Agentic Systems
Senior Applied AI Researcher - End-to-End Agentic Systems

Articul8 • Dublin

On-site
EUR 120,000 - 180,000