AI Data Ingestion Engineer for Scalable Pipelines

Clutch Canada

Monterrey

Presencial

MXN 1.200.000 - 1.800.000

Jornada completa

14 días+
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Competitive salaries
Fast-growing environment
Asynchronous culture

Descripción de la vacante

Speechify is hiring for its data side of the AI team. The role focuses on data collection, ingestion, and supporting model training by building high-quality datasets at petabyte scale while controlling costs through close infra, engineering, and research collaboration.

The ideal candidate has a background in software development (5+ years) and strong scripting, Docker, and cloud skills, with a passion for scalable data infrastructure and AI-driven products.

Formación

  • BS/MS/PhD in Computer Science or related field.
  • 5+ years of industry experience in software development.
  • Proficiency with bash/Python scripting in Linux environments.
  • Proficiency in Docker and Infrastructure-as-Code concepts and experience with GCP.
  • Experience with web crawlers and large-scale data processing workflows is a plus.
  • Ability to handle multiple tasks and adapt to changing priorities.
  • Strong communication skills, both written and verbal.

Responsabilidades

  • Be scrappy to find new sources of audio data and bring it into our ingestion pipeline.
  • Operate and extend the cloud infrastructure for our ingestion pipeline, currently running on GCP and managed with Terraform.
  • Collaborate closely with our Scientists to shift the cost/throughput/quality frontier, delivering richer data at bigger scale and lower cost to power our next-generation models.
  • Collaborate with others on the AI Team and Speechify Leadership to craft the AI Team’s dataset roadmap to power Speechify’s next-generation consumer and enterprise products.

Conocimientos

bash/Python scripting
Docker
Infrastructure as Code
GCP
Web crawlers
Data processing workflows
Strong communication

Educación

BS/MS/PhD in Computer Science

Herramientas

Terraform

Descripción del empleo

Speechify is hiring for its data side of the AI team. The role focuses on data collection, ingestion, and supporting model training by building high-quality datasets at petabyte scale while controlling costs through close infra, engineering, and research collaboration.

The ideal candidate has a background in software development (5+ years) and strong scripting, Docker, and cloud skills, with a passion for scalable data infrastructure and AI-driven products.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote Data Infrastructure Engineer for AI Pipelines
Remote Data Infrastructure Engineer for AI Pipelines

Speechify • Región Centro

Presencial
MXN 900.000 - 1.200.000
Competitive salaries
Asynchronous culture
Diversity & inclusion
+1
Senior Data Engineer - Build AI-Driven Pipelines (Remote)
Senior Data Engineer - Build AI-Driven Pipelines (Remote)

AgileEngine • Puebla de Zaragoza

Presencial
MXN 900.000 - 1.300.000
Growth without limits
Competitive compensation
100% remote with flexible hours
+3
Senior AI Data Engineer: Build Scalable AI Data Pipelines
Senior AI Data Engineer: Build Scalable AI Data Pipelines

Strategic Systems International • Ciudad de México

Presencial
MXN 900.000 - 1.500.000
Senior AI Data Engineer
Senior AI Data Engineer

Strategic Systems International • Ciudad de México

Presencial
MXN 900.000 - 1.500.000
Data Infrastructure Engineer: Ingest & Scale AI Datasets
Data Infrastructure Engineer: Ingest & Scale AI Datasets

Clutch Canada • Ciudad de México

Presencial
MXN 1.040.000 - 1.387.000
Fast-growing environment
Entrepreneurial-minded team
Hands-off management approach
+2
Software Engineer, Data Infrastructure & Acquisition - Mexico City, Mexico
Software Engineer, Data Infrastructure & Acquisition - Mexico City, Mexico

Clutch Canada • Ciudad de México

Presencial
MXN 1.040.000 - 1.387.000
Fast-growing environment
Entrepreneurial-minded team
Hands-off management approach
+2
Data Engineer – Spark Pipelines & Databricks | Flexible
Data Engineer – Spark Pipelines & Databricks | Flexible

NielsenIQ • México

Híbrido
MXN 600.000 - 1.000.000
Flexible working environment
Volunteer time off
Remote Data Engineer — Scalable Pipelines & Lakehouse
Remote Data Engineer — Scalable Pipelines & Lakehouse

PROGRAMMING.COM • Azcapotzalco

Presencial
MXN 520.000 - 760.000
AI Data Architect & Solutions Engineer for Scalable Data
AI Data Architect & Solutions Engineer for Scalable Data

BMC Software • México

Híbrido
MXN 1.000.000 - 1.400.000
Data Engineer: Build Scalable AI Data Pipelines
Data Engineer: Build Scalable AI Data Pipelines

McKinsey & Company, Inc. • Monterrey

Presencial
MXN 1.014.000 - 1.521.000