Global Data Ml Engineer For Multilingual Speech & Ai

Cartesia

San Francisco (CA)

On-site

USD 103,000 - 171,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Deloitte is seeking a Global Data ML Engineer for Multilingual Speech AI in the San Francisco area to design and manage scalable data pipelines on AWS and Snowflake. You will transform source data into analytics-ready datasets and build robust ETL/ELT processes across teams.

The role requires 7+ years of experience in data engineering, Python/SQL, and experience with Lakehouse architectures, with some travel to client sites. Hybrid/onsite collaboration is common in this engagement.

Qualifications

  • 7+ years of experience delivering production-grade data pipelines and curated datasets.
  • 7+ years of hands-on experience with SQL and Python, including Snowflake and/or PySpark for scalable data processing and ELT.
  • 7+ years of experience designing, building, and operating batch and near-real-time data pipelines on cloud platforms (AWS preferred; Azure/GCP acceptable).
  • Experience with data integration frameworks and orchestration tools.
  • Proficiency in designing and implementing Lakehouse/warehouse architectures and ELT patterns.
  • Knowledge of DevOps principles: CI/CD pipelines, version control, Infrastructure-as-Code.
  • Ability to optimise data storage, partitioning, file formats (Delta, Parquet), and performance.
  • Understanding of data quality, data governance, and metadata management.

Responsibilities

  • Architect, build, and operate scalable batch and near-real-time data pipelines on AWS.
  • Design robust ingestion patterns from source systems into S3 and into Snowflake.
  • Develop transformation layers and curated datasets in Snowflake, including dimensional/data product modeling for analytics and downstream applications.
  • Implement orchestration and workflow automation on AWS with retries, backfills, and idempotency.
  • Build reusable Python components for ingestion, validation, and transformations; enforce standards via code reviews and testing.
  • Optimize Snowflake performance and cost warehouse sizing, concurrency patterns, query tuning, clustering/micro-partition considerations, and workload isolation.
  • Partner with stakeholders to translate requirements into well-defined datasets and data contracts.
  • Communicate regularly with Engagement Managers (Directors), project team members, and representatives from various teams, escalating as needed.
  • Independently and collaboratively lead client engagement workstreams focused on improvement, optimisation, and transformation of processes.

Skills

SQL
Python
Snowflake
PySpark
AWS
ETL
Data Pipelines
DevOps
CI/CD
Data Governance

Education

Bachelor's degree

Tools

Airflow

Job description

Global Data Ml Engineer For Multilingual Speech AI

Are you an experienced, passionate pioneer in technology who wants to work in a collaborative environment? As an experienced Data Engineer you will have the ability to share new ideas and collaborate on projects as a consultant without the extensive demands of travel. If so, consider an opportunity with Deloitte under our Project Delivery Talent Model. Project Delivery Model (PDM) is a talent model that is tailored specifically for long-term, onsite client service delivery.

Recruiting for this role ends on Sep, 30th 2026.

Responsibilities
  • Architect, build, and operate scalable batch and near-real-time data pipelines on AWS.
  • Design robust ingestion patterns from source systems into S3 and into Snowflake.
  • Develop transformation layers and curated datasets in Snowflake, including dimensional/data product modeling for analytics and downstream applications.
  • Implement orchestration and workflow automation on AWS with retries, backfills, and idempotency.
  • Build reusable Python components for ingestion, validation, and transformations; enforce standards via code reviews and testing.
  • Optimize Snowflake performance and cost warehouse sizing, concurrency patterns, query tuning, clustering/micro‑partition considerations, and workload isolation.
  • Partner with stakeholders to translate requirements into well‑defined datasets and data contracts.
  • Communicate regularly with Engagement Managers (Directors), project team members, and representatives from various functional and / or technical teams, including escalating any matters that require additional attention and consideration from engagement management.
  • Independently and collaboratively lead client engagement workstreams focused on improvement, optimisation, and transformation of processes including implementing leading practice workflows, addressing deficits in quality, and driving operational outcomes.
Team

AI & Data – AI & Engineering leverages cutting‑edge engineering capabilities to build, deploy, and operate integrated/verticalized sector solutions in software, data, AI, network, and hybrid cloud infrastructure. These solutions are powered by engineering for business advantage, transforming mission‑critical operations. We enable clients to stay ahead with the latest advancements by transforming engineering teams and modernising technology & data platforms. Our delivery models are tailored to meet each client's unique requirements.

Qualifications
Required
  • 7+ years of experience as a Data Engineer delivering production‑grade data pipelines and curated datasets.
  • 7+ years of hands‑on experience with SQL and Python, including Snowflake and/or PySpark for scalable data processing and ELT.
  • 7+ years of experience designing, building, and operating batch and near‑real‑time data pipelines on cloud platforms (AWS preferred; Azure/GCP acceptable).
  • Experience with data integration frameworks and orchestration tools.
  • Proficiency in designing and implementing Lakehouse/warehouse architectures and ELT patterns.
  • Knowledge of DevOps principles: CI/CD pipelines, version control, Infrastructure‑as‑Code.
  • Ability to optimise data storage, partitioning, file formats (Delta, Parquet), and performance.
  • Understanding of data quality, data governance, and metadata management.
  • Bachelor's degree, preferably in Computer Science, Information Technology, Computer Engineering, or related IT discipline; or equivalent experience.
  • Limited immigration sponsorship may be available.
  • Ability to travel 10%, on average, based on the work you do and the clients and industries/sectors you serve.
Preferred
  • Agile delivery experience (5-10 years).
  • Analytical ability to manage multiple projects and prioritise tasks into manageable work products.
  • Can operate independently or with minimum supervision.
  • Excellent written and communication skills.
  • Ability to deliver technical demonstrations.
Wage Range

The wage range for this role takes into account the wide range of factors that are considered in making compensation decisions including but not limited to skill sets; experience and training; licensure and certifications; and other business and organisational needs. The disclosed range estimate has not been adjusted for the applicable geographic differential associated with the location at which the position may be filled. At Deloitte, it is not typical for an individual to be hired at or near the top of the range for their role and compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range is $102,750 to $171,250.

You may also be eligible to participate in a discretionary annual incentive programme, subject to the rules governing the programme, whereby an award, if any, depends on various factors, including, without limitation, individual and organisational performance.

Equal Opportunity Statement

All qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, sexual orientation, gender identity, national origin, age, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

PowerToFly • Alabama

On-site
USD 103,000 - 171,000
Genai Forward-Deployed Engineer: Data Infra Impact
Genai Forward-Deployed Engineer: Data Infra Impact

Scale Ai • San Francisco (CA)

On-site
USD 103,000 - 171,000
Data Engineer II
Data Engineer II

PowerToFly • Alabama

On-site
USD 80,000 - 128,000
Lead Data Engineer (secret clearance required)
Lead Data Engineer (secret clearance required)

PowerToFly • Maryland

On-site
USD 111,000 - 175,000
Lead Data Engineer II
Lead Data Engineer II

Relha LLC • Philadelphia

On-site
USD 130,000 - 180,000
Disability accommodations
Comprehensive benefits
Lead Agentic Software Engineer II
Lead Agentic Software Engineer II

PowerToFly • Alabama

On-site
USD 128,000 - 253,000
AI and Data Solution Architect
AI and Data Solution Architect

Relha LLC • Jacksonville (TX), Northern (KY)

On-site
USD 140,000 - 210,000
Benefits package
Generative AI Engineer III - State and Local Government
Generative AI Engineer III - State and Local Government

Latitude • Nashville (TN), Northern (KY)

Hybrid
USD 111,000 - 218,000
Discretionary annual incentive program
Senior Global Data ML Engineer – Multilingual Speech AI
Senior Global Data ML Engineer – Multilingual Speech AI

Cartesia • San Francisco (CA)

On-site
USD 103,000 - 171,000
AWS Forward Deployed Engineer - GPS
AWS Forward Deployed Engineer - GPS

Relha LLC • Northern (KY)

On-site
USD 135,000 - 265,000