Data Engineer

Azumo

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

PTO
US Holidays
AI Training
Mentored career development
Profit sharing
$US remuneration

Job summary

Azumo is seeking a senior Data Engineer to own ingestion, transformation, storage design, and the retrieval layer for production AI systems. You will build batch and streaming pipelines, manage data quality, and ensure cost-aware, scalable architecture across Snowflake, BigQuery, Redshift or Databricks, while collaborating across client environments.

Join a remote Latin America-wide team that leverages AI tools, maintains SOC 2/HIPAA-related standards, and delivers for high-profile clients like

Qualifications

  • 5+ years building and operating production data pipelines with Python and SQL
  • Deep expertise in designing data warehouses/lakehouses and dimensional modeling
  • Distributed processing at production scale with Spark, Kafka, Flink
  • Orchestration as engineering discipline: Airflow, Dagster, Prefect
  • Transformation under version control with dbt and lineage
  • Cloud deployment with Azure/AWS, Docker, CI/CD & IaC
  • Cost-aware pipeline design and throughput optimization
  • Experience with AI-assisted coding tools in real delivery work
  • Strong written and spoken English (C1+)
  • Bachelor's degree in CS/DS or equal experience

Responsibilities

  • Build ingestion and transformation pipelines with batch and streaming data
  • Design and optimize storage, warehouse and lakehouse models
  • Develop retrieval layers for AI systems (embedding/indexing)
  • Ensure data quality, lineage, tests and alerting
  • Implement data governance, privacy controls and audit trails
  • Deploy containers and manage CI/CD, orchestration and cost budgets
  • Collaborate in client environments within SOC 2/HIPAA contexts

Skills

Python
SQL
CI/CD
Testing
Git
Containers
Orchestration
Spark
Kafka
Flink
Airflow
Dagster
Prefect
dbt
Azure
AWS
Docker
GitHub Actions
Terraform
IaC
Cost optimization
AI coding tools
English proficiency
Bachelor's degree or equivalent

Education

Bachelor's degree in Computer Science, Data Science, or related field

Tools

dbt
Snowflake
BigQuery
Redshift
Databricks
Azure SQL DB

Job description

Azumo builds and operates production AI systems for companies ranging from seed-stage startups to Meta. We are hiring a Data Engineer to own the layer everything else depends on: ingestion and transformation pipelines, storage and warehouse design, and the retrieval infrastructure that AI systems query. The role is fully remote across Latin America, aligned to your client's working day.

Where this role sits

Azumo's engineering organization is built around four lanes. The Data Scientist lane owns the question and the method. The AI Engineer lane owns production behavior. The Software Engineer lane owns AI-augmented product delivery. This role is the Data Engineer lane, and it owns pipelines, storage, and the retrieval layer.

One question places the boundary: when the output is wrong, whose problem is it? "The data was missing, stale, or wrong by the time it arrived" is yours. "The system did the wrong thing with data that was correct" is the AI Engineer's.

Not quite your profile? Check our other openings:
  • Data Scientist
  • AI Engineer
  • AI-Augmented Software Engineer
  • Forward Deployed Engineer
What you will build
  • Ingestion and transformation pipelines. Batch and streaming ingestion on Spark, Kafka, dbt and Airflow, with idempotency, backfills, schema evolution and late-arriving data handled by design rather than by hand.
  • Storage and modeling. Warehouse and lakehouse design on Snowflake, BigQuery, Redshift or Databricks, with partitioning, file layout and query cost treated as engineering decisions.
  • The retrieval layer. The chunking, embedding and indexing pipelines that feed RAG systems on pgvector, Pinecone, Qdrant or Azure AI Search, and the freshness, deduplication and permission problems that come with them.
  • Data quality as a contract. Tests, expectations, lineage and alerting. If a pipeline is wrong, the people downstream should hear it from you and not from the client.
  • Sensitive data by default. PII classification, masking, row and column level access, retention and deletion, and an audit trail that holds up when a client asks who read what.
  • Production operation. Containerized deployment on Azure or AWS, CI/CD, orchestration, observability, and explicit cost and runtime budgets that you own rather than discover after the invoice.
  • Work inside the client's environment. Their repositories, their standups, sometimes their customer calls. Azumo is SOC 2 certified, client code stays in client repositories, and some engagements carry additional requirements such as HIPAA.
How we work

Our engineers build with AI every day. Claude Code, Codex, and similar tools are part of the standard toolchain here, not an experiment. We run an automated audit across the whole codebase on day one and every day after, grading security, cost, and architecture findings by severity with the exact file and line, so a small team can move quickly without quality drifting. We stay vendor-neutral across OpenAI, Anthropic, and open-weight models, and we run Valkyrie, our own production layer, when a single interface to any model is the right call.

About Azumo

Azumo is a San Francisco based software development company that has been building intelligent applications since 2016. We provide nearshore AI engineering teams to organizations that need production AI faster than they can hire for it: as an embedded engineering team, as AI staff augmentation alongside an existing team, or as a full project build. Our engineers work from Latin America, aligned to United States time zones, and have delivered for Twitter, Meta, Discovery Channel, Omnicom, UnitedHealth, and CENTEGIX.

We hire for seniority and test for it before anyone joins a client team. We support engineers in going deep on the modern AI stack, and we give time back to open-source work, community teaching, and philanthropy.

Basic qualifications
  • 5+ years building and operating production data pipelines, with Python and SQL as your primary languages, plus the engineering fundamentals that go with it: testing, code review, CI/CD, Git, containers, and orchestration.
  • Deep expertise in designing and building data warehouses or lakehouses, including dimensional modeling, incremental processing, and the cost and performance trade-offs behind each choice.
  • Distributed processing at production scale with Spark, Kafka, Flink or equivalent, including the failure modes that only appear under load.
  • Orchestration as an engineering discipline rather than a cron replacement: Airflow, Dagster or Prefect, with retries, idempotency and backfill strategy you can defend.
  • Transformation under version control, with tests and lineage: dbt or something you built yourself.
  • Cloud deployment experience, Azure preferred and AWS acceptable, with Docker, CI/CD pipelines, and infrastructure as code (GitHub Actions, Terraform, or Bicep).
  • Working discipline around pipeline cost, runtime and throughput. You can explain what a pipeline costs to run and what you did about it.
  • Active use of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in real delivery work.
  • Clear written and spoken English, C1 or above, and the confidence to explain a technical trade-off directly to a client.
  • Bachelor's degree in Computer Science, Data Science, or a related field, or equivalent professional experience.
Preferred qualifications
  • Vector and retrieval infrastructure: pgvector, Pinecone, Qdrant, FAISS or Azure AI Search, and the retrieval-quality problems that come with it.
  • Streaming, real-time or high-throughput workloads.
  • Experience with cloud-based managed services like Airflow, Glue, Elastic stack, Amazon Redshift, Snowflake, BigQuery, Azure SQL Db, EMR, Databricks.
  • Prior experience with notebooks using Jupyter, Google Collab, or similar.
  • Delivery under a compliance regime such as SOC 2 or HIPAA.
  • Contributions to open-source data libraries, published technical writing, or active participation in the data engineering community.
  • Paid time off (PTO)
  • U.S. Holidays
  • AI Training
  • Mentored career development
  • Profit sharing
  • $US remuneration
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI-Augmented Software Engineer
AI-Augmented Software Engineer

Azumo • United States

Remote
USD 120,000 - 180,000
Paid time off (PTO)
U.S. Holidays
AI Training
+3
Forward Deployed Engineer
Forward Deployed Engineer

Azumo • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
PTO
US Holidays
AI Training
+3
Senior Security Engineer
Senior Security Engineer

Azumo LLC. • Northern (KY)

Hybrid
USD 120,000 - 180,000
Remote-first culture in LATAM
Paid time off (PTO)
U.S. Holidays
+5
Data Engineer Remote Latin America
Data Engineer Remote Latin America

Fractal River • United States

Remote
USD 70,000 - 120,000
Personal development plan
Access to a reference library
Unlimited access to AI tools
+3
Head of Delivery - Latin America - Remote
Head of Delivery - Latin America - Remote

Azumo • United States

Remote
USD 140,000 - 200,000
Paid Time Off
Mentored Career Development
U.S. Holidays
+4
Sr Data Engineer, AI
Sr Data Engineer, AI

Constellation • De Pere (WI)

On-site
USD 146,000 - 162,000
Marketing AI Intern
Marketing AI Intern

Azul Systems, Inc. • Northern (KY)

On-site
USD 28,000 - 34,000
Comprehensive compensation and health"
Referral Program
Remote-first, PTO, holidays
Data Engineer (UA/RU Language speaking)
Data Engineer (UA/RU Language speaking)

Neurons Lab • Spain (TX)

On-site
USD 60,000 - 120,000
Sr Data Engineer, AI
Sr Data Engineer, AI

Constellation • Baltimore (MD)

On-site
USD 146,000 - 162,000
Bonus program
401(k) with company match
Employee stock purchase program
+4
AI Data Architect REQ_22
AI Data Architect REQ_22

3Pillar • United States

On-site
USD 140,000 - 190,000
Medical Insurance
Dental insurance
Vision insurance
+6