Data Engineer, Amazon Traffic Engineering

Amazon

Vancouver

On-site

CAD 110,000 - 170,000

Full time

7 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Amazon is seeking a Data Engineer to build and operate the Core Data Infrastructure that underpins ML and science initiatives for Bot Management. You will design and own production-grade data pipelines ingesting billions of events from diverse sources and transform raw signals into ML-ready feature groups used by Science and ML Platform teams.

You will partner with Applied Scientists and across data engineering, software engineering, and science teams to deliver accurate, timely, and

Qualifications

  • 3+ years of data engineering experience.
  • Experience with data modeling, warehousing and ETL pipelines.
  • Experience with large-scale, high-throughput, 24x7 data systems.
  • Experience in Python/Java/Scala/NodeJS.

Responsibilities

  • Build and operate batch and near real-time data pipelines.
  • Ingest billions of events from diverse sources and transform into ML-ready features.
  • Collaborate with scientists to translate model data requirements into reliable pipelines.
  • Ensure data quality, governance, and reliability of pipelines.

Skills

Data Engineering
ETL Pipelines
Batch & Streaming
Python/Java/Scala/NodeJS
Data Modeling
Warehousing
SDLC & Best Practices

Tools

AWS Glue
Redshift
S3
Kinesis
Kafka
OpenSearch

Job description

Data Engineer, Amazon Traffic Engineering

We are seeking an experienced Data Engineer to build and operate the Core Data Infrastructure that underpins our ML and Science initiatives for Bot Management. You will design and own production-grade data pipelines that ingest billions of events from diverse source systems and transform raw signals into ML-ready feature groups that Science and ML Platform teams depend on for training, evaluation, and inference.


This is a hands-on engineering role at the center of a fast-moving ML organization. Our Science teams build increasingly sophisticated models, each requiring different data formats, latencies, and serving patterns. You will build the pipelines and feature infrastructure that make this possible: ingesting from Trails and Non-Trails sources, transforming disparate datasets into versioned feature groups, and operating the Feature Store that serves features consistently across all model types. You will partner directly with Applied Scientists to translate model data requirements into reliable pipelines, and work across data engineering, software engineering, and science teams to deliver data that is accurate, timely, and well-governed.


Data Pipelines & Ingestion — Build and own batch and near real-time pipelines spanning Trails (raw and aggregated) and Non-Trails sources (Clickstream, Customer Segmentations, OPS). Evolve pipelines from Cradle/POC to production-grade using AWS Glue. Implement data quality checks, drift detection, and governance frameworks.


Feature Engineering & Serving — Build versioned feature groups across multiple storage backends (S3 for tabular data, OpenSearch for embeddings). Develop production pipelines that transform raw signals into ML-ready features, and help operate the Feature Store that serves them consistently to Science teams.


Streaming & Real-Time Systems — Develop and operate Apache Flink applications and stream processing for near real-time feature computation. Build event-driven data flows leveraging Kinesis and Kafka to support low-latency bot detection signals.


Science Partnership — Partner with Applied Scientists and ML Platform engineers to define data contracts and SLAs, understand model data requirements, and ensure feature pipelines integrate cleanly with training and inference systems.


Operational Excellence — Own the reliability, monitoring, and cost efficiency of the pipelines you build. Participate in on-call, root-cause data issues, and drive improvements that reduce operational load.


About the team Traffic Engineering's Bot Management organization protects Amazon's ecosystem by detecting and mitigating automated threats at scale. Our Core ML Data Infrastructure team is responsible for building and operating the foundational data infrastructure that powers bot detection, AI agent identification, and content exfiltration defense. We are building a unified, model-agnostic, production-grade ML platform that brings together training, evaluation, and inference pipelines into a cohesive system serving multiple model types across the organization.


Basic Qualifications


  • 3+ years of data engineering experience

  • Experience with data modeling, warehousing and building ETL pipelines

  • Experience building large-scale, high-throughput, 24x7 data systems

  • Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS

  • Knowledge of batch and streaming data architectures like Kafka, Kinesis, Flink, Storm, Beam

  • Knowledge of professional software engineering & best practices for full software development life cycle, including coding standards, software architectures, code reviews, source control management, continuous deployments, testing, and operational excellence


Preferred Qualifications


  • Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions

  • Experience with non-relational databases / data stores (object storage, document or key-value stores, graph databases, column-family databases)


Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.


Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.


The base salary range for this position is listed below. As a total compensation company, Amazon's package may include other elements such as sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon offers comprehensive benefits including health insurance (medical, dental, vision, prescription, basic life & AD&D insurance), Registered Retirement Savings Plan (RRSP), Deferred Profit Sharing Plan (DPSP), paid time off, and other resources to improve health and well-being. We thank all applicants for their interest, however only those interviewed will be advised as to hiring status.


Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Engineer, WW Pricing Data and Insights
Sr. Data Engineer, WW Pricing Data and Insights

Amazon • Vancouver

On-site
CAD 120,000 - 180,000
Health insurance
RRSP
DPSP
+2
Machine Learning Engineer , Amazon Customer Service
Machine Learning Engineer , Amazon Customer Service

Amazon • Vancouver

On-site
CAD 120,000 - 170,000
Health insurance
RRSP
DPSP
+1
Software Development Manager, Amazon Customer Service
Software Development Manager, Amazon Customer Service

Amazon • Vancouver

On-site
CAD 171,000 - 286,000
Sr. Applied Scientist, Workforce Solutions
Sr. Applied Scientist, Workforce Solutions

Amazon • Vancouver

On-site
CAD 140,000 - 190,000
Sr. Applied Scientist, Workforce Solutions
Sr. Applied Scientist, Workforce Solutions

Amazon Science • Vancouver

On-site
CAD 196,000 - 327,000
Data Engineer II, Alexa Audio
Data Engineer II, Alexa Audio

Socket.dev • Vancouver

On-site
CAD 103,000 - 173,000
Health insurance
RRSP
DPSP
+2
Principal Worldwide Specialist - GenAI, Data Strategy for AI
Principal Worldwide Specialist - GenAI, Data Strategy for AI

Amazon • Vancouver

On-site
CAD 193,000 - 324,000
Medical Benefits
Financial Benefits
Equity Options
+2
Sr. Applied Scientist, Workforce Solutions
Sr. Applied Scientist, Workforce Solutions

Socket.dev • Vancouver

On-site
CAD 196,000 - 327,000
Health insurance (medical, dental, and
RRSP
DPSP
+1
Applied Scientist, Tax Engine
Applied Scientist, Tax Engine

Amazon • Vancouver

On-site
CAD 120,000 - 180,000
Health insurance
RRSP
Paid time off
Sr. Data Scientist, Alexa Connections
Sr. Data Scientist, Alexa Connections

Amazon • Vancouver

On-site
CAD 143,000 - 239,000