Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title And Summary
Senior Data Engineer
Role: Senior Data Engineer
Role Overview
We are seeking a Senior Data Engineer to drive AI‑led transformation across our analytics and data platforms. This role is designed for a hands‑on engineer who can ideate, design, and implement AI‑driven solutions, modernize Python and data pipelines, and maintain best‑in‑class analytics environments.
The role combines agentic AI concepts, cloud‑based AI services, strong Python/PySpark engineering, and deep analytics expertise across Databricks and Snowflake, while applying solid ETL and DBA fundamentals in AWS.
Skill Priority & Core Responsibilities
- AI – Agentic Systems, GenAI & Intelligent Workflow Automation
- Lead AI ideation by identifying opportunities to eliminate manual effort, reduce operational friction, and improve decision‑making.
- Design And Implement Agentic AI Solutions, Including
- Multi‑agent orchestration for task delegation and workflow execution
- AI agents for monitoring, diagnostics, reconciliation, and data reasoning
- Build end‑to‑end GenAI / LLM‑powered workflows, including prompt engineering, chaining, tool use, and evaluation.
- Design and maintain data pipelines specifically for AI workloads, supporting:
- Training, fine‑tuning, and inference
- Vector storage, retrieval‑augmented generation (RAG), and embeddings
- Integrate AI solutions with existing data platforms while ensuring governance, observability, and cost control.
- Apply responsible AI principles, access controls, and auditability for AI‑driven systems.
- Python Engineering, APIs & PySpark Development
- Develop production‑grade Python applications supporting data, AI, and automation use cases.
- Design and expose REST APIs using Flask or FastAPI for model inference, AI services, and data access.
- Improve and refactor existing Python codebases for maintainability, performance, and scalability.
- Migrate legacy logic and scripts into PySpark‑based implementations on Databricks.
- Follow Best Practices In
- Modular architecture and design patterns
- Logging, monitoring, and exception handling
- Unit and integration testing
- Analytics Platforms – Databricks & Snowflake Expertise
- Act as an expert practitioner for Databricks and Snowflake, supporting both analytics and AI‑driven workloads.
- Perform Advanced Problem Solving Related To
- Performance tuning
- Cost optimization
- Query optimization and data layout
- Define and enforce standards and best practices for analytics and AI workloads on these platforms.
- Support Administration And Daily Management Activities, Including
- Workspace and resource governance
- User access and role management
- Platform usage optimization
- Enable analytics teams through reusable patterns, templates, and documentation.
- ETL & Cloud Data Engineering (AWS)
- Architect and implement scalable ETL pipelines using modern cloud‑native patterns.
- Apply expert knowledge of ETL frameworks and design principles, including incremental processing and fault tolerance.
- Build and maintain pipelines using AWS Glue, interacting with S3, IAM, and related AWS services.
- Ensure reliable orchestration, monitoring, and error handling across data pipelines.
- Optimize ETL workloads for performance, scalability, and cost efficiency.
- DBA Concepts & Data Storage Expertise
- Apply strong DBA fundamentals to analytics and operational systems.
- Work With Relational And NoSQL Databases, Including
- Amazon RDS (MySQL, PostgreSQL)
- MongoDB Atlas
- Apply Best Practices In
- Indexing and query optimization
- Backup, recovery, and high availability
- Capacity planning and performance troubleshooting
- Collaborate with infrastructure and database teams on design and operational improvements.
- Cloud-Based AI Services & Platforms (Cross‑Cutting Skill Area)
- Hands‑on experience or strong familiarity with cloud AI and compute services, including:
- AWS Bedrock for foundation models and GenAI applications
- AWS SageMaker for training, experimentation, and deployment
- AWS Lambda for event‑driven AI and data automation
- Databricks AI for ML/GenAI workloads and platform integration
- Microsoft Copilot and related AI tooling for productivity enablement and integration scenarios
- Evaluate, integrate, and operationalize AI services based on cost, security, scalability, and use‑case fit.
- Must‑Have Skills
- 6+ years of experience in Data Engineering / Analytics Engineering roles.
- Strong hands‑on experience with GenAI, LLMs, agentic AI, and multi‑agent orchestration.
- Advanced Python development experience, including Flask or FastAPI.
- Strong experience in PySpark and distributed data processing.
- Expert‑level proficiency with Databricks and Snowflake.
- Deep understanding of ETL architecture, AWS Glue, and cloud‑native data pipelines.
- Good‑to‑Have Skills
- Experience with vector databases and RAG architectures.
- Exposure to ML lifecycle management, MLOps, and experiment tracking.
- Infrastructure-as-Code (Terraform or similar).
- Working knowledge of DBA concepts and databases on AWS.
- Experience working in regulated domains (banking, fintech, healthcare).
- CI/CD practices for data, AI, and analytics pipelines.
Corporate Security Responsibility
- Abide by Mastercard’s security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.