Title: Data Engineer II
Location: Arlington, VA (Onsite)
Position Type: W2 contract
Industry: Banking/Financials/Payments
Required Education: Bachelor’s degree in a quantitative discipline such as Engineering, Mathematics, Finance, Business, or a related field. Equivalent practical experience may also be considered.
Qualifications & Skills
- 9+ years of hands‑on data engineering experience building large‑scale ETL pipelines using Apache Spark, Hadoop, Python, and SQL.
- Strong background in the payments or financial sector.
- Proficiency in writing and optimizing SQL queries to retrieve, manipulate, and analyze data efficiently.
- Hands‑on experience with Apache Spark (PySpark, Spark SQL, Spark Streaming) and the Hadoop ecosystem (HDFS/Ozone, Hive, YARN).
- Understanding data modeling concepts and database design to support scalable data solutions.
- Familiarity with Python.
- Ability to analyze and troubleshoot data issues and provide solutions with minimal supervision.
- Basic knowledge of testing and validating data to ensure accuracy and consistency in data pipelines.
- Excellent verbal and written communication skills, with the ability to articulate complex ideas clearly to both technical and non‑technical stakeholders.
Responsibilities
- Design, implement, and maintain scalable enterprise ETL processes and robust data pipelines for a global client base.
- Develop and optimize code for data processing, ensuring timely availability and accessibility of data.
- Leverage big data processing frameworks such as Apache Spark and Hadoop to build and optimize data pipelines.
- Collaborate with senior engineers to address data challenges and maintain high data quality.
- Assist in data delivery, working alongside Data Engineers and Analysts to support accurate, high‑value data solutions across various clients and industries.
- Build strong working relationships with team members and clients, contributing to local and global projects.
- Apply industry best practices, including version control, code reviews, and data validation, to ensure quality in data processes.
- Use SQL and other database technologies to optimize data processing and reduce the time required to handle large data sets.
- Design, implement, and maintain data pipelines using ETL frameworks, orchestration tools, and distributed data processing engines.
- Participate in automation of routine data tasks and streamlining processes.
- Comply with all Mastercard internal policies and adhere to external regulations.