Data Engineer

mizuho

United States

Hybrid

USD 88,000 - 105,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
401K plan
Discretionary bonus

Job summary

Mizuho is seeking a Data Engineer to design and develop enterprise data products on the Databricks Lakehouse. You will ingest, transform, and validate large data sets, ensuring accuracy and scalability. You will collaborate with global teams and participate in code reviews while learning the Medallion architecture.

The role emphasizes hands-on development with SQL, Python, and PySpark, in a hybrid environment with remote options and opportunities across departmental needs.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field.
  • Proven data engineering or software development experience.
  • Proficiency in Python, SQL, and data modeling.

Responsibilities

  • Assist in building and maintaining ingestion pipelines landing raw data into Bronze layer.
  • Support Silver layer transformations: cleansing, deduplication, schema enforcement.
  • Write SQL and PySpark for defined transformations.

Skills

SQL
Python
Git
Databricks
Spark

Education

Bachelor's degree in Computer Science or related
Master's degree preferred

Tools

Databricks
Azure Data Factory
Kafka
Spark

Job description

Join Mizuho as a Data Engineer!

The IT Data team is responsible for design and development of Data products for the entire firm. The team is embarking on an ambitious new project/implementation. "DEAL". It is the abbreviation for "Data Exchange and Abstraction Layer". It's the next generation, Data Mesh based platform implemented in the Mizuho Azure Cloud on Databricks. Data Mesh is a decentralized data architecture where data is owned and managed by the domain-specific teams that produce the data i.e. Banking, Finance etc., and curate it for downstream consumption. It emphasizes domain-oriented ownership, treating data as a product, providing a self-serve data platform, and using limited federated computational governance from the Data Architecture and Data Management Office.

In this role you will be responsible for development of data solutions for the enterprise using innovative and cutting- edge technologies like AI tools / models. The solutions and software developed will be used for reporting and analytics by the entire firm globally. The data solutions developed will have to be accurate, timely and highly scalable. In this role you will be managing large data-sets with complex interdependencies. This is a hands-on software development role. You will be collaborating with teams firmwide to develop solutions.

You’ll support the development and maintenance of data pipelines on the Databricks Lakehouse platform using the medallion architecture (Bronze/Silver/Gold). You will work under the guidance of senior engineers to ingest, transform, and validate data, growing your skills across the modern data stack.

Key Responsibilities
  • Assist in building and maintaining ingestion pipelines that land raw data into the Bronze layer.
  • Support Silver layer transformations under guidance: cleansing, deduplication, and schema enforcement.
  • Write SQL and PySpark for defined transformation tasks.
  • Run and monitor scheduled jobs; help investigate and resolve pipeline failures.
  • Document pipeline logic, transformations, and fixes.
  • Participate in code reviews as a reviewer-in-training and incorporate feedback on your own work.
  • Learn team standards for version control, testing, and deployment.
Required Qualifications
  • 0-2 years of experience in data engineering, analytics, or a related technical role (internships and academic projects count).
  • Foundational SQL skills (joins, aggregations, filtering).
  • Python proficiency with basic OOPS knowledge .
  • Understanding of core data concepts (tables, schemas, relational data).
  • Willingness to learn Databricks, Spark, and cloud technologies.
  • Familiarity with Git or a demonstrated ability to learn version control quickly.
Preferred / Nice-to-Have
  • Exposure to Databricks, Apache Spark, or PySpark (coursework or hands-on).
  • Awareness of the medallion architecture and Delta Lake basics.
  • Experience with any cloud platform (Azure, AWS, or GCP).
  • Relevant coursework, bootcamp, or a Databricks certification (e.g., Data Engineer Associate).
  • Any experience with data visualization or BI tools.
Competencies
  • Eagerness to learn and take feedback.
  • Attention to detail and care for data accuracy.
  • Basic problem-solving and logical thinking.
  • Communicate issues across teams
Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field. Master's degree preferred.
  • Proven experience in data engineering, software development, or related roles.
  • Proficiency in programming languages commonly used in data engineering (e.g., Python, Scala, etc.).
  • Strong knowledge of database systems, data modeling techniques, and SQL proficiency.
  • Proficiency with ETL tools commonly used in data engineering (e.g., SSIS, Databricks, Azure Data Factory).
  • Experience with big data technologies and frameworks (e.g., Spark, Kafka, etc.).
  • Familiarity with cloud platforms and services (e.g., Azure).
  • Excellent problem-solving skills and attention to detail.
  • Effective communication and collaboration skills in a team-oriented environment.
  • Ability to adapt to evolving technologies and business requirements

The expected base salary ranges from $88,000 - $105,000. Salary offers are based on a wide range of factors including relevant skills, training, experience, education, and, where applicable, certifications and licenses obtained. Market and organizational factors are also considered. In addition to salary and a generous employee benefits package, including Medical, Dental and 401K plans, successful candidates are also eligible to receive a discretionary bonus.

Other requirements

Mizuho has in place a hybrid working program, with varying opportunities for remote work depending on the nature of the role, needs of your department, as well as local laws and regulatory obligations. Roles in some of our departments have greater in-office requirements that will be communicated to you as part of the recruitment process.

Company Overview

Mizuho Financial Group, Inc. is the 15th largest bank in the world as m

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Mizuho • Woodbridge Township (NJ)

Hybrid
USD 88,000 - 105,000
Medical benefit package
401K plan
Data Engineer
Data Engineer

Mizuho Financial Group Inc. • United States

Hybrid
USD 88,000 - 105,000
Medical
Dental
401K plans
+1
Cloud Data Engineer
Cloud Data Engineer

Mizuho Financial Group Inc. • United States

Hybrid
USD 160,000 - 200,000
Hybrid working program
Discretionary bonus
Cloud Data Engineer
Cloud Data Engineer

Mizuho • Woodbridge Township (NJ)

Hybrid
USD 160,000 - 200,000
Hybrid work program
Data Architecture & Engineering — Market Risk Technology
Data Architecture & Engineering — Market Risk Technology

Mizuho • New York (NY)

Hybrid
USD 150,000 - 240,000
Discretionary bonus
Hybrid work program
Data Integration Operations Manager
Data Integration Operations Manager

Mizuho Financial Group Inc. • United States

Hybrid
USD 111,000 - 170,000
Medical insurance
Dental insurance
401K plan
+1
Data Integration Operations Manager
Data Integration Operations Manager

Mizuho • Woodbridge Township (NJ)

Hybrid
USD 111,000 - 170,000
Medical
Dental
401K
+1
Data Architecture & Engineering — Market Risk Technology
Data Architecture & Engineering — Market Risk Technology

Mizuho Financial Group Inc. • New York (NY)

Hybrid
USD 150,000 - 240,000
Base salary range
Discretionary bonus
Comprehensive benefits
Data Engineer - Databricks & Data Mesh (Hybrid)
Data Engineer - Databricks & Data Mesh (Hybrid)

Mizuho Financial Group Inc. • United States

Hybrid
USD 88,000 - 105,000
Medical
Dental
401K plans
+1
Data Engineer: Databricks & Data Mesh (Hybrid)
Data Engineer: Databricks & Data Mesh (Hybrid)

mizuho • United States

Hybrid
USD 88,000 - 105,000
Medical insurance
401K plan
Discretionary bonus