Data Engineer

The Carlyle Group

New York (NY)

On-site

USD 150,000 - 200,000

Full time

8 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive benefits package
Discretionary incentive program

Job summary

The Carlyle Group seeks a Data Engineer to build and maintain robust data pipelines from diverse vendors to gold tables for the ML team, using Snowflake and Databricks. The role requires Python expertise, strong financial data knowledge, and a proactive approach to data aggregation and enrichment from private and public sources.

Responsibilities include end-to-end data solutions, scalable ETL/ELT pipelines, and collaboration with stakeholders to productionize services.

Qualifications

  • Excellent communication skills required.

Responsibilities

  • Build, scale, and maintain robust data solutions to support the firm's objectives.
  • Implement and optimize high-performance data pipelines designed for scalability, reliability, maintainability, and speed.
  • Lead software development projects end to end involving large language models (LLMs), retrieval-augmented generation (RAG) frameworks, and other AI technologies.
  • Champion modern software engineering practices such as CI/CD, infrastructure-as-code, containerization, and cloud-native deployments.
  • Collaborate with business stakeholders to transform use cases into production-ready services and solutions, owning the system from concept to production.
  • Implement rigorous testing and monitoring to maintain data quality and integrity.
  • Mentor and develop junior team members, fostering a culture of excellence and continuous learning.

Skills

Communication skills

Education

CS/Math/Physics/STEM concentration

Tools

Python
SQL
Databricks/PySpark
dbt
Airflow
Temporal
LLMs/RAG
Vector databases

Job description

Position Summary

We are seeking a highly skilled Data Engineer to join our dynamic team. The ideal candidate will be responsible for creating robust data pipelines from various data vendors to gold tables, primarily for our Machine Learning (ML) team, utilizing Snowflake and Databricks platforms. The role demands expertise in Python, deep familiarity with financial data sources, and the ability to deploy complex data pipelines efficiently. This position requires a proactive approach to analyzing, aggregating, and enriching financial data from both private and public companies.

Primary Responsibilities
  • Build, scale, and maintain robust data solutions to support the firm's objectives.
  • Implement and optimize high-performance data pipelines -- extraction, loading, transformation, and orchestration - that are designed for scalability, reliability, maintainability, and speed.
  • Lead software development projects end to end involving large language models (LLMs), retrieval-augmented generation (RAG) frameworks, and other AI technologies.
  • Champion modern software engineering practices as CI/CD, infrastructure-as-code, containerization, and cloud-native deployments
  • Collaborate closely with business stakeholders to transform use cases into production-ready services and solutions, owning the system from concept to production.
  • Implement rigorous testing and monitoring practices to maintain superior data quality and integrity.
  • Mentor and develop junior team members, fostering a culture of excellence and continuous learning within the team.
  • Be willing to travel up to 20% of the time to collaborate with distributed team members across locations.
Requirements
Education & Certificates
  • Concentration in Computer Science, Math, Physics, STEM or other engineering related field, preferred
Professional Experience
  • At least 6 years of experience in data engineering or a related discipline, with a proven track record of success, required
  • Experience in the financial services or private equity industry, preferred
  • Expertise in Python and SQL, with a strong foundation in data manipulation and analysis.
  • Proficient with Databricks/PySpark and dbt for data warehousing and data transformation tasks.
  • Experience with workflow orchestration tools e.g. Airflow, Temporal
  • Experience working with large language models (LLMs) especially prompt engineering, retrieval-augmented generation (RAG)s, and/or vector databases.
  • Knowledge of fundamental principles of machine learning, feature engineering, and knowledge graphs are pluses.
  • Demonstrated experience in designing and implementing complex data systems from the ground up.
  • Proficient in handling large-scale data projects, including data cleaning, ETL, and information retrieval.
  • Previous experience in a product development or financial services environment is highly desirable.
  • Excellent communication skills required, both verbal and written.
Benefits/Compensation

The compensation range for this role is specific to New York, NY and takes into account a wide range of factors including but not limited to the skill sets required/preferred; prior experience and training; licenses and/or certifications.
The anticipated base salary range for this role is $150,000 to $200,000.
In addition to the base salary, the hired professional will enjoy a comprehensive benefits package spanning retirement benefits, health insurance, life insurance and disability, paid time off, paid holidays, family planning benefits and various wellness programs. Additionally, the hired professional may also be eligible to participate in an annual discretionary incentive program, the award of which will be dependent on various factors, including, without limitation, individual and organizational performance.
Due to the high volume of candidates, please be advised that only candidates selected to interview will be contacted by Carlyle.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Atlas Search • New York (NY)

Hybrid
USD 140,000 - 165,000
Data Engineer, Global Credit Technology
Data Engineer, Global Credit Technology

1P284 THE CARLYLE GROUP EMPLOYEE CO., LLC • United States

On-site
USD 150,000 - 170,000
Data Analytics Engineer
Data Analytics Engineer

EXL • San Francisco (CA)

On-site
USD 140,000 - 160,000
Data Engineer II
Data Engineer II

DriveWealth • New York (NY)

Hybrid
USD 145,000 - 165,000
Data Engineer
Data Engineer

Solomon Page • New York (NY)

Hybrid
USD 175,000 - 240,000
Senior Data Engineer
Senior Data Engineer

Syndesus, Inc. • New York (NY)

On-site
USD 170,000 - 190,000
Sr. Data Engineer
Sr. Data Engineer

Kinect • New York (NY)

On-site
USD 150,000 - 210,000
Data Engineer
Data Engineer

Henderson Scott US • New York (NY)

Hybrid
USD 80,000 - 100,000
Data Engineer - Finance
Data Engineer - Finance

Weekday AI (YC W21) • San Francisco (CA)

On-site
Data Engineer
Data Engineer

Greenthumbindustries • Chicago (IL)

Hybrid
USD 95,000 - 115,000