Data Engineer Lead

Capital Group Companies

Los Angeles (CA)

On-site

USD 202,000 - 342,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Generous time-away
Company-funded retirement contribution
Professional development resources
Annual performance bonus
Bonuses

Job summary

Capital Group Companies in Los Angeles is seeking a senior hands-on Data Engineering Lead to set the technical direction for the CSGT data platform and data products. You will own the data engineering strategy, design robust pipelines, and lead cross-team initiatives using Databricks on AWS, PySpark, and dbt.

The role emphasizes AI-first engineering, data governance, and cost-aware scalability, with responsibilities spanning from standards development to production support and incident response.

Qualifications

  • 10+ years of experience in data or software engineering with leadership of complex production data platforms

Responsibilities

  • Own data engineering strategy and roadmap for CSGT data platform and data products
  • Shape standards across teams and Capital Group engineering practice
  • Design and build ingestion, transformation, and serving pipelines using Databricks, PySpark, Delta Lake, dbt, and Airflow
  • Lead cross-team data initiatives from requirements to production support, including estimates, sequencing, dependencies, and cost
  • Advance AI-first engineering by translating business outcomes for agents and AI applications
  • Develop reusable agent workflows for data engineering tasks and establish governance around agent actions

Skills

Python
SQL
Distributed data processing
Data modeling
Data governance
Airflow
Databricks

Education

Bachelor's degree in Computer Science or related field

Tools

Databricks
AWS
PySpark
Delta Lake
Unity Catalog
dbt
Airflow
Astronomer
Datadog
Terraform
Harness
Databricks Genie

Job description

Capital Group Companies is hiring a senior, hands-on Data Engineering Lead to set technical direction for the CSGT data platform and data products in Los Angeles.

Responsibilities
  • Own data engineering strategy and roadmap for CSGT, including Lakehouse architecture on Databricks and AWS; make and explain decisions on scalability, security, reliability, and cost
  • Shape standards and practices across adjacent teams and the broader Capital Group technology organization, contributing to centers of excellence and firm-wide engineering practices
  • Design and build ingestion, transformation, and serving pipelines using Databricks, PySpark, Delta Lake, dbt, and Airflow; establish reusable patterns and frameworks for consistent, maintainable data products
  • Evaluate new structured and unstructured datasets at the business-capability level, and integrate them into platform data domains and subject areas
  • Lead complex, cross-team data initiatives from requirements through production support, including estimates, work breakdown, sequencing, dependencies, and cost; surface risks early while balancing durable platform needs with near-term business priorities
  • Advance AI-first engineering by translating business outcomes into specifications, engineering business and architectural context for agents, and directing agents to plan, build, test, and document changes in small, reviewable increments
  • Develop reusable agent workflows and skills for data engineering work including profiling, source-to-target mapping, pipeline and test generation, schema-change analysis, and incident investigation
  • Define which agent actions can be executed autonomously versus require human approval, and specify how agent activity is reviewed and traced
  • Improve AI-readiness by delivering governed, understandable data through Unity Catalog metadata, lineage, business definitions, semantic models, and access controls; connect datasets to Databricks Genie and other AI applications used by investment professionals
  • Define evaluation datasets and acceptance criteria for agents and AI-generated SQL and code
  • Set the testing strategy across platform layers (including performance, stability, and availability) and review/approve quality metrics before release
  • Build data quality, reconciliation, freshness, observability, and recovery controls into automated testing and CI/CD, with security and policy checks embedded from the start
  • Partner with investment professionals and product managers to drive shared product vision and ownership of business outcomes, demonstrating how data and AI support portfolio construction and research at scale
  • Raise the engineering bar through design and code reviews, direct day-to-day engineering execution for initiatives, and guide through the most complex data and performance challenges
  • Teach engineers to inspect and challenge AI-generated work; share reusable patterns and context through internal and external forums; help managers identify strengths and development needs
Requirements
  • 10+ years of experience in data or software engineering, including technical leadership of complex production data platforms delivered across multiple teams
  • Strong hands-on Python and SQL skills, sound software design judgment, and deep understanding of distributed data processing, query performance, and automated testing
  • Production experience with Databricks on AWS, including PySpark, Delta Lake, Unity Catalog, Databricks Jobs, Databricks SQL, and Databricks Asset Bundles, including security, access, and cost implications
  • Experience orchestrating production pipelines with Apache Airflow (including Astronomer) and building tested transformations with dbt, with reliable retries, backfills, and dependency management
  • Strong data modeling and governance experience, including dimensional and time-series models, slowly changing dimensions, bi-temporal history, data contracts, lineage, and semantic metadata
  • Implemented data quality, observability, and CI/CD for data platforms using tools such as Deequ, dbt tests, Lakehouse Monitoring, Datadog, Terraform, and Harness
  • Experience preparing governed data for AI using natural-language-to-SQL tools such as Databricks Genie, semantic metadata, or other governed data-access patterns
  • Use AI coding agents beyond code completion: write specifications, supply context, run tests, and review generated changes via source control
  • Evaluate AI-generated output with representative test cases, regression tests, execution traces, and human review; distinguish plausible output from verified results
  • Understand prompt injection, sensitive data handling, and least-privilege access; design approval boundaries and audit trails for agents operating against enterprise systems
  • Lead architecture discussions, influence without formal authority, develop other engineers, and clearly explain technical choices and trade-offs to investment professionals and technology leaders
  • Act as an agent of change with urgency: question how work gets done and remove or automate low-value steps while respecting existing implementations
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
Technologies
  • Databricks
  • AWS
  • Python
  • SQL
  • PySpark
  • Delta Lake
  • Unity Catalog
  • Databricks Jobs
  • Databricks SQL
  • Databricks Asset Bundles
  • Apache Airflow
  • Astronomer
  • dbt
  • Airflow
  • Deequ
  • Lakehouse Monitoring
  • Datadog
  • Terraform
  • Harness
  • Databricks Genie
  • CI/CD
Preferred Qualifications
  • Experience with investment management data such as portfolios, positions, returns, exposures, benchmarks, and attribution, or with multi-asset portfolio construction
  • Experience building LLM applications or agent workflows that call tools and APIs, including context management, retrieval, state, and error handling
  • Familiarity with Model Context Protocol (MCP)
  • Experience with PostgreSQL, SQL Server, or Lakebase, or modernizing legacy data platforms onto a Lakehouse
Benefits
  • Generous time-away and health benefits from day one, with opportunity for flexible work options
  • 2-for-1 matching gifts for charitable contributions
  • Opportunity to secure annual grants for organizations you love
  • Access on-demand professional development resources
  • Competitive salary, bonuses and benefits
  • Company-funded retirement contribution
  • Individual annual performance bonus
  • Capital’s annual profitability bonus
  • Retirement plan where Capital contributes 15% of eligible earnings
Location and Salary
  • Location: Los Angeles, CA (onsite)
  • Salary: USD 201,683 - 342,072 per yearly
  • Southern California base salary range: $201,683-$322,693
  • New York base salary range: $213,795-$342,072
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer Lead
Data Engineer Lead

504 CGCG-US CG Companies Global-US • Los Angeles (CA)

Hybrid
USD 202,000 - 323,000
Data Engineer Lead
Data Engineer Lead

Capital Group • Los Angeles (CA)

On-site
USD 202,000 - 342,000
Time-away from day one
Flexible work options
2-for-1 matching gifts
+1
Data Engineer Lead
Data Engineer Lead

Capital-Group-1 • Los Angeles (CA)

On-site
USD 202,000 - 323,000
Bonuses
Retirement plan contribution
Health benefits
Data and AI Engineering Manager
Data and AI Engineering Manager

504 CGCG-US CG Companies Global-US • Los Angeles (CA)

Hybrid
USD 202,000 - 323,000
Data and AI Engineering Manager
Data and AI Engineering Manager

Capital Group • Los Angeles (CA)

On-site
USD 202,000 - 323,000
Generous time away
Flexible work options
2-for-1 matching gifts
+2
Data Engineer Lead
Data Engineer Lead

Capital Group • New York (NY)

On-site
USD 180,000 - 260,000
Generous time-off and health benefits
Company-funded retirement contribution
2-for-1 charitable gifts matching
Data and Analytics Specialist
Data and Analytics Specialist

Capital Group • Los Angeles (CA), Northern (KY)

On-site
USD 154,000 - 287,000
Competitive salary and bonuses
Retirement contribution (15%)
Flexible work options
+3
Principal Data Product Manager - Enterprise Data Office
Principal Data Product Manager - Enterprise Data Office

504 CGCG-US CG Companies Global-US • Los Angeles (CA)

Hybrid
USD 208,000 - 354,000
Lead Data Engineer
Lead Data Engineer

Capital Technology Group • United States

Hybrid
USD 150,000 - 200,000
Remote Work
Medical, Dental, Vision
401(k) with 4% matching
+2
Data and AI Engineering Manager
Data and AI Engineering Manager

Capital-Group-1 • Los Angeles (CA)

Hybrid
USD 202,000 - 323,000
Flexible work options
Health benefits
Time off and wellness programs
+1