Sr Data Engineer, Data Analytics & Intelligence, NA

Vantage Data Centers Management Company LLC

Denver (CO)

On-site

USD 130,000 - 155,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision coverage
Life and AD&D
401k with company match
Paid time off
Employee assistance

Job summary

Vantage Data Centers Management Company LLC in Denver, CO is seeking a data engineer to build and scale governed Azure data foundations for enterprise reporting, operational intelligence, and AI-enabled consumption across North America. You will design, implement, and operate data pipelines on Azure, focusing on reliability and scalability.

You will work with PySpark, Python, SQL, and Azure Data Factory, maintain lakehouse structures, data dictionaries, lineage, and governance, collaborating

Qualifications

  • Bachelor’s degree in a related field or equivalent experience.
  • 5–8 years of data engineering / analytics engineering experience.
  • Proficiency in Python for data pipelines and PySpark-based transformations.
  • Proficiency in SQL for querying and data quality checks.
  • Experience with Azure data services including Data Factory, Synapse, and Data Lake Storage Gen2.
  • Experience with CI/CD using GitHub or Azure DevOps.

Responsibilities

  • Build, maintain, and optimize data pipelines on Azure for enterprise reporting and AI-enabled consumption.
  • Develop batch and incremental pipelines with Azure Data Factory; manage lakehouse datasets and semantic-model inputs.
  • Implement Spark- and SQL-based transformations to produce curated datasets.
  • Own pipelines and datasets including monitoring, troubleshooting, and production support.
  • Collaborate across teams to translate requirements into scalable data solutions and ensure governance.

Education

Bachelor’s degree in Engineering, Computer Science, Data Analytics, or a related field

Tools

Python
PySpark
Spark
SQL
Azure
Azure Data Factory
Azure Synapse
Microsoft Fabric / Lakehouse patterns
Fabric Lakehouse
GitHub
Azure DevOps
Jira

Job description

Build and scale governed Azure data foundations for enterprise reporting, operational intelligence, and AI-enabled consumption for Operations in North America.

Responsibilities
  • Design, build, and maintain reliable, scalable data pipelines using Python and PySpark on the Microsoft Azure data platform
  • Develop and operate batch and incremental pipelines using Azure Data Factory for orchestration and Azure Data Lake Storage Gen2 as the primary data store
  • Create and maintain curated lakehouse / gold-layer datasets and semantic-model inputs for governed operational insights and AI-enabled consumption
  • Independently implement SQL- and Spark-based transformations to produce curated datasets for enterprise reporting, analytics, and downstream use
  • Own assigned pipelines and datasets, including monitoring, troubleshooting, performance optimization, documentation, and production support
  • Work with Azure Synapse and Microsoft Fabric / Lakehouse patterns where applicable, plus related Azure analytics services
  • Prepare structured operational data for AI-enabled use cases by documenting business rules, source lineage, data reliability constraints, known quality limitations, and data dictionary definitions
  • Support source visibility and confidence context, including Data Reliability & Trust Indicator integration where applicable
  • Contribute to ontology, taxonomy, semantic model, and data dictionary alignment across operational domains (KPIs, incidents, work orders, and more)
  • Collaborate with business analysts, operations SMEs, data stewards, IT Global, and cross-functional stakeholders to translate requirements into working data solutions
  • Apply data governance, security, access-control, data classification, and engineering standards for compliant, maintainable, scalable solutions
  • Identify, document, and route data-quality issues to accountable owners to improve source correction rather than masking defects downstream
  • Participate in code reviews, technical discussions, sprint planning, and platform improvement initiatives
  • Proactively identify data quality issues, pipeline risks, platform dependencies, and improvement opportunities, and communicate them clearly
  • Develop and maintain PySpark notebooks and jobs to ingest, transform, validate, and curate data within the enterprise data platform
  • Build and modify Azure Data Factory pipelines for batch and incremental ingestion
  • Implement Spark transformations to write curated datasets to Azure Data Lake Storage Gen2 and/or Fabric Lakehouse patterns using established folder structures, naming conventions, and governance standards
  • Create and maintain SQL views, tables, lakehouse objects, and semantic-model inputs to support analytics, operational intelligence, and AI-enabled consumption patterns
  • Prepare datasets for Fabric Data Agent / AI agent use cases with documentation including business rules, joins, grain, quality limitations, source lineage, and operational definitions
  • Respond to pipeline failures, data validation issues, operational alerts, and data-quality escalations with root-cause analysis and remediation steps
  • Perform performance tuning for Spark jobs and SQL workloads (partitioning, filtering, incremental logic, query optimization, and resource-aware design)
  • Validate outputs with business partners, operations SMEs, and data stewards; address defects or discrepancies through documented correction paths
  • Support observability, logging, and auditability practices for pipelines and AI-consumable datasets where applicable
  • Commit code using Git, follow branching standards, support pull request reviews, and participate in CI/CD using GitHub, Azure DevOps, or similar tools
  • Update documentation for pipelines, datasets, data contracts, data dictionaries, business rules, and operational runbooks as changes are made
  • Execute assigned backlog items within sprint timelines; raise risks, dependencies, or blockers early
  • Additional duties as assigned by management
Requirements
  • Bachelor’s degree in Engineering, Computer Science, Data Analytics, or a related field, or equivalent experience
  • Minimum 5–8 years of experience in data engineering, analytics engineering, or a closely related technical data role
  • Proficiency in Python for data pipelines, automation, and data processing workflows, including PySpark-based transformations
  • Proficiency in SQL for querying, transformation, analytical processing, model validation, and data quality checks
  • Solid understanding of ETL/ELT pipelines, data transformation patterns, data integration concepts, incremental processing, and production support practices
  • Experience analyzing enterprise data sources to identify relationships, transformations, business rules, grain, ownership, and quality constraints
  • Experience building solutions on the Microsoft Azure platform, including exposure to Azure Data Factory, Azure Synapse, Azure Data Lake Storage Gen2, Microsoft Fabric / Lakehouse patterns, and related analytics services
  • Working knowledge of data modeling fundamentals, including fact/dimension tables and semantic models
  • Experience supporting governed data products, including metadata, lineage, issue documentation, access-control awareness, data-quality validation, and operational runbooks
  • Experience with source control and CI/CD workflows using GitHub or Azure DevOps
  • Strong communication and collaboration skills across IT Global, business SMEs, data governance partners, platform teams, and operations stakeholders
  • Experience in Agile environments and using collaboration or tracking tools such as Jira or similar tools
  • Travel required is expected to be up to 10% but may increase over time
Technologies
  • Python, PySpark, Spark
  • Microsoft Azure, Azure Data Factory, Azure Data Lake Storage Gen2, Azure Synapse
  • Microsoft Fabric, Fabric Lakehouse, Microsoft Fabric / Lakehouse patterns
  • SQL
  • Git, GitHub, Azure DevOps
  • Jira, CI/CD
Benefits
  • Medical, dental, and vision coverage
  • Life and AD&D
  • Short and long-term disability coverage
  • Paid time off
  • Employee assistance
  • 401k participation with company match
  • Above market total compensation package
  • Comprehensive suite of health and welfare, retirement, and paid leave benefits
Salary Range
  • USD 130,000 - 155,000 per year
  • Range is based on Colorado market data and may vary in other locations
  • Compensation may depend on qualifications, skills, competencies, and experience and may fall outside the range shown
Additional Details
  • On-site role in Denver, CO (3 days on site required, 2 days flexible)
  • Eligible for company benefits including medical, dental, and vision coverage, life and AD&D, short and long-term disability coverage, paid time off, employee assistance, and 401k with company match
Physical Demands and Special Requirements
  • Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions
  • Occasionally required to stand, walk, sit; use hands to handle or feel objects; reach with hands and arms; climb stairs; balance; stoop or kneel; talk and hear
  • Occasionally required to lift and/or move up to 25 pounds
Desired Qualifications
  • Experience with distributed data processing frameworks, including Apache Spark
  • Experience preparing governed data products for AI-enabled use cases, including Microsoft Fabric Lakehouse, semantic models, Fabric/Data Agent patterns, ontology or taxonomy alignment, and explainable AI outputs
  • Familiarity with data observability, metadata management, lineage, data contracts, reliability indicators, and operational best practices in production environments
  • Familiarity with additional Azure services such as Azure Functions or Logic Apps in support of data workflows
  • Experience supporting data platform enhancement, refactoring, modernization, or reusable architecture initiatives
  • Experience working with structured and unstructured operational sources, such as enterprise applications, operational workflows, documents, dashboards, and knowledge assets
  • Experience working in a scaling or fast-paced organization where priorities evolve quickly and practical delivery discipline is required
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Loyal Source Government Services • Orlando (FL)

On-site
USD 113,000 - 138,000
Senior Data Engineer – 8+ Years Experience
Senior Data Engineer – 8+ Years Experience

Hudson Manpower • New Jersey

On-site
USD 140,000 - 190,000
Senior Data Engineer – 8+ Years Experience
Senior Data Engineer – 8+ Years Experience

Hudson Manpower • New York (NY)

On-site
USD 140,000 - 190,000
Data Engineer
Data Engineer

Sakata Seed America, Inc. • Woodland (CA)

On-site
USD 110,000 - 155,000
Medical, Dental & Vision Insurance
401(k) with Company Match
Paid Vacation & Holidays
Principal Engineer I - Senior Data Engineer
Principal Engineer I - Senior Data Engineer

Software Resources, Inc. • Phoenix (AZ)

Hybrid
USD 120,000 - 160,000
Medical, dental, and vision coverage
401(k) with company match
Short-term disability
+1
Sr Data Engineer, Data Analytics & Intelligence, NA
Sr Data Engineer, Data Analytics & Intelligence, NA

Vantage Data Centers • Denver (CO)

On-site
USD 130,000 - 155,000
Medical insurance
Dental insurance
Vision insurance
+4
Cloud Data Platform Architect
Cloud Data Platform Architect

AdventHealth • Town of Florida (NY)

On-site
USD 96,000 - 180,000
Medical, Dental, and Vision Insurance
Paid Time Off
Retirement Plan
+2
Data Engineer
Data Engineer

Jobtailor • Town of Florida (NY)

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

San Antonio Spurs • San Antonio (TX)

On-site
USD 80,000 - 110,000
Cloud Data Platform Architect
Cloud Data Platform Architect

AdventHealth • United States

On-site
USD 96,266 - 179,046