Data Engineer

Vitol

Houston (TX)

On-site

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical
Vision
Paid Vacation
401k with company contributions
Life insurance
Short-term & Long-term disability

Job summary

Vitol is seeking a Data Engineer to build and run pipelines and data models behind our core data platform, supporting energy traders and analysts in real time.

You will write Python, migrate pipelines to a Prefect-based framework, and collaborate across the stack with senior engineers and desk users to deliver reliable, scalable data solutions.

Qualifications

  • 3+ years of hands-on data engineering experience building and operating production data pipelines.
  • Strong Python with pandas and data-handling expertise.
  • SQL modeling, querying, and optimization awareness.
  • Snowflake or equivalent cloud warehouse; dbt experience is a plus.
  • Broad data tech knowledge across streaming, caching, timeseries, relational and warehouse systems.
  • Pipeline orchestration experience; Prefect preferred.
  • Solid software engineering fundamentals and version control habits.
  • Knowledge of AWS data services and reliable design patterns.
  • Build-to-operate mindset including monitoring and alerting.
  • Clear communication; ability to discuss trade-offs with analysts.
  • Terraform, Docker, or API development exposure.
  • Experience in financial services, commodities, or energy trading data is a plus.

Responsibilities

  • Build, test, and operate production data pipelines in Python on Prefect.
  • Maintain Snowflake warehouse models with dbt; ensure quality and cost awareness.
  • Migrate legacy pipelines and Oracle components to current frameworks.
  • Work across Kafka, Redis, InfluxDB, Oracle, and Snowflake following established patterns.
  • Acquire data from vendor APIs, files, feeds, and web sources reliably.
  • Implement data quality, freshness, and reconciliation checks.
  • Leverage AWS data services where patterns require them.
  • Monitor, alert, and troubleshoot components in production.
  • Collaborate with analysts and desk users to ensure solutions meet real needs.
  • Contribute to making platform data AI-ready and properly documented.

Skills

Data engineering experience
Python
SQL
Data tech breadth
Pipeline orchestration
Engineering fundamentals
AWS data services
Build-to-operate mindset
Clear communication
Energy trading data

Tools

Snowflake
Prefect
dbt
Docker
Terraform
API development

Job description

Vitol is a leader in energy and commodities. Vitol produces, manages and delivers energy and commodities to consumers and industry worldwide.In addition to its primary business of trading, Vitol is invested in infrastructure globally, with $10+billion invested in long-term assets.

Vitol’s customers include national oil companies, multinationals, leading industrial companies and utilities. Founded in Rotterdam in 1966, today Vitol serves its customers from some 40 offices worldwide. Revenues in 2023 were $400bn.

Our people are our business. Talent is precious to us and we create an environment in which individuals can reach their full potential, unhindered by hierarchy. Our team comprises more than 65+ nationalities and we are committed to developing and sustaining a diverse work force.

We are looking for a Data Engineer to build and run the pipelines and data models behind our core data platform — the centralized data layer that serves global energy traders and analysts in real time.

This is a hands‑on delivery role. You will write the Python that moves market and fundamental data from source through to the analysts, applications, and AI systems that consume it, and you will support what you ship.

Python is the core skill, but it is not the whole job. We are standing up a Snowflake warehouse, migrating legacy pipelines onto a Prefect-based framework, and consolidating how the platform stores and exposes data. The work reaches across more of the stack than pipeline code alone.

The platform is polyglot by design: streaming, caching, timeseries, relational, and warehouse technologies each doing what they are good at. You will build to the established patterns across that stack, and we expect you to understand why each technology sits where it does.

You will work alongside senior engineers who set those patterns, and directly with the analysts and desk users your work serves. The scope grows as you do — this is a role you can build a platform career from.

What You Will Do
  • Build, test, and operate production data pipelines in Python on our modern pipeline framework, orchestrated with Prefect
  • Build and maintain warehouse models in Snowflake using dbt — clean, tested, documented, and cost‑aware
  • Migrate legacy pipelines and Oracle-based components onto current frameworks and standards, retiring technical debt as you go
  • Work across the platform stack — Kafka, Redis, InfluxDB, Oracle, and Snowflake — building to the established pattern for each rather than reaching for the tool you already know
  • Acquire data from external sources — vendor APIs, files, feeds, and web sources — and land it reliably
  • Implement data quality, freshness, and reconciliation checks so problems surface before users find them
  • Use AWS data services where our platform patterns call for them
  • Support what you ship — monitoring and alerting on your components, and investigating when something breaks
  • Contribute to making platform data AI‑ready, and work with the catalog and steward teams so what you build is documented, classified, and findable
  • Engage directly with analysts and desk users to check that what you are building solves the actual problem
Qualifications
  • 3+ years of hands‑on data engineering experience building and operating production data pipelines
  • Strong Python — clean, tested, maintainable code, not scripts that happen to run. Fluency with pandas and the wider data‑handling ecosystem
  • SQL depth — you can model, query, and tune, and you know what makes a query expensive before you run it
  • Snowflake or a comparable cloud warehouse — dimensional modeling, performance tuning, and cost‑aware design. dbt experience is a strong plus
  • Breadth across data technologies — streaming, caching, timeseries, relational, and warehouse, with a view on where each belongs. Kafka, Redis, InfluxDB, Oracle, and Snowflake are what we run; comparable exposure matters more than an exact match
  • Pipeline orchestration experience — Prefect preferred; Airflow, Dagster, or similar considered
  • Sound engineering fundamentals — object‑oriented design, design patterns, testing, code review, and version control as habits rather than requirements
  • Working knowledge of AWS data services and how to compose them into something reliable
  • A build‑to‑operate mindset — monitoring, alerting, and failure recovery are part of how you design, not something added later
  • Clear communication — you can explain a technical trade‑off to an analyst, and turn a vague request into the right set of questions
  • Terraform, Docker, or API development (FastAPI, Flask) exposure is welcome — we use all three
  • Experience in financial services, commodities, or energy trading data is a plus; the data volume, latency requirements, and stakes are real
Additional Information
What Success Looks Like in Year One
  • You own several production data components outright and they run reliably with minimal intervention
  • Warehouse models you built in Snowflake are in active use by analysts and downstream applications
  • A meaningful share of the legacy pipelines you inherited have been migrated onto the current framework and standards
  • Data quality and freshness checks you implemented are catching issues before users report them
  • You build to the right platform pattern without being told which one applies — and you say so when none of them fit
  • Analysts on your projects come to you directly, and trust what you deliver
  • Medical
  • Vision
  • Paid Vacation
  • 401k with company contributions
  • Life insurance
  • Short-term & Long-term disability

All your information will be kept confidential according to EEO guidelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Red Cat Holdings • Salt Lake City (UT)

On-site
USD 110,000 - 150,000
Annual equity package
Bonus potential
Data Engineer - Tier 1 Global Asset Manager
Data Engineer - Tier 1 Global Asset Manager

Mondrian Alpha • New York (NY)

On-site
USD 120,000 - 160,000
Data Support Engineer
Data Support Engineer

Vitol • Houston (TX)

On-site
USD 85,000 - 120,000
Medical insurance
Dental insurance
Vision insurance
+4
Data Engineer
Data Engineer

Sakata Seed America, Inc. • Woodland (CA)

On-site
USD 110,000 - 155,000
Medical, Dental & Vision Insurance
401(k) with Company Match
Paid Vacation & Holidays
Bi-Lingual Senior Data Engineer
Bi-Lingual Senior Data Engineer

GTN Technical Staffing • United States

On-site
USD 100,000 - 130,000
Lead Data Engineer
Lead Data Engineer

K2 Integrity • New York (NY)

On-site
USD 140,000 - 210,000
Lead Data Engineer
Lead Data Engineer

K2 Intelligence, LLC • Northern (KY)

Hybrid
USD 120,000 - 180,000
Data Engineer
Data Engineer

Prodigy Resources • Denver (CO)

On-site
USD 110,000 - 170,000
Sr. Data Engineer
Sr. Data Engineer

Save-A-Lot, Ltd. • Missouri

On-site
USD 100,000 - 140,000
401K match up to 4%
Paid Time Off
Medical Insurance options including FS
+5
Data Engineer
Data Engineer

Errgo • Town of Boston (NY)

Hybrid
USD 120,000 - 180,000
Medical, dental, and vision insurance
401(k)
Equity
+2