Data Engineer, Forward Deployed

Applied Computing

Houston (TX)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Applied Computing in Houston seeks a Data Engineer to architect and maintain high-frequency time-series and historian data pipelines, enabling AI-ready data for deep learning models and real-time LLM workloads.

You will work across AWS (EKS, S3, EBS, IAM, KMS, CloudWatch) and Databricks/PySpark, ensuring data is contextualised, synchronised, and optimised for AI applications in energy operations.

Qualifications

  • Strong Python for data processing, scripting, and orchestration.
  • Expertise in PostgreSQL (partitioning, indexing, performance).
  • Extensive AWS experience for data pipelines (EKS, S3, IAM, KMS).
  • Experience with Databricks and PySpark for large-scale processing.
  • Familiarity with time-series industrial data (DCS/SCADA logs, historians).
  • Knowledge of streaming frameworks (Kafka, Flink) or MLOps data versioning is a plus.

Responsibilities

  • Ingest, contextualise, and move data from diverse sources into Lakehouse.
  • Build and optimise real-time and batch pipelines for AI workloads.
  • Maintain data lineage, governance, and schema drift controls.
  • Ensure performance tuning for PostgreSQL and Spark-based workloads.
  • Collaborate across teams to align data availability with ML/LLM needs.

Skills

Python
PostgreSQL
AWS
Databricks
PySpark
Time-series data

Tools

Databricks
PySpark
AWS
Kubernetes
PostgreSQL

Job description

Applied Computing was founded in 2024 to build Orbital, a physics-informed foundation model for energy operations. We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable
abundance for a growing planet.

The hydrocarbon industry keeps the world running. But its complexity has left operators tied to legacy systems, making critical decisions on less than 10% of available data. We built Orbital to change that. It’s a foundation model built specifically for energy that lets companies use AI at scale, harnessing all of their operational
data and optimising in real time for any metric. Decisions get faster, operations get safer, and carbon intensity falls.

We’ve raised over $32 million, including one of the largest seed rounds for an
AI company in the UK. We’re just getting started

The Role

As our Data Engineer, you’ll architect and maintain pipelines that make high-frequency time-series, lab, and historian data into a scalable Lakehouse architecture, usable for both deep learning models and real-time LLMs. You’ll be working across AWS (EKS, S3, EBS, KMS, CloudWatch) and Databricks/PySpark, ensuring data is contextualised, synchronised, and optimised for both deep learning models and real-time LLM workloads.

This isn’t a traditional ETL role, you’ll be solving problems at the intersection of control systems, industrial data engineering, and AI enablement.

Technical Requirements
  • Deep expertise in PostgreSQL (partitioning, indexing, query optimisation, storage design).
  • Strong proficiency in Python for data processing, scripting, and pipeline orchestration.
  • Hands-on experience with AWS (EKS, S3, EBS, IAM, KMS, CloudWatch, etc.)for secure and scalable data pipelines.
  • Proven ability to work with Databricks and PySpark for large-scale distributed data processing.
  • Familiarity with time-series industrial data (control systems, DCS/SCADA logs, process historians).
  • Experience in unstructured data sync and management within hybrid cloud/on-prem environments.
  • Bonus: Experience working as a data engineer in oil and gas or energy environments
  • Bonus: Knowledge of streaming frameworks (Kafka, Flink, Spark Streaming) or MLOps stacks for data versioning and lineage.
Core Responsibilities
1. Ingest & Contextualise Data
  • Ingest from OPC UA servers, process historians, IoT sensors, LIMS systems, alarms/events, and P&IDs.
  • Map signals to their physical processes (tags, units, hierarchies) for interpretability in AI pipelines.
2. Data Movement & Accessibility
  • Build pipelines that handle real-time streaming and batch ingestion into the Lakehouse.
  • Manage synchronisation between historian archives, unstructured files, and AWS storage (S3/EBS).
  • Orchestrate Databricks Lakeflow/Connectors for integrating data into Lakebase/Lakehouse.
  • Handle secure, high-throughput transfers between historian archives and sandbox/live environments.
3. Change Tracking & Integrity
  • Detect and manage schema changes, signal drift, and inconsistencies acrosstime.
  • Implement lineage and audit trails across Spark/Databricks and AWS pipelines.
4. Data Preparation for AI
  • Build and maintaindual pipelines:
    • Training→ large-scale historical data prep for time-series + LLM training.
    • Inference→ low-latency, real-time pipelines for anomaly detection, optimisation, and LLM search.
  • Support heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs).
5. Database Performance & Optimisation
  • Tune PostgreSQLand sparkfor high-throughput time-series workloads (partitioning, indexing, query optimisation).
  • Optimise pipelines for both fast analytical queries and high-efficiency model training.
  • Deploy and manage data pipelines in AWS EKS (Kubernetes) with persisten tEBS-backed storage.
What Success Looks Like
  • Live data streams are contextualised,queryable, and AI-ready.
  • Schema changes and signal drift are detected and handled without breaking downstream workflows.
  • Training and inference pipelines run smoothly in parallel, optimised for scale and latency.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Oscar • Grand Prairie (TX)

On-site
USD 110,000 - 150,000
Medical coverage
Dental coverage
Vision coverage
+3
Energy AI Data Engineer - Real-Time Lakehouse
Energy AI Data Engineer - Real-Time Lakehouse

Applied Computing • Houston (TX)

On-site
USD 120,000 - 160,000
ML Engineer (Forward Deployed)
ML Engineer (Forward Deployed)

Applied Computing • Houston (TX)

On-site
USD 140,000 - 210,000
Data Engineer
Data Engineer

Ranger Technical Resources • Town of Florida (NY)

On-site
USD 130,000 - 185,000
Data Engineer – AWS Lakehouse (Mandarin Required)
Data Engineer – AWS Lakehouse (Mandarin Required)

Bitus Labs • Irvine (CA)

On-site
USD 120,000 - 180,000
Data Engineer (Spark)
Data Engineer (Spark)

Addepto • Town of Poland (NY)

Hybrid
USD 110,000 - 170,000
Flexible remote or office work
Professional training and conferences
Paid time off
+2
Databricks Data Architect
Databricks Data Architect

Unison Group • Town of Charlotte (NY)

On-site
USD 94,000 - 142,000
Databricks Data Architect
Databricks Data Architect

Unison Group • California City (CA)

On-site
USD 94,000 - 142,000
Databricks Data Architect
Databricks Data Architect

UNISON Group • Charlotte (NC)

On-site
USD 94,000 - 149,000
Data Platform Lead - IT Transformation
Data Platform Lead - IT Transformation

Southwest Generation Valmont Facility • Denver (CO)

On-site
USD 120,000 - 170,000
Health insurance
Life insurance
Retirement savings plan