Data Engineer

RSM Solutions, Inc

Irvine (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RSM Solutions, Inc is seeking a Data Engineer in Irvine, California. The role involves designing data pipelines using Apache Spark, Kubernetes, and integrating various data quality measures. Ideal candidates will have extensive experience in developing data solutions and automating processes. The position emphasizes collaboration with data scientists for enhanced operational efficiency.

This onsite role is suitable for US Citizens or Green Card Holders, requiring strong skills in T-SQL, PowerShell, GitHub, and Azure DevOps.

Qualifications

  • Minimum 6 years experience developing data pipelines in Spark.
  • 2 years deploying workloads on Kubernetes/Kubeflow.
  • Experience with MLflow or similar experiment tracking.

Responsibilities

  • Design and implement batch and streaming pipelines in Spark.
  • Build high throughput ETL/ELT jobs with SSIS, SSAS, and T-SQL.
  • Automate model retraining and deployment triggers within Kubeflow.

Skills

Data pipeline development in Spark
Kubernetes
MLflow or similar tools
T-SQL
Python/Scala for Spark
PowerShell/.NET scripting
GitHub and Azure DevOps
Prometheus and Grafana

Education

6+ years of experience in data engineering

Tools

Apache Spark
Kubernetes
SSIS and SSAS
Azure Monitor

Job description

Overview

The Data Engineer role is similar to the Data Integration role but more Operations-focused, orchestrating deployments and ML flow, configuring data on clusters and managing model performance. This role bridges Data Engineering and MLOps, allowing data scientists to focus on experimentation while the business sees rapid, reliable predictive insight.


Location & Eligibility

This role is located onsite in Irvine, California. I prefer candidates that are local. Relocation is accepted but there are no relocation dollars available. I can only work with US Citizens or Green Card Holders for this role. Candidates with H1, OPT, EAD, F1, H4 or those not a US Citizen or Green Card Holder are not eligible.


Responsibilities


  • Design and implement batch and streaming pipelines in Apache Spark running on Kubernetes and Kubeflow Pipelines to hydrate feature stores and training datasets.

  • Build high throughput ETL/ELT jobs with SSIS, SSAS, and T‑SQL against MS SQL Server, applying Data Vault style modeling patterns for auditability.

  • Integrate source control, build, and release automation using GitHub Actions and Azure DevOps for every pipeline component.

  • Instrument pipelines with Prometheus exporters and visualize SLA, latency, and error budget metrics to enable proactive alerting.

  • Create automated data quality and schema drift checks; surface anomalies to support a rapid incident response process.

  • Use MLflow Tracking and Model Registry to version artifacts, parameters, and metrics for reproducible experiments and safe rollbacks.

  • Work with data scientists to automate model retraining and deployment triggers within Kubeflow based on data freshness or concept drift signals.

  • Develop PowerShell and .NET utilities to orchestrate job dependencies, manage secrets, and publish telemetry to Azure Monitor.

  • Optimize Spark and SQL workloads through indexing, partitioning, and cluster sizing strategies, benchmarking performance in CI pipelines.

  • Document lineage, ownership, and retention policies; ensure pipelines conform to PCI/SOX and internal data governance standards.


Qualifications


  • At least 6 years of experience building data pipelines in Spark or equivalent.

  • At least 2 years deploying workloads on Kubernetes/Kubeflow.

  • At least 2 years of experience with MLflow or similar experiment‑tracking tools.

  • At least 6 years of experience in T‑SQL, Python/Scala for Spark.

  • At least 6 years of PowerShell/.NET scripting.

  • At least 6 years of experience with GitHub, Azure DevOps, Prometheus, Grafana, and SSIS/SSAS.

  • Kubernetes CKA/CKAD, Azure Data Engineer (DP‑203), or MLOps‑focused certifications (e.g., Kubeflow or MLflow) would be great to see.

  • Mentor engineers on best practices in containerized data engineering and MLOps.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Integration Engineer
Data Integration Engineer

RSM Solutions, Inc • Irvine (CA)

On-site
USD 90,000 - 120,000
Data Engineer -- W2 ONLY
Data Engineer -- W2 ONLY

nTech Workforce • Oakbrook Terrace (IL)

On-site
USD 80,000 - 100,000
Data Engineer
Data Engineer

ATC • New York (NY)

On-site
USD 140,000 - 180,000
Data Engineer
Data Engineer

Optomi • Austin (TX)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

Jobtailor • Town of Florida (NY)

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

VTG Defense • McLean (VA)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

Akaasa Technologies • Atlanta (GA)

On-site
USD 96,000 - 152,000
Senior Data Engineer – 8+ Years Experience
Senior Data Engineer – 8+ Years Experience

Hudson Manpower • New Jersey

On-site
USD 140,000 - 190,000
Data Engineer
Data Engineer

Hirewell • Atlanta (GA)

On-site
USD 95,000 - 125,000
Data Engineer
Data Engineer

Full Scope • San Francisco (CA)

On-site
USD 120,000 - 170,000