Data Engineer - AI/ML Engineer

IntraEdge Technology

United States

Remote

USD 120,000 - 170,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

IntraEdge Technology is seeking a Data Engineer - AI/ML Engineer to design and maintain scalable data pipelines for analytics, ML, and operations. You will work with Python, SQL, and modern data engineering frameworks, collaborating with Data Scientists to productionize NLP models.

The role includes building batch and streaming pipelines with Kafka, using GCP services (BigQuery, Dataflow, Pub/Sub, Cloud Storage) and deploying data APIs.

Qualifications

  • 5+ years building data pipelines or backend data workflows.
  • 2+ years designing REST/GraphQL data services or APIs.
  • Hands-on ML/AI model integration in production.
  • Experience with healthcare data, NLP, and structured/unstructured datasets.
  • 2+ years cloud platforms (GCP preferred).
  • 2+ years with Kafka or similar streaming platforms.
  • 2+ years databases (Postgres or equivalent).
  • 2+ years CI/CD tools and IaC (Terraform preferred).
  • 2+ years Kubernetes (GKE/EKS/AKS).
  • Knowledge of cloud networking and IAM workflows.
  • Familiarity with monitoring tools.

Responsibilities

  • Design, build and maintain scalable data pipelines for analytics, ML, and reporting.
  • Develop robust data ingestion, transformation, and integration workflows with Python/SQL.
  • Build batch and streaming pipelines using Kafka and modern frameworks.
  • Leverage GCP services (BigQuery, Dataflow, Pub/Sub, Cloud Storage, Cloud Functions).
  • Develop data APIs (REST/GraphQL) for high-performance access across microservices.
  • Implement CI/CD automation for data pipelines (GitHub Actions, Argo CD).
  • Collaborate with Data Scientists to integrate ML/NLP models into pipelines.
  • Operationalize NLP data pipelines for structured/unstructured data (Rx claims, clinical docs).
  • Enable model retraining workflows using Vertex AI, Kubeflow, or similar tooling.
  • Establish observability/data quality with logs, metrics, and data lake integration.
  • Support reliability, performance, and scalability of data infrastructure.

Skills

Python
SQL
Data pipelines
REST API
GraphQL
Kafka
GCP
Kubernetes
CI/CD
IAM
Monitoring
NLP
MLOps
Vertex AI

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
GitHub Actions
Argo CD
Jenkins
Pulumi

Job description

Job Title: Data Engineer - AI/ML Engineer

Location: Remote/Pune

Job Type: Full Time

Shift timing: 2pm to 11pm


Responsibilities
  • Design, build, and maintain scalable data pipelines to support analytics, ML, and operational reporting.
  • Develop robust data ingestion, transformation, and integration workflows using Python, SQL, and modern data engineering frameworks.
  • Build and maintain batch and streaming data pipelines leveraging technologies such as Kafka (or similar pub/sub tools).
  • Work with Google Cloud Platform (GCP) services, including Cloud Storage, Dataflow, Pub/Sub, BigQuery, Cloud Spanner and Cloud Functions.
  • Develop and manage data APIs and interfaces (REST and GraphQL) to enable high-performance data access across microservices.
  • Implement CI/CD automation for data pipelines using GitHub Actions, Argo CD, or equivalent tools.
  • Collaborate with Data Scientists and MLOps teams to integrate ML/NLP models into data pipelines and production workflows.
  • Build and operationalize NLP data pipelines for structured and unstructured data sources (e.g., Rx claims, clinical documents).
  • Enable continuous learning and model-retraining workflows using Vertex AI, Kubeflow, or similar GCP-native tooling.
  • Implement frameworks for observability and data quality, ensuring ML predictions, confidence scores, and fallback events are logged into data lakes or monitoring systems.
  • Support distributed data systems and ensure reliability, performance, and scalability of data infrastructure.
Infrastructure Engineering Responsibilities
  • Design and provision cloud infrastructure using Infrastructure as Code (IaC) tools such as Terraform or Pulumi for GCP resources including GKE clusters, Cloud SQL, VPC networks, IAM, and storage.
  • Deploy, configure, and manage containerized data workloads using Kubernetes (GKE) including deployments, autoscaling (HPA/VPA), namespaces, resource quotas, and health checks.
  • Architect and maintain network topology for data platform environments VPCs, subnets, firewall rules, private service access, Cloud NAT, and VPC Service Controls.
  • Implement and enforce IAM policies, service account governance, and secrets management (GCP Secret Manager or HashiCorp Vault) to ensure least-privilege access across all data services.
  • Build and maintain infrastructure monitoring and alerting using Cloud Monitoring, Prometheus, Grafana, or equivalent — covering pipeline latency, throughput, error rates, and resource utilization.
  • Establish and maintain CI/CD pipelines for infrastructure changes using Terraform Cloud, GitHub Actions, or Argo CD, ensuring infrastructure drift detection and rollback capability.
  • Manage environment parity (dev/staging/prod) for data platform infrastructure, including environment-specific configuration management and promotion workflows.
  • Drive cloud cost governance — right-sizing compute resources, implementing committed-use discounts, setting up budget alerts, and producing cost attribution reports per workload.
  • Design and implement disaster recovery, backup, and high-availability strategies for data stores, pipeline infrastructure, and ML serving endpoints.
  • Collaborate with Security and Platform teams to ensure data infrastructure compliance with enterprise security policies, SOC 2, HIPAA, and CVS Health regulatory requirements.

Required Qualifications
  • 5+ years of experience building data pipelines or backend data workflows using Python, Java, or similar languages.
  • 2+ years of experience designing REST/GraphQL data services or integrating data APIs.
  • Hands-on experience working with ML/AI model integration in production (e.g., Vertex AI Endpoints, TensorFlow Serving, ML REST APIs).
  • Experience handling structured and unstructured datasets, including healthcare data (Rx claims, clinical documents, NLP text).
  • Familiarity with the end-to-end ML lifecycle: data ingestion, feature engineering, training, deployment, and real-time inference.
  • 2+ years of experience with cloud platforms (GCP preferred; AWS or Azure acceptable).
  • 2+ years working with streaming platforms like Kafka or equivalent.
  • 2+ years of experience with databases (Postgres or similar relational systems).
  • 2+ years of experience with CI/CD tools (GitHub Actions, Jenkins, Argo CD, etc.).
  • 2+ years of hands-on experience with Infrastructure as Code tools (Terraform preferred; Pulumi or CDK acceptable).
  • 2+ years managing containerized workloads using Kubernetes (GKE, EKS, or AKS) — deploying services, configuring autoscaling, and managing resource limits.
  • Solid understanding of cloud networking fundamentals: VPCs, subnets, firewall rules, private connectivity, and DNS resolution in GCP or AWS.
  • Experience designing IAM roles, service accounts, and secrets management workflows to enforce least-privilege access across data services.
  • Familiarity with infrastructure monitoring and alerting tools (Cloud Monitoring, Prometheus/Grafana, or equivalent).

Preferred Qualifications
  • Direct, hands-on experience with Google Cloud Platform, especially BigQuery, Dataflow, GKE, Composer and Vertex AI.
  • Knowledge of Kubernetes concepts and experience running data services or pipelines on GKE.
  • Strong understanding of distributed systems, microservice patterns, and data-centric system design.
  • Experience using Vertex AI, Kubeflow, or other ML orchestration platforms for model training and serving.
  • Knowledge of GenAI pipelines, LLM prompt workflows, and agent orchestration frameworks (e.g., LangChain, transformers).
  • Experience deploying Python-based ML/NLP services into microservice ecosystems using REST, gRPC, or sidecar architectures.
  • Domain experience in healthcare, claim adjudication, or Rx data processing.
  • Experience with Terraform modules, workspaces, and remote state backends for managing multi-environment GCP infrastructure.
  • Familiarity with GCP Shared VPC, VPC Service Controls, or private Google Access configurations for secure data platform networking.
  • Exposure to FinOps practices — cloud cost attribution, showback/chargeback models, and resource tagging strategies.
  • Experience with GitOps workflows using Argo CD or Flux for managing infrastructure and application delivery.
Education
  • Bachelor's degree in Computer Science, Data Engineering, Information Systems, or equivalent experience (High School Diploma + 4 years of relevant experience acceptable).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data AI Engineer
Senior Data AI Engineer

KADIT Innovations • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 170,000
Gcp Data Engineer
Gcp Data Engineer

Reliable Group Paragon • United States

Remote
USD 12,000 - 22,000
Software Engineer- Data Engineering
Software Engineer- Data Engineering

Noctua Technology • Washington, Virginia (MN), California (MO)

Remote
USD 103,000 - 170,000
Data Engineering Manager - (GCP, AI/ML & GenAI) @ Naveera Tech, USA - Remote Work
Data Engineering Manager - (GCP, AI/ML & GenAI) @ Naveera Tech, USA - Remote Work

Naveera Technology LLC • Denver (CO)

Remote
USD 180,000 - 240,000
Data Scientist Engineer
Data Scientist Engineer

Compunnel, Inc. • Town of Florida (NY)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

SZNS Solutions • Reston (VA)

On-site
USD 120,000 - 150,000
Competitive salary and benefits package
Hybrid work environment
Continuous learning and development opportunities
+1
Data Engineering Manager - (GCP, AI/ML & GenAI) @ Naveera Tech, USA - Remote Work
Data Engineering Manager - (GCP, AI/ML & GenAI) @ Naveera Tech, USA - Remote Work

Naveera Technology LLC • United States

Remote
USD 170,000 - 210,000
Flexible remote work
Global clientele exposure
Certification opportunities
Data Engineer
Data Engineer

SZNS • Reston (VA)

On-site
USD 100,000 - 130,000
Competitive salary and benefits package
Hybrid work environment
Collaborative work environment
+1
Senior Staff Engineer-Data and AI Platform
Senior Staff Engineer-Data and AI Platform

H.E.B. • Austin (TX)

On-site
USD 180,000 - 230,000
Senior Staff Engineer-Data and AI Platform
Senior Staff Engineer-Data and AI Platform

H-E-B • Austin (TX)

On-site
USD 180,000 - 240,000