Data Engineer - AI/ML Engineer

IntraEdge, Inc.

Pune District

Remote

INR 2,500,000 - 4,000,000

Full time

20 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

IntraEdge, Inc. is seeking a senior data engineer to design and maintain scalable data pipelines for analytics, ML, and operational reporting. You will ingestion, transformation, and exposure of data via REST/GraphQL APIs, collaborating with Data Scientists and MLOps teams.

The role involves building streaming/batch pipelines on GCP, implementing CI/CD with GitHub Actions, Argo CD, and managing Kubernetes workloads. Experience with NLP data and Vertex AI is valued.

Qualifications

  • 5+ years of experience building data pipelines using Python, Java, or similar languages.
  • 2+ years of experience designing REST/GraphQL data services or data APIs.
  • Hands-on experience with ML/AI model integration in production (Vertex AI endpoints, TF Serving, ML REST).
  • Experience with structured and unstructured data including healthcare data and NLP text.
  • 2+ years of cloud platforms (GCP preferred; AWS/Azure acceptable).
  • 2+ years of CI/CD tools (GitHub Actions, Jenkins, Argo CD).
  • 2+ years IaC tools (Terraform preferred; Pulumi acceptable).
  • 2+ years Kubernetes (GKE) deploying services and autoscaling.

Responsibilities

  • Design, build, and maintain scalable data pipelines for analytics, ML, and reporting.
  • Ingest, transform, and integrate data using Python, SQL, and modern frameworks.
  • Leverage Kafka-like pub/sub for batch and streaming pipelines.
  • Work with GCP services: BigQuery, Dataflow, Pub/Sub, Cloud Storage, Cloud Functions.
  • Develop data APIs (REST/GraphQL) for high-performance access across microservices.

Skills

Python
REST/GraphQL
GCP
Kubernetes
CI/CD
NLP

Education

Bachelor's degree in CS/ Data Eng

Tools

Terraform
Pulumi
GitHub Actions
Argo CD
Kafka
BigQuery

Job description

Location: Remote/Pune

Job Type: Full Time

The shift timing: 2pm to 11pm

Responsibilities

  • Design, build, and maintain scalable data pipelines to support analytics, ML, and operational reporting.
  • Develop robust data ingestion, transformation, and integration workflows using Python, SQL, and modern data engineering frameworks.
  • Build and maintain batch and streaming data pipelines leveraging technologies such as Kafka (or similar pub/sub tools).
  • Work with Google Cloud Platform (GCP) services, including Cloud Storage, Dataflow, Pub/Sub, BigQuery, Cloud Spanner and Cloud Functions.
  • Develop and manage data APIs and interfaces (REST and GraphQL) to enable high-performance data access across microservices.
  • Implement CI/CD automation for data pipelines using GitHub Actions, Argo CD, or equivalent tools.
  • Collaborate with Data Scientists and MLOps teams to integrate ML/NLP models into data pipelines and production workflows.
  • Build and operationalize NLP data pipelines for structured and unstructured data sources (e.g., Rx claims, clinical documents).
  • Enable continuous learning and model-retraining workflows using Vertex AI, Kubeflow, or similar GCP-native tooling.
  • Implement frameworks for observability and data quality, ensuring ML predictions, confidence scores, and fallback events are logged into data lakes or monitoring systems.
  • Support distributed data systems and ensure reliability, performance, and scalability of data infrastructure.
  • Design and provision cloud infrastructure using Infrastructure as Code (IaC) tools such as Terraform or Pulumi for GCP resources including GKE clusters, Cloud SQL, VPC networks, IAM, and storage.
  • Deploy, configure, and manage containerized data workloads using Kubernetes (GKE) — including deployments, autoscaling (HPA/VPA), namespaces, resource quotas, and health checks.
  • Architect and maintain network topology for data platform environments — VPCs, subnets, firewall rules, private service access, Cloud NAT, and VPC Service Controls.
  • Implement and enforce IAM policies, service account governance, and secrets management (GCP Secret Manager or HashiCorp Vault) to ensure least-privilege access across all data services.
  • Build and maintain infrastructure monitoring and alerting using Cloud Monitoring, Prometheus, Grafana, or equivalent — covering pipeline latency, throughput, error rates, and resource utilization.
  • Establish and maintain CI/CD pipelines for infrastructure changes using Terraform Cloud, GitHub Actions, or Argo CD, ensuring infrastructure drift detection and rollback capability.
  • Manage environment parity (dev/staging/prod) for data platform infrastructure, including environment-specific configuration management and promotion workflows.
  • Drive cloud cost governance — right-sizing compute resources, implementing committed-use discounts, setting up budget alerts, and producing cost attribution reports per workload.
  • Design and implement disaster recovery, backup, and high-availability strategies for data stores, pipeline infrastructure, and ML serving endpoints.
  • Collaborate with Security and Platform teams to ensure data infrastructure compliance with enterprise security policies, SOC 2, HIPAA, and CVS Health regulatory requirements.

Required Qualifications

  • 5+ years of experience building data pipelines or backend data workflows using Python, Java, or similar languages.
  • 2+ years of experience designing REST/GraphQL data services or integrating data APIs.
  • Hands-on experience working with ML/AI model integration in production (e.g., Vertex AI Endpoints, TensorFlow Serving, ML REST APIs).
  • Experience handling structured and unstructured datasets, including healthcare data (Rx claims, clinical documents, NLP text).
  • Familiarity with the end-to-end ML lifecycle: data ingestion, feature engineering, training, deployment, and real-time inference.
  • 2+ years of experience with cloud platforms (GCP preferred; AWS or Azure acceptable).
  • 2+ years working with streaming platforms like Kafka or equivalent.
  • 2+ years of experience with databases (Postgres or similar relational systems).
  • 2+ years of experience with CI/CD tools (GitHub Actions, Jenkins, Argo CD, etc.).
  • 2+ years of hands-on experience with Infrastructure as Code tools (Terraform preferred; Pulumi or CDK acceptable).
  • 2+ years managing containerized workloads using Kubernetes (GKE, EKS, or AKS) — deploying services, configuring autoscaling, and managing resource limits.
  • Solid understanding of cloud networking fundamentals: VPCs, subnets, firewall rules, private connectivity, and DNS resolution in GCP or AWS.
  • Experience designing IAM roles, service accounts, and secrets management workflows to enforce least-privilege access across data services.
  • Familiarity with infrastructure monitoring and alerting tools (Cloud Monitoring, Prometheus/Grafana, or equivalent).

Preferred Qualifications

  • Direct, hands-on experience with Google Cloud Platform, especially BigQuery, Dataflow, GKE, Composer and Vertex AI.
  • Knowledge of Kubernetes concepts and experience running data services or pipelines on GKE.
  • Strong understanding of distributed systems, microservice patterns, and data-centric system design.
  • Experience using Vertex AI, Kubeflow, or other ML orchestration platforms for model training and serving.
  • Knowledge of GenAI pipelines, LLM prompt workflows, and agent orchestration frameworks (e.g., LangChain, transformers).
  • Experience deploying Python-based ML/NLP services into microservice ecosystems using REST, gRPC, or sidecar architectures.
  • Domain experience in healthcare, claim adjudication, or Rx data processing.
  • Experience with Terraform modules, workspaces, and remote state backends for managing multi-environment GCP infrastructure.
  • Familiarity with GCP Shared VPC, VPC Service Controls, or private Google Access configurations for secure data platform networking.
  • Exposure to FinOps practices — cloud cost attribution, showback/chargeback models, and resource tagging strategies.
  • Experience with GitOps workflows using Argo CD or Flux for managing infrastructure and application delivery.

Education

  • Bachelor's degree in Computer Science, Data Engineering, Information Systems, or equivalent experience (High School Diploma + 4 years of relevant experience acceptable).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer - AI/ML Engineer
Data Engineer - AI/ML Engineer

Mployee.me • Pune District

Hybrid
INR 1,200,000 - 1,500,000
Google Cloud Data Engineer
Google Cloud Data Engineer

Vation Digital Pvt. Ltd. • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Engineering Manager – AWS to GCP Data Migration, AI/ML & GenAI
Engineering Manager – AWS to GCP Data Migration, AI/ML & GenAI

Naveera Technology LLC • India

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineer (GCP)
Senior Data Engineer (GCP)

Quantiphi • Bengaluru

On-site
INR 350,000 - 800,000
Data Scientist
Data Scientist

Omnicom Global Solutions • Bengaluru Urban

Hybrid
INR 1,500,000 - 4,000,000
Data Engineer- GCP/Databricks
Data Engineer- GCP/Databricks

Zohorecruit • Bengaluru

Hybrid
INR 2,400,000 - 4,800,000
Data Engineer
Data Engineer

LS Digital Group • Navi Mumbai

On-site
INR 1,200,000 - 1,600,000
AI/ML engineer with GCP Cloud
AI/ML engineer with GCP Cloud

Infosys • Dadri

On-site
INR 1,200,000 - 1,800,000
Senior GCP Data Engineer
Senior GCP Data Engineer

Tata Consultancy Services • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineer - GCP
Senior Data Engineer - GCP

AuxoAI Inc. • Bengaluru

On-site
INR 1,200,000 - 2,000,000