4467- Software Development Engineer-III (Data Engineer) (Iceberg/Trino)

Innovaccer

Dadri

On-site

INR 1,800,000 - 3,200,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Generous leave
Parental leave
Sabbatical
Health insurance
Care vouchers
Salary advances

Job summary

Innovaccer is seeking a Senior Software Engineer on the Lakehouse team to build and operate data pipelines at the heart of our on-premise platform. You will work end-to-end on Spark ingestion to Iceberg, and Trino transformations powering analytics and applications.

You will optimize validation, typing, deduplication, and serving table health, while collaborating across teams. Requirements include 5+ years of data engineering, strong SQL, Spark, and experience with Iceberg or equivalent formats.

Qualifications

  • 5+ years of data engineering experience building production pipelines at scale.
  • Strong SQL skills and hands-on experience with Apache Spark for batch processing.
  • Experience with Trino/Presto and open table formats (Iceberg preferred, Delta Lake or Hudi acceptable).
  • Working knowledge of S3-compatible object storage and columnar file formats.

Responsibilities

  • Build Spark ingestion jobs that land high-volume raw files into Iceberg tables with schema handling, bad-record quarantine, and idempotent batch replay.
  • Develop and operate Trino SQL transform pipelines across data layers: validation and typing, business-rule transforms, MERGE-based deduplication, and aggregate builds.
  • Port existing warehouse SQL workloads to Trino and Spark SQL dialects, and validate results against source outputs.
  • Automate Iceberg table maintenance: compaction, snapshot expiry, and orphan-file cleanup as scheduled workflows.
  • Tune query and pipeline performance: partitioning strategy, file sizing, statistics, and resource-group behavior.
  • Instrument pipelines with data-quality checks, reconciliation reports, and alerting, and participate in per-dataset validation during rollout phases.

Skills

SQL
Apache Spark
Trino/Presto
Python
Java
Airflow
CI/CD
HL7/CCDA

Education

B.E./B.Tech./M.Sc. in Computer Science or related

Tools

Apache Iceberg
Parquet
Delta Lake / Hudi

Job description

Engineering at Innovaccer

With every line of code, we accelerate our customers' success, turning complex challenges into innovative solutions. Collaboratively, we transform each data point we gather into valuable insights for our customers. Join us and be part of a team that's turning dreams of better healthcare into reality, one line of code at a time. Together, we're shaping the future and making a meaningful impact on the world.

About the Role

As a Senior Software Engineer on the Lakehouse team, you will build and operate the data pipelines at the heart of Innovaccer's on-premise platform: Spark ingestion jobs landing raw healthcare data into Apache Iceberg, and Trino SQL transforms building the layered tables that power analytics and applications. You will work hands-on across the full pipeline surface, from file validation and quarantine at ingestion to query performance and table health in serving.

A Day in the Life
  • Build Spark ingestion jobs that land high-volume raw files into Iceberg tables with schema handling, bad-record quarantine, and idempotent batch replay.
  • Develop and operate Trino SQL transform pipelines across data layers: validation and typing, business-rule transforms, MERGE-based deduplication, and aggregate builds.
  • Port existing warehouse SQL workloads to Trino and Spark SQL dialects, and validate results against source outputs.
  • Automate Iceberg table maintenance: compaction, snapshot expiry, and orphan-file cleanup as scheduled workflows.
  • Tune query and pipeline performance: partitioning strategy, file sizing, statistics, and resource-group behavior.
  • Instrument pipelines with data-quality checks, reconciliation reports, and alerting, and participate in per-dataset validation during rollout phases.
  • B.E., B.Tech., M.Sc. degree in Computer Science or a related technical field.
  • 5+ years of data engineering experience building production pipelines at scale.
  • Strong SQL skills and hands-on experience with Apache Spark for batch processing.
  • Experience with Trino/Presto (or a comparable distributed SQL engine) and open table formats: Iceberg preferred, Delta Lake or Hudi acceptable.
  • Working knowledge of S3-compatible object storage and columnar file formats (Parquet).
  • Experience with workflow orchestration tools (Airflow or equivalent) and CI/CD for data pipelines.
  • Professional development experience with Python and/or Java.
  • Healthcare data formats (HL7, CCDA, claims files) and regulated-environment experience are pluses.
Here’s What We Offer
  • Generous Leaves: Enjoy generous leave benefits of up to 40 days.
  • Parental Leave: Leverage one of industry's best parental leave policies to spend time with your new addition.
  • Sabbatical: Want to focus on skill development, pursue an academic career, or just take a break? We've got you covered.
  • Health Insurance: We offer comprehensive health insurance to support you and your family, covering medical expenses related to illness, disease, or injury. Extending support to the family members who matter most.
  • Care Program: Whether it’s a celebration or a time of need, we’ve got you covered with care vouchers to mark major life events. Through our Care Vouchers program, employees receive thoughtful gestures for significant personal milestones and moments of need.
  • Financial Assistance: Life happens, and when it does, we’re here to help. Our financial assistance policy offers support through salary advances and personal loans for genuine personal needs, ensuring help is there when you need it most.

Innovaccer is an equal-opportunity employer. We celebrate diversity, and we are committed to fostering an inclusive and diverse workplace where all employees, regardless of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, marital status, or veteran status, feel valued and empowered.

About Innovaccer

Innovaccer builds software that helps hospitals, clinics, and health insurance companies make sense of all the scattered data they deal with every day — patient records, appointments, insurance claims, lab results — and turns it into something useful and actionable.

Think of a typical hospital: patient information often lives in a dozen different systems that don't talk to each other well. This causes delays in scheduling, slows down doctors, and makes it harder to catch health problems early. Innovaccer's platform connects all of that data together, and increasingly, uses AI "agents" to actually do some of the manual work for healthcare teams — things like scheduling appointments, preparing insurance paperwork, or drafting clinical notes — so that staff can spend more time with patients and less time on admin work.

Some of the largest healthcare systems in the US — including CommonSpirit Health, Atlantic Health, and Banner Health — use Innovaccer's platform today.

The company is also investing heavily in this direction: in June 2026, Innovaccer signed a multi-year partnership with AWS to run its AI agents at a much larger scale, using AWS's cloud AI tools (Amazon Bedrock) and healthcare-specific data infrastructure (AWS HealthLake). This is a strong signal of where the platform — and this engineering team — is headed next. For more information, visit www.innovaccer.com and check us out on YouTube, Glassdoor, LinkedIn, Instagram, and the Web.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

4467- Software Development Engineer-III Data Engineer Iceberg Trino
4467- Software Development Engineer-III Data Engineer Iceberg Trino

Innovaccer • Dadri

On-site
INR 1,400,000 - 2,100,000
Generous leave benefits
Parental leave
Sabbatical
+3
4467- Software Development Engineer-III (Data Engineer) (Iceberg/Trino)
4467- Software Development Engineer-III (Data Engineer) (Iceberg/Trino)

Innovaccer Inc. • India

On-site
INR 1,800,000 - 3,000,000
Generous leaves
Parental leave
Sabbatical
+3
Senior Software Engineer
Senior Software Engineer

Innovaccer • Dadri

On-site
INR 1,800,000 - 2,800,000
Health insurance
Sabbatical policy
Pet-friendly office
+2
4350- Staff Engineer - Backend (Gravity)
4350- Staff Engineer - Backend (Gravity)

Innovaccer Inc. • Dadri

On-site
INR 4,000,000 - 7,000,000
Leaves up to 40 days
Parental leave
Sabbatical
+3
4195- Staff Engineer - Healthcare Interoperability (Gravity)
4195- Staff Engineer - Healthcare Interoperability (Gravity)

Innovaccer • Dadri

On-site
INR 4,000,000 - 6,500,000
Health insurance
Generous leaves up to 40 days
Parental leave
+3
4195- Staff Engineer - Backend (Gravity)
4195- Staff Engineer - Backend (Gravity)

Innovaccer Analytics • Dadri

On-site
INR 4,000,000 - 8,000,000
40 days leave
Parental leave
Sabbatical
+3
4403 - Software Development Engineer-II (Finance - AI Engineering)
4403 - Software Development Engineer-II (Finance - AI Engineering)

Innovaccer • Dadri

On-site
INR 4,000,000 - 9,000,000
Generous leaves
Parental leave
Sabbatical
+3
Software Development Engineer-II (Frontend)
Software Development Engineer-II (Frontend)

Innovaccer Analytics • Dadri

On-site
INR 1,500,000 - 2,100,000
Generous Leave
Parental Leave
Sabbatical Leave
+3
4448 -Associate Director- Implementation Engineering
4448 -Associate Director- Implementation Engineering

Innovaccer • Dadri

On-site
INR 1,800,000 - 3,200,000
Generous leaves
Parental leave
Sabbatical
+3
0001-Vice President, Implementation Engineering
0001-Vice President, Implementation Engineering

Innovaccer • Dadri

On-site
INR 3,000,000 - 6,000,000
Leaves up to 40 days
Parental leave policy
Sabbatical for skill development
+3