MyInternships.in
Quantiphi logo — Quantiphi Senior Data Engineer at Quantiphi
Quantiphi

Senior Data Engineer

Job · Full-timeIn OfficeMumbai3-5 years1 opening

About this role

Quantiphi is hiring for Senior Data Engineer in Mumbai. This opening was published by Quantiphi on their official careers board (Workday) on 11 May 2026 and was confirmed live on 14 September 2026. Job details • Company: Quantiphi • Role: Senior Data Engineer • Location: Mumbai (as listed: IN MH…

Similar jobs hiring now

Not quite right? These jobs in Mumbai match the same skills — apply to a few to improve your chances.

Skills you'll use

PythonJavaSQLExcelGitKotlinKafkaSparkSnowflakeCI/CDAgile

What you'll do

  • Build and maintain the ingestion pipelines — Apache Flink streaming jobs on Dataproc for HL7v2 and FHIR feeds, PySpark batch jobs for CCDA and CSV bulk loads. Implement against the CDM_ingest_mapping shared Python library that defines SourceSpec instances per source × format combination
  • Implement source-format parsers (HL7v2, CCDA, CSV, FHIR) as Python classes per the parser component spec. Write fixture-driven tests covering both well-formed inputs and DLQ-routing scenarios
  • Implement and maintain the synchronous Informatica MDM call pattern in the ingestion path — batched calls, timeout handling, circuit breaker behavior, DLQ routing for MDM failures. Implement the asynchronous MDM event consumer that applies ECI changes to CDM
  • Build the dbt transformation layer end to end — staging models per source, intermediate models that union and resolve REFs, CDM target models (DIM/FACT/BRIDGE/REF) that apply SCD2 via shared macros, and data product models. Write the dbt YAML schemas, tests, and documentation that accompany every model
  • Implement and maintain the shared dbt macro library — hash_key, scd2_merge, attribute_hash, restate_merge, audit_columns. Macros are the most-reused code; their correctness is non-negotiable and they require golden tests
  • Build the FHIR serialization layer — flat FHIR Iceberg tables (one per resource type) materialized via dbt, the PySpark bundling pipeline that produces FHIR Bundles for Kafka publication, and the FHIR validator integration that gates publication on US Core 6.1 conformance
  • Build and maintain FHIR-Repository integration components — the Java/Kotlin egress interceptor that captures client-originated FHIR changes, the Flink loopback consumer that merges those changes into CDM, the bundle consumer that ingests CDM-originated bundles into FHIR-Repository. Implement origin-tag-based loop prevention
  • Implement Cloud Composer DAGs to orchestrate dbt runs, batch ingestion jobs, maintenance operations (Iceberg compaction, snapshot expiration, orphan file cleanup), and data product refresh schedules
  • Work within the spec-driven development framework — draft unit specs for new components, work with peers on spec review, generate implementation and tests using Code agents with the spec as primary context, iterate until tests pass, and submit code review packages that include the spec, tests, and implementation together
  • Implement and monitor data quality checks at every layer — DBT tests for staging and CDM, FHIR validator output for serialization, Iceberg metadata observations for storage health, freshness monitors at the source-to-CDM boundary

Who can apply

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related quantitative field
  • 3+ years of hands-on data engineering experience
  • Strong proficiency in Python and SQL. PySpark and PyFlink familiarity strongly preferred. Some Java or Kotlin exposure useful for FHIR-Repository interceptor work (one or two engineers on the team will lead this; the rest contribute as needed)
  • Hands-on experience with Google Cloud Platform — Cloud Storage, Dataproc, Cloud Composer, Cloud Run or GKE for container workloads, Secret Manager, IAM. Experience with the Dataproc Flink optional component is a strong plus
  • Production experience with dbt — incremental materialization strategies, custom macros, tests, sources, sequencing, and project organization for large model graphs. dbt-trino adapter experience is a plus
  • Production experience with Apache Iceberg — table creation, partitioning, compaction, snapshot operations, schema evolution. Familiarity with reading and writing Iceberg from multiple engines (Spark, Flink, Trino) is valuable
  • Experience with Apache Kafka — producers, consumers, partitioning, consumer-group semantics, retention and compaction, and integration with stream processors. Confluent Cloud experience preferred
  • Experience with streaming data processing — Apache Flink in production preferred; Apache Spark Structured Streaming acceptable as adjacent experience
  • Familiarity with healthcare data standards — at minimum, HL7v2 message structure and FHIR R4 resource shapes. Hands-on parsing experience for one or both is preferred
  • Experience with version control (Git), branch-based development workflows, pull request reviews, and CI/CD pipelines (GitHub Actions, GitLab CI, or Cloud Build)

About Quantiphi

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth. If working in an environment that encourages you to innovate…

Searching for “Quantiphi jobs” or “jobs at Quantiphi”? You’re in the right place.

Sponsored — Deals for Professionals

Sponsored

Similar Jobs

Hand-picked roles that match this job's skills

Sponsored — Deals for Professionals

Sponsored

Similar Jobs Based on Your Skills

Recently Posted Jobs in Mumbai

Fresh jobs posted in Mumbai — apply early

Ready to apply for Senior Data Engineer?

Free to apply · takes under 2 minutes · Quantiphi reviews on a rolling basis.

Apply Now