New
New
Senior Data Engineer - Databricks & Streaming - Healthcare AI (Onsite, Evening Shift, Lahore, PKR Salary)
Data EngineerData
0 views0 saves0 applied
Quick Summary
Key Responsibilities
Build and maintain resilient ingestion pipelines for third-party vendor REST APIs. Work primarily with voice-AI observability and telephony platforms. Handle different pagination schemes,
Requirements Summary
4+ years of experience in data engineering, with substantial production experience in Databricks. Strong experience with Spark SQL, PySpark, Delta Lake, Medallion Architecture,
Technical Tools
Data EngineerData
Requirements
~2 min read- 4+ years of experience in data engineering, with substantial production experience in Databricks.
- Strong experience with Spark SQL, PySpark, Delta Lake, Medallion Architecture, and Delta Live Tables (DLT).
- Hands-on experience with Structured Streaming or equivalent production-grade streaming ingestion using Azure Event Hubs, Kafka, or Kinesis.
- Strong understanding of checkpoint recovery, watermarking, and deduplication strategies.
- Demonstrated experience debugging source-to-warehouse data discrepancies.
- Ability to walk through a real-world incident involving mismatched record counts and explain how the root cause was identified and resolved.
- Proven experience integrating third-party REST APIs in production.
- Experience handling pagination edge cases, rate and row limits, retries, and schema drift.
- Experience with entity resolution or data matching involving messy, real-world text data.
- Experience with a metrics or semantic layer such as Holistics AML/AQL, dbt Metrics, or LookML.
- Working understanding of why non-additive measures cannot be reliably calculated from pre-aggregated rollups.
- Strong SQL and Python skills, with the ability to own data pipelines end-to-end with minimal oversight.
- Strong written and spoken English, with the ability to collaborate effectively with a US-based team asynchronously.
- Healthcare data experience, including referrals, payer taxonomy, claims/eligibility, or other PHI-adjacent datasets.
- Familiarity with HIPAA handling expectations.
- Experience with voice-agent, call-center, or telephony/conversation data.
- Familiarity with call transcripts and containment or outcome metrics.
- Hands-on experience with Holistics, specifically AML/AQL modeling.
- Experience with the broader Azure ecosystem beyond Event Hubs, including ADLS, ADF, and Key Vault.
Responsibilities
~2 min read- →Build and maintain resilient ingestion pipelines for third-party vendor REST APIs.
- →Work primarily with voice-AI observability and telephony platforms.
- →Handle different pagination schemes, including offset/limit and page/cursor models.
- →Manage row and rate limits through time-windowing and adaptive bisection.
- →Implement robust schema-drift handling through contract and column-presence checks.
- →Ensure alerts are triggered when fields are renamed, moved, or removed rather than silently propagating null values.
- →Own streaming ingestion from Azure Event Hubs into Databricks using Structured Streaming and/or Auto Loader.
- →Manage checkpoints and offsets, watermarking, and at-least-once deduplication.
- →Perform source-parity reconciliation across ingested and production data.
- →Investigate row counts, dropped or duplicated events, late-arriving data, and schema mismatches.
- →Identify and resolve the root cause when ingested data does not match production sources.
- →Develop and maintain Delta Lake pipelines using a Medallion Architecture (Bronze Silver Gold).
- →Use Spark SQL and PySpark to build and maintain production data pipelines.
- →Implement idempotent MERGE upserts.
- →Work with Delta Live Tables and materialized-view constraints, including CREATE OR REFRESH and LIVE references.
- →Understand and manage differences between DLT and job execution contexts.
- →Build entity-resolution pipelines for dirty, free-text data.
- →Normalize practice, provider, and payer names using regex, canonical dictionaries, fuzzy matching, confidence-scored crosswalks, and override tables.
- →Maintain the semantic and metrics layer with rigorous metric definitions.
- →Define and maintain accurate denominators, data grain, and cohort boundaries.
- →Ensure the correct handling of non-additive aggregates, including medians and percentiles that cannot be reliably supported through aggregate-aware pre-aggregation.
- →Ensure every metric remains accurate and reproducible.
- →Instrument data quality across the entire pipeline.
- →Monitor data freshness, source parity, data contracts, and other critical quality checks.
- →Build alerting mechanisms that identify data issues before they reach dashboards.
Location & Eligibility
Where is the job
Lahore, Pakistan
On-site at the office
Who can apply
PK
Listing Details
- First seen
- September 19, 2026
- Last seen
- September 19, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 51%
- Scored at
- September 19, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application · ~5 min on hr-pod-hiring-talent-globally's site
Please let hr-pod-hiring-talent-globally know you found this job on Jobera.
3 other jobs at hr-pod-hiring-talent-globally
View all →Explore open roles at hr-pod-hiring-talent-globally.
Similar Data Engineer jobs
View all →C
CompassxRemoteSenior Data Engineer (IICS & Redshift) (Independent Contractor)
USD 60–75
ContractRemote
Graduate Data Engineer
Data Engineer
Remote
Senior Lead Data Engineer
USD 37-61
Remote
D
DoiTRemoteData Engineer - Cloud & SaaS Integrations
Remote
Re-advertisement: Digital Impact Specialist (Data Engineer), P-3, Digital Impact Division, Digital Foundations - Technology, Valencia, Spain, #00137299, Temporary Position, 364 days
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.