Wpp
Wpp13h ago
New

Senior Data Engineer

IndiaIndiasenior
Data EngineerData
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Design, build, and deploy robust ETL/ELT pipelines within the lakehouse platform (Google BigQuery or Databricks) using SQL, Python, PySpark, and Spark SQL.

Requirements Summary

Education: Minimum of a bachelor’s degree in computer science, Engineering, Mathematics, or a related technical field preferred. Experience: 8+ years of relevant experience in data engineering,

Technical Tools
Data EngineerData

Why we're hiring:

We are seeking a highly skilled and experienced Senior Data Engineer to join our growing data team. In this critical role, you will be instrumental in designing, building, and optimizing our scalable data lakehouse platform using Google BigQuery or Databricks. You will be a key player in developing robust data pipelines that ingest data from various sources, including Google Analytics 4 (GA4), and transform it into reliable, analysis-ready datasets within the lakehouse environment. This role requires deep expertise in modern lakehouse platforms – Google BigQuery and/or Databricks – together with strong skills in SQL, Python, and Apache Spark (PySpark), along with strong hands-on experience across Azure, AWS, and GCP cloud environments, as our data ecosystem spans multiple cloud platforms. You will be responsible for the entire data lifecycle within the lakehouse, from ingestion and transformation to governance and optimization, ensuring data quality and performance. You should be adept at analyzing performance bottlenecks in Spark jobs and BigQuery workloads, providing enhancement recommendations, and collaborating effectively with both technical and non-technical stakeholders.

What you'll be doing:

  • Design, build, and deploy robust ETL/ELT pipelines within the lakehouse platform (Google BigQuery or Databricks) using SQL, Python, PySpark, and Spark SQL.
  • Implement and manage the Medallion Architecture (Bronze, Silver, Gold layers) using Delta Lake or BigQuery datasets to ensure data quality and progressive data refinement.
  • Leverage native ingestion tooling – such as BigQuery Data Transfer Service, Pub/Sub streaming, or Databricks Auto Loader – for efficient, scalable, and incremental ingestion of data from sources like GA4 into the Bronze layer.
  • Develop, schedule, and monitor complex, multi-task data workflows using Cloud Composer (Airflow), BigQuery scheduled queries, or Databricks Workflows.
  • Optimize BigQuery tables (partitioning, clustering, materialised views) and Spark jobs / Delta Lake tables (using techniques like OPTIMIZE, Z-ORDER, and partitioning) for high performance and cost efficiency.
  • Implement data governance, security, and discovery using Dataplex / BigQuery policy tags or Unity Catalog, including managing access controls and data lineage.
  • Write complex, customized SQL queries to manipulate data and support ad-hoc analytical requests from business teams.
  • Develop strategies for data ingestion from multiple sources, using various techniques including streaming, API consumption, and replication.
  • Document data engineering processes, data models, and technical specifications for the lakehouse platform.
  • Conform to agile development practices, including version control (Git), continuous integration/delivery (CI/CD), and test-driven development.
  • Provide production support for data pipelines, actively monitoring and resolving issues to ensure the continuous flow of critical data.
  • Collaborate with analytics and business teams to understand data requirements and deliver well-modelled, performant datasets in the gold layer

What you'll need:

  • Education: Minimum of a bachelor’s degree in computer science, Engineering, Mathematics, or a related technical field preferred.
  • Experience: 8+ years of relevant experience in data engineering, with a significant focus on building data pipelines on distributed systems.
  • Lakehouse Platform Expertise (Google BigQuery and/or Databricks)
  • BigQuery: Deep, hands-on experience with BigQuery architecture, including partitioning, clustering, materialised views, slot/cost optimisation, and diagnosing query performance using query plans and INFORMATION_SCHEMA.
  • Apache Spark / Delta Lake: Strong experience with Spark architecture, writing and optimising PySpark and Spark SQL jobs, and building reliable pipelines on Delta Lake. Proficient with ACID transactions, time travel, schema evolution, and DML operations (MERGE, UPDATE, DELETE).
  • Data Ingestion: Experience with modern ingestion tools, such as BigQuery Data Transfer Service, Pub/Sub / Dataflow streaming, Databricks Auto Loader, and COPY INTO for scalable file processing.
  • Data Governance: Strong understanding of data governance concepts and practical experience implementing security, lineage, and discovery using Dataplex, BigQuery IAM and policy tags, or Unity Catalog.
  • Programming: 5+ years of strong, hands-on experience in Python, with an emphasis on PySpark for large-scale data transformation.
  • SQL: 6+ years of advanced SQL experience, including complex joins, window functions, and CTEs.
  • Cloud Platforms: 8+ years of strong, hands-on experience working across all three major cloud platforms – Azure, AWS, and GCP – including expertise in cloud storage (ADLS Gen2, S3, Google Cloud Storage), security and identity management (Azure AD/Entra ID, AWS IAM, GCP IAM), and cloud networking. Proven ability to design, deploy, and support data solutions in multi-cloud environments.
  • Data Modeling: Experience designing star schemas and applying data warehouse methodologies to build analytical models (Gold layer).
  • CI/CD & DevOps: Hands-on experience with version control (Git) and CI/CD pipelines (e.g., GitHub Actions, Azure DevOps) for automating the deployment of BigQuery and Databricks assets.
  • Primary Data Platform: Google BigQuery or Databricks
  • Cloud Platforms: Azure, AWS, and GCP (strong hands-on experience across all three required)
  • Data Warehouses (Integration): Snowflake
  • Orchestration/Transformation: Cloud Composer (Airflow), Databricks Workflows, dbt (data build tool)
  • Version Control: Git/GitHub or similar repositories
  • Infrastructure as Code (Bonus): Terraform
  • BI Tools (Bonus): Looker or Power BI

Who you are:

What We Offer

~1 min read

Location & Eligibility

Where is the job
India
On-site within the country
Who can apply
IN

Listing Details

Posted
September 8, 2026
First seen
September 8, 2026
Last seen
September 8, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
67%
Scored at
September 8, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Wpp
Wpp
greenhouse
Employees
10,000+
Founded
1985
Domain
wpp.com
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Wpp Senior Data Engineer