Quick Summary
Design, build, and deploy robust ETL/ELT pipelines within the lakehouse platform (Google BigQuery or Databricks) using SQL, Python, PySpark, and Spark SQL.
Education: Minimum of a bachelor’s degree in computer science, Engineering, Mathematics, or a related technical field preferred. Experience: 8+ years of relevant experience in data engineering,
Why we're hiring:
We are seeking a highly skilled and experienced Senior Data Engineer to join our growing data team. In this critical role, you will be instrumental in designing, building, and optimizing our scalable data lakehouse platform using Google BigQuery or Databricks. You will be a key player in developing robust data pipelines that ingest data from various sources, including Google Analytics 4 (GA4), and transform it into reliable, analysis-ready datasets within the lakehouse environment. This role requires deep expertise in modern lakehouse platforms – Google BigQuery and/or Databricks – together with strong skills in SQL, Python, and Apache Spark (PySpark), along with strong hands-on experience across Azure, AWS, and GCP cloud environments, as our data ecosystem spans multiple cloud platforms. You will be responsible for the entire data lifecycle within the lakehouse, from ingestion and transformation to governance and optimization, ensuring data quality and performance. You should be adept at analyzing performance bottlenecks in Spark jobs and BigQuery workloads, providing enhancement recommendations, and collaborating effectively with both technical and non-technical stakeholders.
What you'll be doing:
- Design, build, and deploy robust ETL/ELT pipelines within the lakehouse platform (Google BigQuery or Databricks) using SQL, Python, PySpark, and Spark SQL.
- Implement and manage the Medallion Architecture (Bronze, Silver, Gold layers) using Delta Lake or BigQuery datasets to ensure data quality and progressive data refinement.
- Leverage native ingestion tooling – such as BigQuery Data Transfer Service, Pub/Sub streaming, or Databricks Auto Loader – for efficient, scalable, and incremental ingestion of data from sources like GA4 into the Bronze layer.
- Develop, schedule, and monitor complex, multi-task data workflows using Cloud Composer (Airflow), BigQuery scheduled queries, or Databricks Workflows.
- Optimize BigQuery tables (partitioning, clustering, materialised views) and Spark jobs / Delta Lake tables (using techniques like OPTIMIZE, Z-ORDER, and partitioning) for high performance and cost efficiency.
- Implement data governance, security, and discovery using Dataplex / BigQuery policy tags or Unity Catalog, including managing access controls and data lineage.
- Write complex, customized SQL queries to manipulate data and support ad-hoc analytical requests from business teams.
- Develop strategies for data ingestion from multiple sources, using various techniques including streaming, API consumption, and replication.
- Document data engineering processes, data models, and technical specifications for the lakehouse platform.
- Conform to agile development practices, including version control (Git), continuous integration/delivery (CI/CD), and test-driven development.
- Provide production support for data pipelines, actively monitoring and resolving issues to ensure the continuous flow of critical data.
- Collaborate with analytics and business teams to understand data requirements and deliver well-modelled, performant datasets in the gold layer
What you'll need:
- Education: Minimum of a bachelor’s degree in computer science, Engineering, Mathematics, or a related technical field preferred.
- Experience: 8+ years of relevant experience in data engineering, with a significant focus on building data pipelines on distributed systems.
- Lakehouse Platform Expertise (Google BigQuery and/or Databricks)
- BigQuery: Deep, hands-on experience with BigQuery architecture, including partitioning, clustering, materialised views, slot/cost optimisation, and diagnosing query performance using query plans and INFORMATION_SCHEMA.
- Apache Spark / Delta Lake: Strong experience with Spark architecture, writing and optimising PySpark and Spark SQL jobs, and building reliable pipelines on Delta Lake. Proficient with ACID transactions, time travel, schema evolution, and DML operations (MERGE, UPDATE, DELETE).
- Data Ingestion: Experience with modern ingestion tools, such as BigQuery Data Transfer Service, Pub/Sub / Dataflow streaming, Databricks Auto Loader, and COPY INTO for scalable file processing.
- Data Governance: Strong understanding of data governance concepts and practical experience implementing security, lineage, and discovery using Dataplex, BigQuery IAM and policy tags, or Unity Catalog.
- Programming: 5+ years of strong, hands-on experience in Python, with an emphasis on PySpark for large-scale data transformation.
- SQL: 6+ years of advanced SQL experience, including complex joins, window functions, and CTEs.
- Cloud Platforms: 8+ years of strong, hands-on experience working across all three major cloud platforms – Azure, AWS, and GCP – including expertise in cloud storage (ADLS Gen2, S3, Google Cloud Storage), security and identity management (Azure AD/Entra ID, AWS IAM, GCP IAM), and cloud networking. Proven ability to design, deploy, and support data solutions in multi-cloud environments.
- Data Modeling: Experience designing star schemas and applying data warehouse methodologies to build analytical models (Gold layer).
- CI/CD & DevOps: Hands-on experience with version control (Git) and CI/CD pipelines (e.g., GitHub Actions, Azure DevOps) for automating the deployment of BigQuery and Databricks assets.
- Primary Data Platform: Google BigQuery or Databricks
- Cloud Platforms: Azure, AWS, and GCP (strong hands-on experience across all three required)
- Data Warehouses (Integration): Snowflake
- Orchestration/Transformation: Cloud Composer (Airflow), Databricks Workflows, dbt (data build tool)
- Version Control: Git/GitHub or similar repositories
- Infrastructure as Code (Bonus): Terraform
- BI Tools (Bonus): Looker or Power BI
Who you are:
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- September 8, 2026
- First seen
- September 8, 2026
- Last seen
- September 8, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 67%
- Scored at
- September 8, 2026
Signal breakdown
Please let Wpp know you found this job on Jobera.
3 other jobs at Wpp
View all →Explore open roles at Wpp.
Similar Data Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.
