Senior Data Engineer
Quick Summary
deployment, DAG packaging, scheduler and worker capacity, upgrades, observability, and incident response. Build paved paths for authoring, testing, deploying, and operating workflows,
Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.
In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.
What We Offer
~1 min readAbout the Role
~1 min readHandshake is hiring a Senior Data Platform Engineer to build and operate the infrastructure that moves data reliably across our career marketplace and AI products. On the Data and ML Platform team, you will own the systems that orchestrate batch workloads and deliver timely, trustworthy data to engineers, analysts, and product teams.
This is a hands-on technical leadership role with the opportunity to shape Handshake's data infrastructure strategy. You will set the technical direction for our workflow orchestration platform (Airflow), data ingestion, data pipelines, our data warehouse (BigQuery) and streaming data infrastructure. You'll collaborate closely with ML engineers, data scientists, and other engineers to maximize velocity on our data platform. This infrastructure will also support Handshake's growing AI development, from model and agent workflows to evaluation and production data access.
Responsibilities
~1 min read- →
Lead the technical direction and roadmap for our Airflow and Astronomer platform while owning its production operations: deployment, DAG packaging, scheduler and worker capacity, upgrades, observability, and incident response.
- →
Build paved paths for authoring, testing, deploying, and operating workflows, including reusable operators, CI checks, local and staging environments, and clear runbooks.
- →
Design and operate streaming and change data capture pipelines using tools such as Pub/Sub, Dataflow or Beam, and Datastream to serve analytics and product use cases.
- →
Make data delivery resilient to retries, duplicates, schema changes, late events, backfills, and replay; define useful latency, freshness, and reliability targets.
- →
Improve the cloud foundation behind data workloads with Kubernetes, Terraform, IAM, secrets, CI/CD, and cost-aware capacity management.
- →
Lead cross-team design decisions on data contracts, interfaces, and operational ownership; mentor engineers and align application, analytics, ML, and cloud partners on scalable platform approaches.
- →
Participate in on-call, troubleshoot production failures across systems, and turn incidents into durable improvements.
- →
Build dependable data and orchestration foundations for AI workloads, including batch inference, evaluation datasets, and agent-facing data.
Strong software engineering skills in Python and experience building and operating production data or distributed systems.
Deep hands-on Airflow experience beyond writing DAGs: scheduling and execution behavior, deployment, scaling, upgrades, monitoring, and debugging failures.
Experience with event-driven or streaming data infrastructure, including a message broker or managed event bus and a stream processing system.
Sound understanding of CDC, delivery guarantees, idempotency, ordering, schema evolution, and recovery or replay in production pipelines.
Experience with cloud infrastructure and infrastructure as code; comfortable working with containers, Kubernetes, CI/CD, access controls, and production observability.
A track record of leading ambiguous infrastructure initiatives: setting technical direction, making pragmatic architecture tradeoffs, and driving adoption across teams.
Clear communication and a track record of partnering across teams while owning systems through production support.
GCP services such as Pub/Sub, Dataflow, Datastream, BigQuery, GKE, and Cloud Storage.
Astronomer, Apache Beam, Terraform, Spacelift, Datadog, dbt, Spark or Dataproc.
Supporting both batch and low-latency consumers, including product-facing data services or ML features.
Experience supporting ML or AI workloads such as inference pipelines, reproducible evaluations, or governed data access.
Location & Eligibility
Listing Details
- Posted
- September 28, 2026
- First seen
- September 28, 2026
- Last seen
- September 29, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 57%
- Scored at
- September 29, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.