Lead Data Engineer
Quick Summary
Design, build, and optimize production-grade batch and streaming data pipelines in a multi-cloud (GCP, AWS, Azure) ecosystem to unify enterprise data assets into a high-performance AI, analytics,
Are you interested in making a world of difference in cancer care?
Cancer strikes more than 10 million people worldwide each year. As the leading medical society representing doctors who care for people with cancer, the American Society of Clinical Oncology (ASCO) is committed to conquering cancer through research, education, and promotion of the highest quality care.
What We Offer
~2 min readThis position is hybrid with a primary location at our headquarters in Alexandria, VA. The hire must reside within 75 miles of our headquarters. We anticipate the hire to be onsite approximately 1-2 days per week.
This is an exempt position. The hiring salary range for this position is $164,000-$200,000 annualized.
The range displayed reflects the minimum and maximum annualized salary ASCO expects to provide for a new hire for the position across the U.S. We thoughtfully determine and design offers based on the selected candidate’s relevant experience and qualifications.
Responsibilities
~2 min read- →Architect & Deploy Scalable Data Pipelines/Products: Design, build, and optimize production-grade batch and streaming data pipelines in a multi-cloud (GCP, AWS, Azure) ecosystem to unify enterprise data assets into a high-performance AI, analytics, and operational foundation.
- →Simple to Complex Data Pipelines: Design, build, and optimize scalable batch and streaming data ingestion pipelines for both structured and unstructured data.
- →Trusted Context Foundation: Ensure all developed solutions meet high standards for security, quality, reliability, explainability, and maintainability.
- →Engineer Agentic AI & RAG Workflows: Build and operationalize modern AI capabilities—including RAG, semantic search, vector databases, and autonomous multi-agent workflows, grounded in clean enterprise context.
- →Establish MLOps & LLMOps Rigor: Build CI/CD, containerization, and observability frameworks to streamline LLM service integrations for fast, reliable, and cost-effective application delivery.
- →Write Production-Grade Code & APIs: Deliver clean, maintainable Python code and microservices that integrate data products directly into AI agents, business applications, and operational workflows.
- →Optimize Platform Cost & Performance: Continually refine database designs, storage tiers, and compute workloads to maximize query speed while actively optimizing cloud operating costs (OpEx).
- →Translate Strategy into Execution: Partner with product leads and business stakeholders to translate mission goals into scalable technical solutions that accelerate the digital roadmap.
- →Lead Through Code & Technical Mentorship: Set engineering standards, drive rigorous code reviews, and elevate team capability by actively building alongside junior and mid-level data engineers.
- →Ensure Operational Reliability & Incident Health: Perform root-cause analysis and rapid remediation for complex data platform incidents to maintain continuous system availability and data integrity.
- Bachelor’s degree in Computer Science, Data Engineering, Software Engineering, Applied Mathematics, or a related technical field (or equivalent practical experience).
- 9+ years of hands-on experience architecting, building, and maintaining enterprise-grade batch and streaming data pipelines in modern cloud environments.
- 2+ years of hands-on experience building and deploying production-grade AI/ML solutions, including RAG architectures, vector databases, LLM integrations, and agentic workflows.
- Deep mastery of Python and modern data/software engineering practices.
- Proven expertise with cloud-native data architectures (GCP/BigQuery, AWS, or Azure), distributed computing, enterprise data modeling, and visualizing end-to-end integration flows.
- Direct experience establishing MLOps/LLMOps pipelines, containerization (Docker), real-time evaluation frameworks, and automated CI/CD deployments.
- Demonstrated ability to tune physical data models and compute workloads to increase system performance while controlling cloud operating costs (OpEx).
- Proven ability to drive technical strategy, establish coding standards, and mentor engineering teams through direct, active code contribution.
Nice to Have
~1 min read- Hands-on exposure to specialized AI orchestration tooling (e.g., Google Vertex AI Agent Builder).
- Master’s degree in Computer Science, Data Science, or Artificial Intelligence.
- Ability to design, build, and evolve multi-cloud data architectures that balance high performance with operational cost efficiency (OpEx optimization), ensuring data platforms scale seamlessly with enterprise growth.
- Hands-on expertise in moving modern AI concepts (agentic workflows, RAG, and LLM integrations) from experimental concepts into reliable, secure, high-availability operational environments.
- A "lead from the code" mindset that elevates team capability, enforces modern data engineering practices, and establishes high standards for code quality, documentation, and maintainability.
- Strategic ability to connect complex data architectures directly to measurable enterprise value, translating business objectives into high value technical initiatives.
- Deep commitment to building a system of interconnected data sources with built-in observability and improvement mechanisms by embedding data quality, automated governance, security, and real-time observability across all pipeline workflows.
- Strong collaborative approach that bridges the gap between raw data infrastructure, product leads, software engineers, and business stakeholders to accelerate the digital product roadmap.
Requirements
~1 min readExtended periods seated or standing at a desk.
High use of computer and other office technology equipment.
1-5 days/yr
Location & Eligibility
Listing Details
- Posted
- August 19, 2026
- First seen
- August 20, 2026
- Last seen
- August 20, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 49%
- Scored at
- August 20, 2026
Signal breakdown
Please let asco know you found this job on Jobera.
3 other jobs at asco
View all →Explore open roles at asco.
Similar Data Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.