5h ago
↻ Repost

Senior Data Engineer (Web Scraping)

IndiaIndiaRemoteFull-timesenior
Data EngineerData
2 views0 saves0 applied

Quick Summary

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer (Web Scraping) based in India. This is a senior,

Technical Tools
Data EngineerData

This is a senior, hands-on opportunity to own the next generation of reliable web-data acquisition within a modern data platform.
You will design, build, and operate production-grade Python scrapers and scalable ingestion pipelines that support trusted research and data products.
The role focuses on improving the maturity, reliability, and scalability of web-scraping capabilities across a growing data environment.
You will establish reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling.
From investigating new data sources to deploying and supporting production workloads, you will have substantial ownership over the full engineering lifecycle.
The role combines deep Python engineering with cloud infrastructure, data pipelines, observability, and practical problem-solving in a remote-first environment.
You will work independently while collaborating with a distributed engineering team through code reviews, documentation, and structured development workflows.

  • Own the development, deployment, and ongoing operation of web-scraping and web-data ingestion pipelines.
  • Design and establish scalable web-scraping frameworks with reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling.
  • Build, maintain, and improve reliable production scrapers for both new and existing data sources.
  • Investigate websites and determine the most appropriate acquisition method, including APIs, direct HTTP requests, HTML parsing, browser automation, or third-party tooling.
  • Evaluate build-versus-buy options for scraping infrastructure and external services, considering capabilities, reliability, cost, operational complexity, and risk.
  • Ensure web-data acquisition activities appropriately account for internal policies, website terms, robots.txt, access restrictions, privacy, and intellectual-property considerations, escalating unclear situations when required.
  • Diagnose and resolve scraping challenges related to website changes, dynamic content, authentication, sessions, rate limits, concurrency, and other operational constraints.
  • Integrate scraping workloads into scalable data-platform and lakehouse architectures.
  • Improve scheduling, monitoring, storage, validation, and operational support for scraping workloads.
  • Use AI-assisted engineering tools where appropriate while maintaining a thorough understanding of, and accountability for, the code being delivered.
  • Support production workloads through monitoring, debugging, maintenance, and continuous improvement.
  • Contribute to a remote engineering environment through code reviews, documentation, ticket-based workflows, and knowledge sharing.

Requirements

~2 min read
  • Demonstrated professional experience building and operating production web-scraping systems at scale.
  • Proven ability to independently take substantial scraping projects from initial investigation through implementation, deployment, and ongoing production support.
  • Strong production-level Python engineering skills, with experience developing maintainable applications rather than standalone scripts.
  • Hands-on experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright, or Selenium.
  • Strong practical understanding of HTTP, HTML, APIs, JavaScript-rendered websites, and browser/network behavior.
  • Experience addressing common scraping challenges including pagination, authentication, sessions, retries, rate limiting, concurrency, and proxies.
  • Strong understanding of data pipelines, data quality, and how collected data should be validated, stored, and consumed by downstream systems.
  • Experience deploying, monitoring, and supporting production workloads in a cloud environment.
  • Strong debugging, analytical, and problem-solving abilities, with the judgment to make effective engineering decisions independently.
  • Comfortable working within a remote engineering team and participating in code reviews, documentation, and ticket-based development workflows.
  • Experience with AWS is desirable.
  • Familiarity with lakehouse or data-lake architectures, particularly Apache Iceberg, is a plus.
  • Experience with PySpark or other distributed data-processing technologies is beneficial.
  • Familiarity with Docker and containerized workloads is advantageous.
  • Experience with Terraform or other infrastructure-as-code tools is a plus.
  • Familiarity with Grafana or comparable observability platforms is desirable.
  • Experience operating high-volume or distributed crawling systems is beneficial.
  • Experience evaluating or operating commercial scraping, proxy, or browser-infrastructure services is a plus.
  • Experience implementing automated scraper testing, canary runs, or source-drift detection is desirable.
  • Exposure to legal, compliance, privacy, or data-governance processes related to web-data acquisition is advantageous.
  • Strong ownership, autonomy, documentation, communication, and collaboration skills.

What We Offer

~2 min read
✓Fully remote position within a remote-first technology team.
✓Opportunity to take ownership of a critical web-data acquisition capability and influence its architecture and operating standards.
✓Senior, hands-on role with substantial autonomy across investigation, engineering, deployment, and production support.
✓Work on scalable data pipelines and modern lakehouse architectures supporting research and data products.
✓Exposure to cloud infrastructure, distributed processing, observability, browser automation, APIs, and production scraping technologies.
✓Opportunity to establish reusable engineering patterns and improve the reliability and scalability of data ingestion.
✓Collaboration with a distributed engineering team through code reviews, documentation, and structured workflows.
✓Environment that supports independent problem-solving, technical ownership, and continuous improvement.
✓Fully remote setup available across the relevant distributed team environment.

Location & Eligibility

Where is the job
India
Remote within one country
Who can apply
IN

Listing Details

Posted
September 30, 2026
First seen
September 30, 2026
Last seen
September 30, 2026

Posting Health

Days active
-1
Repost count
1
Trust Level
62%
Scored at
September 30, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Senior Data Engineer (Web Scraping)