Software Engineer, Data Mining
Quick Summary
Python, Go, PostgreSQL Infrastructure: Redis, Docker, Kubernetes Frontend: React, TypeScript AI / ML: LLMs, agents Our stack will evolve, you’ll help decide how.
NationGraph is building the data and intelligence layer for the public sector.
More than 110,000 state and local government agencies across the U.S. independently publish information about:
How they operate
What they buy
Who they work with
What problems they are trying to solve
That information is fragmented across millions of websites, documents, databases, procurement systems, meeting records, and public records.
NationGraph turns that information into structured, connected, actionable intelligence for businesses selling to government.
Founded in 2024, NationGraph is dedicated to making uncommon knowledge common, because public data should actually be public.
We’re looking for a Software Engineer, Data Mining to own one of the most important technical problems at NationGraph: building the systems that acquire public-sector information from across the internet at massive scale.
Our goal is to operate hundreds of thousands, and eventually millions, of scrapers covering every level of government across the U.S. and Canada, and eventually worldwide.
This is not a role focused on manually building individual scrapers. You’ll own the infrastructure, abstractions, and automation that allow us to create, deploy, monitor, and maintain an enormous fleet of scrapers reliably.
You’ll work across:
Web crawling and scraping
Browser automation
Distributed systems
Data extraction
Infrastructure and orchestration
LLMs and agents
Monitoring and observability
Responsibilities
~1 min read- →
Build systems for creating, deploying, scheduling, monitoring, and maintaining hundreds of thousands of scrapers.
Design abstractions that allow us to scale toward millions of sources without scaling engineering effort linearly.
Work across government websites, APIs, procurement systems, PDFs, spreadsheets, meeting records, and legacy systems.
Handle changing websites, undocumented APIs, rate limits, broken sources, and countless edge cases.
Build for orchestration, concurrency, retries, backfills, change detection, observability, cost management, and failure recovery.
Ensure we know when sources break, data disappears, or extraction silently becomes incorrect.
Work with our ML Research team to use LLMs and agents to:
Discover new sources
Understand unfamiliar websites
Generate scraping logic
Detect source changes
Diagnose and repair failures
Validate extracted data
Identify common platforms and patterns that can unlock thousands of government agencies at once.
Make new sources increasingly cheap and automated to onboard.
Help comprehensively map public-sector information across the U.S. and Canada.
Build the foundation to eventually acquire public-sector information worldwide.
You’re an unusually strong engineer who enjoys figuring out how things work.
You’ve built production web crawlers, scraping systems, browser automation, or large-scale external data pipelines.
You’re strong in Python, Go, TypeScript, or another backend/systems language.
You understand the realities of scraping modern websites, including:
JavaScript rendering
Sessions and cookies
Rate limits
Proxies
Authentication
Changing schemas and websites
You understand distributed systems, including:
Orchestration
Queues and concurrency
Idempotency
Retries
Backfills
Observability
Failure recovery
You care deeply about data quality, correctness, and reliability.
You’re excited about using LLMs and agents to automate traditionally manual scraping work.
You naturally think about leverage: not how to scrape one website, but how to build a system capable of scraping the next 10,000.
You thrive in ambiguity and would rather build the system than be handed one.
We’re particularly interested in backgrounds spanning:
Alternative data
Quantitative research infrastructure
Search and crawling
AI data infrastructure
Knowledge graphs
Large-scale document processing
Data aggregation
None of these are requirements.
A large part of this architecture still needs to be invented.
You’ll have significant ownership over how NationGraph discovers, acquires, represents, and serves public-sector information.
There is no single API for American government.
There are tens of thousands of institutions, millions of sources, inconsistent schemas, and enormous amounts of information buried in systems never designed for machines.
We believe a major long-term advantage in applied AI will come from proprietary context and data.
Government contains enormous amounts of valuable information that is technically public but practically inaccessible.
Your job is to change that.
You’ll work closely with the CEO, CTO, and a small engineering and research team.
The team has backgrounds spanning high-scale infrastructure, quantitative finance, AI, and startups.
We move quickly.
We operate with very little bureaucracy.
Engineers have significant ownership over technical decisions and product outcomes.
Location & Eligibility
Listing Details
- Posted
- August 27, 2026
- First seen
- September 25, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 28%
- Scored at
- September 26, 2026
Signal breakdown
Similar Software Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.