Research Engineer
Quick Summary
Salary range - $275k - $325k | Equity - 0.25% | In-person NYC About Datalab Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs,
Salary range - $275k - $325k | Equity - 0.25% | In-person NYC
Datalab trains models that read documents reliably at scale. The world's most important information is trapped in PDFs, scans, and files that can't easily be parsed, and getting it out correctly matters. From frontier AI labs processing training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction has to be right.
We’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our tools, Chandra, Surya, Marker, and Lift, have 70,000+ GitHub stars and broad developer mindshare. We're backed by founding members of OpenAI, FAIR, and Hugging Face.
We're looking for a Research Engineer to own problems end to end across our models, inference service, and product. You won't just train a model and hand it off. You'll take it from training through benchmarking, into our inference stack, and work with the team to integrate it into our products.
We're a small team that has shipped the current state of the art OCR model, Chandra. Our models collectively have 70k+ Github stars. Our tools are used internally at frontier AI labs like Anthropic, and Fortune 500 enterprises like Siemens.
Our team focuses on training small, efficient models that outperform much larger LLMs on domain-specific tasks (like OCR, structured extraction, tables). We move fast, prioritize practical results, and build tools that are open, reproducible, and built to last. You'll test hypotheses quickly, iterate on results, and balance experimental rigor with shipping to customers.
A typical project might look like: identify a gap in extraction quality on long documents, train and benchmark a new model, optimize it for inference, and work with the team to ship it to users. Concretely:
Nice to Have
~1 min readHave experience with OCR, document AI, or structured extraction
Have published work, whether that's a paper, a benchmark report, or a deep technical blog post
Have been a major contributor to open-source projects, especially in ML, vision, or NLP
Enjoy writing about your work and sharing learnings with the community
A 30-minute video call to evaluate fit
90-minute live architecture discussion
Culture fit interview/team meeting
At this stage of the company, every interview is somewhat custom, so these phases may be rearranged slightly.
Datalab is an equal opportunity employer. We do not discriminate on the basis of any characteristic protected by federal, state, or local law.
If you need an accommodation to participate in our hiring process, or if you believe you have experienced discrimination or harassment at any point in it, contact operations@datalab.to.
Location & Eligibility
Listing Details
- Posted
- July 6, 2026
- First seen
- September 29, 2026
- Last seen
- September 29, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 30%
- Scored at
- September 29, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.