Data/ML Engineer
Quick Summary
Work closely with the Accelerator leadership team to align data engineering and machine learning initiatives with overarching goals and long-term vision.
Work effectively in a modern, professional software and data engineering environment with a strong understanding of Agile concepts and practices.
The Accelerator seeks a Data/ML Engineer to strengthen our data team and advance the engineering, enrichment, and provisioning of the data we collect.
The Accelerator at Princeton includes a portfolio of multiple planned independent and intersecting tools, built on a shared data and compute platform serving computational social scientists at research institutions across North America, Europe, and Africa. The Data/ML Engineer will work within our team to help drive data engineering and machine learning initiatives and collaborations. They will play a crucial role in building and operating the pipelines that transform large-scale social media and web behavior data into research-ready data products, and in developing the machine learning and enrichment capabilities that extend their value. They will work on problems that have no precedent and little source material, requiring novel solutions. They will also be responsible for working with the other teams within the Accelerator and our external partners to help foster collaboration and create an impactful environment for our users.
Responsibilities
~1 min readWhat We Offer
~1 min read- Design, develop, and operate ML and NLP enrichment pipelines over large-scale text and behavioral data, including language identification, translation, and topic and content classification.
- Own the full lifecycle of enrichment models: selection, evaluation against labeled data, batch inference architecture, cost efficiency, and reprocessing and versioning strategy.
- Develop ML-ready feature layers and data products to support advanced research use cases.
- Evaluate and apply large language model workflows and other emerging AI methods where they demonstrably improve outcomes, with attention to their validity for downstream scientific analysis.
- Apply statistical analysis and modeling to characterize datasets, estimate coverage, and support research design.
- Contribute to cost attribution, visibility, and governance across institutional workspaces, including cluster policies, budget controls, and storage lifecycle management.
- Design data and ML workloads to operate within the platform's cost governance framework.
- Develop automation for workspace and project provisioning as institutions and research projects onboard.
- Operate within Unity Catalog governance, multi-tenant isolation, and research data security requirements.
- Work effectively in a modern, professional software and data engineering environment with a strong understanding of Agile concepts and practices.
- Modern Software Engineering Foundations: agile (Scrum), DevOps, CI/CD, code review, and pair programming, with working knowledge of cloud compute platforms to support collaborative, scalable, and efficient development.
- Author and maintain researcher-facing documentation and provide direct technical support to research users of the platform.
- Collaborate with research teams to define data products, sampling frames, and enrichment requirements, and apply state-of-the-art techniques to ongoing scientific challenges.
- Stay current with the latest advancements in data engineering, machine learning, and relevant fields to continuously innovate.
- Build strong relationships with external partners, driving collaborations that enhance the Accelerator's scientific impact.
Requirements
~2 min read- 3+ years of relevant experience as a data engineer, machine learning engineer, or data scientist, which may include graduate research and internship experience, with a record of building production systems that operate reliably at scale. Experience working in a remote, agile environment.
- Bachelor's degree or equivalent in a relevant field.
- Strong proficiency in Python and SQL, and hands-on experience with distributed data processing (e.g., Apache Spark) on large data volumes.
- Experience building, evaluating, and operating machine learning or NLP pipelines, including batch inference.
- Working knowledge of cloud data platforms.
- Strong communication and interpersonal skills to effectively collaborate with researchers in the field, other engineers at various levels of experience, and administrative and leadership team members.
- Experience with Azure and Databricks, including Unity Catalog.
- Experience with infrastructure-as-code (e.g., Terraform), containers, and CI/CD tooling.
- Experience with large-scale social media, web behavior, or text-as-data research.
- Familiarity with large language model annotation workflows and their evaluation.
- Publications in reputable scientific journals or conferences is desirable.
Princeton University is an Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to age, race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability status, protected veteran status, or any other characteristic protected by law.
The University considers factors such as (but not limited to) scope and responsibilities of the position, candidate's qualifications, work experience, education/training, key skills, market, collective bargaining agreements as applicable, and organizational considerations when extending an offer. The posted salary range represents the University's good faith and reasonable estimate for a full-time position; salaries for part-time positions are pro-rated accordingly.
If the salary range on the posted position shows an hourly rate, this is the baseline; the actual hourly rate may be higher, depending on the position and factors listed above.
The University also offers a comprehensive benefit program to eligible employees. Please see this link for more information.
Location & Eligibility
Listing Details
- Posted
- October 7, 2026
- First seen
- October 7, 2026
- Last seen
- October 7, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 55%
- Scored at
- October 7, 2026
Signal breakdown
Similar Machine Learning Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.