Research Engineer, Privacy and Anonymization
Quick Summary
Strong proficiency in Python and experience building reliable production data or ML systems Experience with information extraction, named-entity recognition, classification,
HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.
About the Role
~1 min readWe’re looking for a Research Engineer to build the privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training. You’ll develop methods to detect and remove PII, secrets, and other sensitive information from raw data before it enters our processing and synthetic data pipelines. You’ll own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable for training agents.
Responsibilities
~1 min read- →
Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information and design transformations based on the data type and downstream use case
- →
Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods
- →
Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows
- →
Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts
- →
Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields
- →
Work with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards
Strong proficiency in Python and experience building reliable production data or ML systems
Experience with information extraction, named-entity recognition, classification, or related methods for detecting rare or sensitive content
Strong experimental instincts and the ability to compare approaches across recall, precision, latency, cost, and downstream data utility
An understanding of the difference between redaction, masking, pseudonymization, anonymization, and synthetic data—and when each is appropriate
High attention to detail and the ability to reason about subtle leakage paths, edge cases, and adversarial failure modes
Built data processing pipelines end-to-end without a fully prescribed roadmap
Hands-on experience with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, format-preserving encryption, etc.
Worked with sensitive data in areas such as healthcare, finance, or security
Built low-latency or high-throughput ML inference and data-processing systems
Worked in unstructured problem spaces and take ownership from early research through production deployment
Early-stage startup experience and strong communication skills for collaboration across teams and time zones
We prioritize technical aptitude and learning potential over years of experience. Motivated candidates are encouraged to apply even if they don't meet all criteria.
Visa Sponsorship: We provide support for relocation and visas for strong full-time candidates to the US or Singapore.
Timeline: Applications are rolling. The process is 2 technical interviews and a 2-3 day work trial.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- September 16, 2026
- First seen
- September 26, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 38%
- Scored at
- September 26, 2026
Signal breakdown
Similar Research Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.