Member of Technical Staff, Data Infrastructure
Bay Areafull-timelead
OtherMember Of Technical Staff
0 views0 saves0 applied
Quick Summary
Key Responsibilities
distributed compute, data orchestration, and storage across modalities. Develop high-throughput systems for data ingestion, processing, and transformation — including training data catalogs,
Technical Tools
OtherMember Of Technical Staff
The Role
We seek experienced engineers to architect and scale the core infrastructure behind distributed training pipelines and petabyte-scale data catalogs. You'll work directly with researchers to accelerate experiments, develop new datasets, improve infrastructure efficiency, and enable key insights across our data assets.
Key Responsibilities
- Design, build, and operate scalable, fault-tolerant infrastructure for LLM research: distributed compute, data orchestration, and storage across modalities.
- Develop high-throughput systems for data ingestion, processing, and transformation — including training data catalogs, deduplication, quality checks, and search.
- Build systems for web crawling, data ingestion, and real-time data processing to support model training operations.
- Develop tools and frameworks for efficient data storage, retrieval, and versioning across distributed systems.
- Ensure data collection adheres to privacy regulations.
Qualifications
- BS/MS/PhD in Computer Science, Machine Learning, or a related field (or equivalent experience).
- 3+ years of experience building data processing pipelines at scale, particularly with AI/ML applications.
- Strong proficiency in Python and experience with data processing frameworks (Apache Spark, Beam, Airflow).
- Familiarity with synthetic data generation techniques and data augmentation strategies.
- Familiarity with web scraping, crawling technologies, and Common Crawl datasets.
- Solid understanding of machine learning fundamentals and experience with ML frameworks (PyTorch, TensorFlow).
- Experience with SQL and NoSQL databases for managing structured and unstructured data.
Preferred Skills
- Experience with large language models and understanding of tokenization, embeddings, and model architectures.
- Experience managing human annotation workflows and quality control processes.
- Experience with vector databases and embedding-based retrieval systems.
- Knowledge of data privacy regulations and ethical AI practices.
- Experience with distributed computing and large-scale data storage systems (HDFS, S3, BigQuery).
What We Offer
~1 min readThe annual base salary range for this role is $200,000 – $350,000 USD. Final compensation is determined based on experience, skills, and qualifications. Equity and benefits are included in the total package.
Location & Eligibility
Where is the job
Bay Area
On-site at the office
Who can apply
Same as job location
Listing Details
- Posted
- March 10, 2026
- First seen
- September 26, 2026
- Last seen
- October 5, 2026
Posting Health
- Days active
- 9
- Repost count
- 0
- Trust Level
- 20%
- Scored at
- October 6, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application
Similar Member Of Technical Staff jobs
View all →Anaplan Solution Architect
ETIC, SAP Security Associate - Cyber Security
Technicien supérieur en électronique - conception et maintenance des testeurs F/H (MRF)
Senior Associate - Economic & Policy (Sustainability and Climate Change)
Gestionnaire de marchés - F/H (DSI/ TSI)
顧問類(台北)-顧問/資深顧問(電腦審計/數據分析/風險管理)
Browse Similar Jobs
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.