AI Engineer – LLM Data
Quick Summary
About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As a Data Engineer specializing in Natural Language Processing (NLP) and large-scale data processing, you will quickly and effectively gather, curate, and prepare high-quality datasets to support cutting-edge NLP research.
About the Institute of Foundation Models
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
The Role
As an AI Engineer specializing in LLM data, you will build and improve high-quality training data for foundation models across pre-training, mid-training, and post-training. Your work will include large-scale data curation and processing, LLM-based data synthesis, data quality evaluation, and experimentation to understand how different data choices affect model performance.
Build, curate, and improve large-scale datasets for LLM pre-training, mid-training, and post-training.
Rapidly support time-sensitive data and model development tasks in a fast-moving research environment.
Develop and improve data processing pipelines including data extraction, cleaning, filtering, deduplication, quality scoring, transformation, and dataset composition.
Design and implement LLM-based data synthesis and augmentation pipelines, including prompt-based generation, filtering, refinement, and quality control of synthetic data.
Research and apply methods for improving training data quality, diversity, coverage, and efficiency.
Design experiments to understand the relationship between training data and model performance, and use model evaluation results to guide data improvements.
Develop scalable tools and workflows for processing and analyzing large datasets efficiently.
Analyze datasets using both statistical and model-based methods to identify quality issues, biases, duplication, distributional gaps, and opportunities for improvement.
Collaborate closely with researchers, model engineers, and other data teams to translate model development needs into effective data solutions.
Document datasets, experiments, data processing methodologies, and key findings clearly to support reproducibility and knowledge sharing.
Experience preparing data for large-scale LLM pre-training, continued/mid-training, supervised fine-tuning, preference optimization, reinforcement learning, or other post-training workflows.
Experience with synthetic data generation using LLMs, including generation, filtering, verification, or quality evaluation.
Experience designing or running LLM evaluations, benchmarks, model training, or fine-tuning experiments.
Understanding of how data quality, mixture, diversity, and scaling affect foundation model performance.
Experience with large-scale or distributed data processing and compute infrastructure.
Experience working with research teams on rapidly evolving foundation model or generative AI projects.
Contributions to open-source AI/ML projects, relevant publications, or demonstrated hands-on work with modern foundation models are a plus.
Location & Eligibility
Listing Details
- Posted
- June 17, 2025
- First seen
- March 26, 2026
- Last seen
- September 9, 2026
Posting Health
- Days active
- 167
- Repost count
- 0
- Trust Level
- 31%
- Scored at
- September 9, 2026
Signal breakdown
Please let Ifm Us know you found this job on Jobera.
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.