Principal Data Engineer, LLM/AI Platforms
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal Data Engineer, LLM/AI Platforms based in United States.
This is a principal-level engineering role focused on building the data infrastructure that powers next-generation LLM and AI platforms at massive scale.
You will architect and optimize data platforms supporting LLMs, Retrieval-Augmented Generation (RAG), and sophisticated agentic systems.
The role combines deep hands-on engineering with technical leadership, platform architecture, and continuous innovation.
You will work with extremely large datasets and distributed systems while emphasizing scalability, resilience, performance, and cost efficiency.
A major focus is turning advanced AI and data science concepts into reliable, production-grade services.
You will collaborate with data scientists, product managers, and engineering teams while mentoring engineers and raising technical standards.
This is an opportunity to shape AI platform engineering practices across critical, large-scale systems in a highly autonomous environment.
- Architect, implement, and optimize data platforms and pipelines designed for LLMs, RAG, and advanced AI agentic systems at Exabyte scale.
- Drive the adoption and deployment of agentic workflows and agent-harnessing techniques to support autonomous, data-driven capabilities.
- Design highly scalable, fault-tolerant, secure, and cost-effective data solutions that enable rapid iteration without compromising engineering quality.
- Develop production-ready code with strong attention to performance, maintainability, testing, and operational reliability.
- Provide technical leadership in data modeling, normalization, semantic cataloging, and data architecture for AI and machine learning workloads.
- Establish MLOps and DataOps best practices for LLM platforms, including monitoring, observability, automated recovery, and service reliability.
- Own the end-to-end lifecycle of critical data services, including development, testing, deployment, monitoring, and continuous optimization.
- Collaborate with data scientists, product managers, and engineering teams to transform research prototypes into robust, production-ready services.
- Lead technical workshops, design reviews, and knowledge-sharing initiatives while mentoring engineers and strengthening organizational expertise in AI platform technologies.
- Champion DevSecOps practices and engineering standards across large-scale distributed data environments.
- Identify opportunities to improve platform performance, reliability, scalability, and developer productivity through new technologies and engineering practices.
Requirements
~2 min read- Master’s degree or PhD in Computer Science, Data Engineering, or a related STEM discipline, or equivalent practical experience.
- 10+ years of progressive experience in Data Engineering or Platform Engineering, including at least 3 years architecting and building AI/ML or Data Science platforms at massive scale.
- 3+ years of experience in a Principal or Staff-level engineering capacity, with demonstrated technical leadership and mentorship experience.
- Hands-on expertise with LLM engineering, including fine-tuning, prompt engineering, deployment, RAG, and agentic workflow development.
- Proven experience designing and delivering large-scale distributed systems, including sharding, partitioning, concurrency, and fault-tolerant architectures.
- Expert-level proficiency in Python or JVM-based technologies, with a strong ability to write clean, performant, maintainable, and well-tested production code.
- Deep experience with distributed data processing frameworks such as Spark, Dask, or Flink.
- Strong knowledge of cloud platforms such as AWS, GCP, or OCI and their associated data services.
- Expertise with containerization and orchestration technologies including Docker and Kubernetes.
- Experience with messaging and streaming technologies such as Kafka or Pulsar.
- Familiarity with data warehousing and orchestration platforms such as Snowflake, BigQuery, Airflow, and Kubeflow.
- Experience with MLOps technologies such as MLflow, SageMaker, or Vertex AI.
- Familiarity with agentic AI frameworks such as LangChain or LlamaIndex.
- Strong understanding of engineering practices including peer code reviews, resilient architecture, comprehensive testing, and secure development methodologies.
- Demonstrated ability to use AI technologies to improve decision-making, automate workflows, increase efficiency, and support measurable business outcomes.
- Strong communication and collaboration skills, with the ability to influence technical direction and mentor engineers across teams.
- Direct experience deploying and managing LLMs in production is a plus.
- Experience in cybersecurity, intelligence, or highly regulated industries is a plus.
- Contributions to open-source data or AI/ML projects are a plus.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 26, 2026
- First seen
- September 27, 2026
- Last seen
- September 29, 2026
Posting Health
- Days active
- 1
- Repost count
- 0
- Trust Level
- 80%
- Scored at
- September 29, 2026
Signal breakdown
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.