Sr Machine Learning Engineer - AI
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Sr Machine Learning Engineer - AI based in India.
This role focuses on developing and deploying efficient small language models (SLMs) for real-world AI applications.
You’ll work across model fine-tuning, optimization, evaluation, and production deployment.
The position combines machine learning engineering with practical MLOps and performance engineering.
You’ll help make models smaller, faster, and more efficient through techniques such as quantization, pruning, and knowledge distillation.
Your work will extend to edge devices, mobile environments, and local infrastructure where latency and resource efficiency are critical.
You’ll also build reliable pipelines and monitoring systems that support models throughout their production lifecycle.
The role offers an opportunity to contribute to advanced AI systems while solving challenging performance and deployment problems.
-
Fine-tune and train small language models using Hugging Face, TRL, and adapter-based techniques such as LoRA, QLoRA, and PEFT.
-
Optimize models for efficient inference through quantization, pruning, knowledge distillation, and other model-compression approaches.
-
Deploy machine learning models to edge devices, mobile platforms, and local servers while meeting demanding latency and resource constraints.
-
Build end-to-end MLOps pipelines covering data ingestion, experimentation, model development, evaluation, deployment, and production operations.
-
Establish and maintain model evaluation frameworks, benchmarking processes, and custom test suites to measure model quality and performance.
-
Monitor production models for accuracy, inference latency, CPU/GPU utilization, and other relevant operational metrics.
-
Contribute to continuous improvements in AI deployment workflows, model efficiency, reliability, and scalability.
Requirements
~1 min read-
Hands-on experience developing, training, and fine-tuning small language models or other transformer-based models using Hugging Face and related tooling.
-
Strong knowledge of adapter-based fine-tuning methods, including LoRA, QLoRA, and PEFT.
-
Practical experience with model optimization techniques such as quantization, pruning, and knowledge distillation.
-
Experience deploying machine learning models to edge devices, mobile environments, or local/on-premises infrastructure.
-
Ability to design and implement end-to-end MLOps pipelines from data ingestion through production deployment.
-
Experience monitoring machine learning systems in production, including model accuracy, latency, and hardware utilization.
-
Strong understanding of model evaluation, benchmarking, and performance optimization.
-
Experience with experiment tracking, model registries, and ML-focused CI/CD practices is valuable.
-
Knowledge of ONNX export and cross-platform inference is an advantage.
-
Strong problem-solving skills and the ability to work effectively in a collaborative engineering environment.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- First seen
- September 28, 2026
- Last seen
- September 28, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 28, 2026
Signal breakdown
Similar Machine Learning Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.