Quick Summary
We are looking for an AI Specialist Engineer to enhance the performance of large language and vision models for on-device inference. Your expertise will be crucial in developing and deploying cutting-edge AI solutions, ensuring optimal efficiency across diverse hardware architectures.
Compress and optimize large language and vision models for on-device inference. Develop pipelines for model distillation and hardware-specific compilation. Benchmark performance across various NPU/GPU architectures.
Expertise in model distillation, pruning, and 4-bit/8-bit quantization techniques. Hands-on experience with TensorRT, ONNX Runtime, and edge deployment. Strong C++ and Python skills.
We are looking for an AI Specialist Engineer to enhance the performance of large language and vision models for on-device inference. Your expertise will be crucial in developing and deploying cutting-edge AI solutions, ensuring optimal efficiency across diverse hardware architectures.
Responsibilities
~1 min read- →Compress and optimize large language and vision models for on-device inference.
- →Develop pipelines for model distillation and hardware-specific compilation.
- →Benchmark performance across various NPU/GPU architectures.
Requirements
~1 min read- Expertise in model distillation, pruning, and 4-bit/8-bit quantization techniques.
- Hands-on experience with TensorRT, ONNX Runtime, and edge deployment.
- Strong C++ and Python skills.
Location & Eligibility
Listing Details
- Posted
- April 24, 2026
- First seen
- April 24, 2026
- Last seen
- August 29, 2026
Posting Health
- Days active
- 164
- Repost count
- 0
- Trust Level
- 21%
- Scored at
- October 6, 2026
Signal breakdown
Similar Ai Specialist jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.