Quick Summary
3–7 years of relevant professional experience in software engineering, artificial intelligence, machine learning, or related technical fields,
This is an opportunity to build next-generation generative AI applications for enterprise search, summarization, content generation, and other language-driven use cases.
You will work at the intersection of natural language processing, large language models, and modern software engineering.
The role involves designing LLM application architectures, optimizing prompts, integrating embedding models, and developing production-ready AI solutions.
You will work with both commercial and open-source models while exploring fine-tuning and domain-specific adaptation techniques.
A strong focus on scalability, performance, cost efficiency, and response latency will be central to the role.
You will also collaborate across backend and AI engineering environments to integrate LLM capabilities into reliable software systems.
This position is well suited to an engineer who combines strong Python development skills with deep curiosity and practical expertise in generative AI.
- Design and develop high-performance generative AI applications using commercial and open-source large language models, including models such as GPT-4, Claude, Llama 3, and Mistral.
- Develop, test, and systematically refine prompt templates, system instructions, and few-shot learning examples to improve model accuracy, consistency, relevance, and compliance.
- Prepare custom training datasets and implement parameter-efficient fine-tuning approaches such as PEFT and LoRA to adapt language models for domain-specific tasks and use cases.
- Integrate LLM services into scalable backend and microservices architectures using efficient API frameworks, asynchronous processing, message queues, and supporting infrastructure.
- Implement strategies such as semantic caching, token optimization, and model routing to improve production performance while managing inference costs and response latency.
- Work with embedding models, vector indexes, and retrieval-oriented architectures to support robust enterprise search and other knowledge-intensive AI applications.
- Establish systematic approaches for evaluating LLM outputs through automated evaluation suites, structured testing, and human-in-the-loop validation frameworks.
Requirements
~1 min read- 3–7 years of relevant professional experience in software engineering, artificial intelligence, machine learning, or related technical fields, with strong hands-on exposure to generative AI and LLM applications.
- Strong Python programming skills and extensive experience with LLM development frameworks such as LangChain, LlamaIndex, and/or Semantic Kernel.
- Deep understanding of transformer architectures, tokenization methods, quantization techniques, and the execution and integration of open-source language models.
- Practical experience building scalable REST APIs using technologies such as FastAPI or Flask and integrating AI services with message brokers, vector indexes, or other backend infrastructure.
- Experience with LLM application architecture, prompt engineering, model adaptation, and production deployment of AI-powered applications.
- Strong analytical and testing skills, with the ability to systematically evaluate model outputs using automated metrics, evaluation suites, and human validation approaches.
- Bachelor’s degree or equivalent technical education, combined with strong problem-solving skills and the ability to work independently in a fast-moving AI engineering environment.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 21, 2026
- First seen
- September 27, 2026
- Last seen
- September 27, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 46%
- Scored at
- September 27, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.