LLM Inference & GPU Systems Consultant

USUS·Charlottemid
OtherConsultant
0 views0 saves0 applied

Quick Summary

Key Responsibilities

Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.

Technical Tools
OtherConsultant

We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.

Responsibilities

~1 min read

NVIDIA GPU Runtime Optimization: Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.

Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.

Hardware Utilization: Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.

Model Lifecycle Management: Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.

Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.

Requirements

~1 min read

8+ years experience working as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.

8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode).

Proficiency in OpenShift AI and GPU orchestration tools like RunAI.

Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM.

Proven track record managing the Hugging Face deployment lifecycle.


Location & Eligibility

Where is the job
Charlotte, US
On-site at the office

Listing Details

Posted
August 6, 2026
First seen
October 3, 2026
Last seen
October 4, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
13%
Scored at
October 4, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

LLM Inference & GPU Systems Consultant