JOBSEARCHER

LLM Inference & GPU Systems Engineer

Hi,Greetings of the day!We are looking to Hire a Talented Professional for the below Job opportunity with one of our clients,If you have any relevant candidate profiles, please share your updated resumes at your earliest convenience, and I'll be happy to provide more details about the role.Job Description :Position: LLM Inference & GPU Systems Engineer Location: charlotte, NC Duration: Long Term Job Description: We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.Key Responsibilities: NVIDIA GPU Runtime Optimization: Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.Hardware Utilization: Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using Run AI and Kubernetes GPU orchestration. Model Lifecycle Management: Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.Required Qualifications 1-3 years' experience as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.Built a multi-agent AI analytics platform using DeepAnalyze-8B, Gemini and vLLM with LLM inference and execution workflows.Experience developing scalable ML pipelines and deployment-ready AI architectures at Dassault Systèmes.Strong foundation in REST APIs, backend development, ML pipelines, and containerized AI application. Regards,surekha.v