LLMOps & Inference Optimization Engineer
LLMOps & Inference Optimization EngineerLocation: LATAM (Remote)🚀 About VariacodeAt Variacode, we connect top-tier tech talent with cutting-edge global projects. We are currently looking for an ambitious, challenge-seeking LLMOps & Inference Optimization Engineer to join a high-performing team building next-generation, high-throughput AI applications and proprietary large language model architectures.Our client's technology is used in 80+ countries by Fortune 500 companies (like Home Depot), allowing millions of users to preview products in their own spaces in real-time. If you want to make a massive impact and work on high-traffic, scalable systems, this is your place!🎯 The RoleWe are seeking a visionary technical authority to lead our foundational AI infrastructure strategy. As a LLMOps & Inference Optimization Engineer, you will own the architectural design, deployment, and optimization of our self-hosted model ecosystem, directly resolving massive GPU resource constraints and setting the global standard for low-latency, enterprise-grade AI production systems.⚙️ What you will do (Responsibilities):Architect the end-to-end deployment and scaling strategy for proprietary and open-source foundation models across highly distributed production environments.Define global optimization standards for model inference speed and GPU utilization using advanced quantization (AWQ, GPTQ) and memory management techniques.Lead the design of high-throughput serving architectures utilizing frameworks like vLLM, TensorRT-LLM, and Triton Inference Server.Establish the blueprints for advanced, highly reliable Retrieval-Augmented Generation (RAG) pipelines incorporating hybrid search and enterprise vector databases.Govern cloud-native GPU orchestration, budget allocation, and distributed inference topologies leveraging Ray.Mentor senior engineers and collaborate closely with executive leadership, Product, and Data Science teams to align the AI engineering roadmap with long-term business goals.🔍 What we are looking for: Must-haves (Required):Experience: 8+ years of software, data, or ML engineering experience, with a deep, proven track record of architecting and deploying self-hosted AI models in high-scale production.Technical Stack & Tools: vLLM, TensorRT-LLM, Triton Inference Server, Ollama, PyTorch, Quantization (AWQ/GPTQ), FlashAttention, and distributed systems via Ray.Key Skills: High-level systems architecture design, strategic technical leadership, advanced mathematical/GPU resource optimization, and deep technical ownership.Language: Fluent English proficiency with exceptional technical communication and consultative skills.Nice-to-haves (Pluses):Active contributions to open-source LLMOps or deep learning frameworks.Hands-on CUDA programming and custom GPU kernel design.Expertise in fine-tuning methodologies (LoRA/QLoRA) and complex Vector Databases (Pinecone, Qdrant, Milvus).Experience driving DevSecOps compliance and model lineage governance within highly regulated industries.🎁 What we offer:Opportunity to work with cutting-edge tech stacks and Fortune 500 clients.A highly collaborative, innovation-driven, and remote-first global culture.Continuous professional growth and career development pathways.Competitive remote-first compensation package in USD, tailored to world-class principal-level AI talent.