JOBSEARCHER

Staff Machine Learning Engineer - LLM Quantization & Deployment

Overview In this role, you will develop production-grade LLM inference and quantization solutions to power XPENG’s VLA stack. You will collaborate with research, systems, and product teams to deliver reliable, low-latency models and calibration pipelines for on-vehicle deployment. You will help define metrics and benchmarks to evaluate LLM performance and guide optimization decisions. This is a hands-on, cross-functional role at the intersection of AI, software engineering, and autonomous driving, offering impact on transport innovation. Compensation / Benefitsfun, supportive environmentcomputational resourcescutting-edge technologiesopportunity to impact transportation revolutioncompetitive compensationsnacks, lunches, dinners, and activities ResponsibilitiesDevelop VLA inference models and productionize PTQ/QAT and lower-bit techniques (INT8, FP4, mixed-precision)Write production-grade Python with strong testing, observability, and reproducibilityBuild export, calibration, benchmarking, validation, and deployment pipelinesCollaborate with VLA research to estimate performance and prove feasibilityCurate evaluation datasets and establish a comprehensive benchmarking metric suiteAnalyze numerical errors, regressions, and performance trade-offsDevelop PTQ and QAT orchestration workflowsCoordinate with field-testing/simulation teams for autonomy performance sign-offWork with in-vehicle software on latency analysis and issue triagePartner with training infra to advance QAT and model distillation Key requirementsMaster in CS/CE/EE, or equivalent, with 3-5 years of industry experienceStrong understanding of Transformer architectures and LLM inferenceHands-on experience quantizing or deploying DL models in productionProficiency with PyTorch and at least one inference or compilation stackStrong Python programming and software engineering skillsAbility to collaborate across research, systems, infra, and product teamsExcellent communication and problem-solving skills in a fast-paced, collaborative environmenteffective cross-functional collaborationproblem-solving under pressureclear communication of technical conceptsquantization methods PTQ, QAT, AWQ, GPTQ, SmoothQuantweight-only/activation/KV-cache/mixed-precision quantizationLLM runtimes such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, custom runtimes