JOBSEARCHER

AI Engineer

Role: AI EngineerLocation: Richardson, TX/ Raleigh, NC/ Phoenix AZ Job Description: Benchmark Al models (LLMs, vision, multimodal) across hardware configurations; measure latency, throughput, utilization, memory behavior, and scaling efficiencyProfile workloads end-to-end using tools such as Nsight Systems/Compute, PyTorch Profiler, and system telemetry (nvidia-smi, DCGM) to isolate bottlenecksBuild roofline/performance models to quantify achieved vs. theoretical performance and prioritize the highest-impact optimizationsApply and evaluate optimizations: quantization, pruning, distillation, operator/kernel fusion, graph compilation, KV-cache management, batching strategies, speculative decodingRecommend hardware/system configurations (GPU selection, memory sizing, interconnect, storage/network I/O) for given model workloadsEstablish performance baselines, SLAs, and regression testing so models stay fast as they evolveWrite clear analyses and recommendations for engineering and leadership audiences.Required qualificationsBS/MS in CS, Computer Engineering, EE, or equivalent practical experienceStrong Python; working proficiency in at least one systems language (C++/Rust/C)Hands-on experience with a deep-learning framework (PyTorch preferred), including model execution, export, and profilingDemonstrated experience delivering measurable performance improvements in DL training or inferenceSolid grounding in computer architecture: memory hierarchy, bandwidth vs. compute limits, parallelismAbility to reason quantitatively about latency, throughput, batching, memory footprint, and utilization under real workloadsFluency with Linux and GPU computing environments Preferred qualificationsGPU programming (CUDA, Triton, ROCm/HIP) and low-level libraries (cuBLAS, cuNN, CUTLASS)Inference runtimes/serving engines: TensorRT(-LLM), ONNX Runtime, vLLM, SLang, Triton Inference ServerLLM inference mechanics: attention, KV caching, prefill vs. decode, continuous batching, speculative decodingDistributed training/inference: data/tensor/pipeline parallelism, NCCL, InfiniBand/RoCEModel compression research or MLPerf-style benchmarking experienceEdge/on-device deployment (Jetson, NPUs, Core ML) if your systems include edge hardware