NKI Kernel Developer for AI Model Evaluation
Role Overview Evaluate Neuron Kernel Interface, NKI, development tasks used to train and assess advanced AI models. You will review CUDA-to-NKI migrations, Trainium-focused performance optimizations, and cross-platform numerical correctness, then deliver clear written feedback using defined rubrics. Key Responsibilities Assess NKI kernel-development tasks for quality, correctness, and suitability for AWS Trainium and Inferentia2 hardware. Evaluate the fidelity of CUDA-to-NKI migrations. Review Trainium-specific performance optimization quality. Evaluate numerical-correctness standards across GPU and Trainium platforms. Provide clear, rubric-based written feedback. Qualifications At least 2 years of hands-on experience developing or optimizing NKI kernels for AWS Trainium or Inferentia2 hardware. Strong knowledge of tile-based computation, SBUF, PSUM, and HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration. Experience evaluating CUDA-to-NKI migration quality. Familiarity with Trainium performance profiling, including NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth bottlenecks. Experience establishing or assessing cross-platform numerical-correctness standards, including GPU versus Trainium accumulation order, rounding behavior, and mixed-precision semantics. Preferred Qualifications Experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel-library contributions. Prior CUDA or Triton kernel-development experience. Familiarity with NeuronCore-v2 architecture, on-chip SRAM topology, and FP32, BF16, FP8, and INT8 data types. Experience benchmarking machine-learning training workloads on Trn1 or Trn2 instances. Work Terms Remote role, open to candidates located in the United States. Hourly engagement. Compensation $70 to $90 per hour.