{"schemaVersion":"jobsearcher.job.v1","id":"f3fd3abb25bc1ffac9a14119","url":"https://jobsearcher.com/jobs/f3fd3abb25bc1ffac9a14119","canonicalUrl":"https://jobsearcher.com/jobs/f3fd3abb25bc1ffac9a14119","title":"NKI Kernel Developer for AI Model Evaluation","description":"Role Overview Evaluate Neuron Kernel Interface, NKI, development tasks used to train and assess advanced AI models. You will review CUDA-to-NKI migrations, Trainium-focused performance optimizations, and cross-platform numerical correctness, then deliver clear written feedback using defined rubrics. Key Responsibilities Assess NKI kernel-development tasks for quality, correctness, and suitability for AWS Trainium and Inferentia2 hardware. Evaluate the fidelity of CUDA-to-NKI migrations. Review Trainium-specific performance optimization quality. Evaluate numerical-correctness standards across GPU and Trainium platforms. Provide clear, rubric-based written feedback. Qualifications At least 2 years of hands-on experience developing or optimizing NKI kernels for AWS Trainium or Inferentia2 hardware. Strong knowledge of tile-based computation, SBUF, PSUM, and HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration. Experience evaluating CUDA-to-NKI migration quality. Familiarity with Trainium performance profiling, including NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth bottlenecks. Experience establishing or assessing cross-platform numerical-correctness standards, including GPU versus Trainium accumulation order, rounding behavior, and mixed-precision semantics. Preferred Qualifications Experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel-library contributions. Prior CUDA or Triton kernel-development experience. Familiarity with NeuronCore-v2 architecture, on-chip SRAM topology, and FP32, BF16, FP8, and INT8 data types. Experience benchmarking machine-learning training workloads on Trn1 or Trn2 instances. Work Terms Remote role, open to candidates located in the United States. Hourly engagement. Compensation $70 to $90 per hour.","company":"Confidential","rawCompany":"confidential","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-31T11:28:06.298Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"NKI Kernel Developer for AI Model Evaluation","description":"Role Overview Evaluate Neuron Kernel Interface, NKI, development tasks used to train and assess advanced AI models. You will review CUDA-to-NKI migrations, Trainium-focused performance optimizations, and cross-platform numerical correctness, then deliver clear written feedback using defined rubrics. Key Responsibilities Assess NKI kernel-development tasks for quality, correctness, and suitability for AWS Trainium and Inferentia2 hardware. Evaluate the fidelity of CUDA-to-NKI migrations. Review Trainium-specific performance optimization quality. Evaluate numerical-correctness standards across GPU and Trainium platforms. Provide clear, rubric-based written feedback. Qualifications At least 2 years of hands-on experience developing or optimizing NKI kernels for AWS Trainium or Inferentia2 hardware. Strong knowledge of tile-based computation, SBUF, PSUM, and HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration. Experience evaluating CUDA-to-NKI migration quality. Familiarity with Trainium performance profiling, including NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth bottlenecks. Experience establishing or assessing cross-platform numerical-correctness standards, including GPU versus Trainium accumulation order, rounding behavior, and mixed-precision semantics. Preferred Qualifications Experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel-library contributions. Prior CUDA or Triton kernel-development experience. Familiarity with NeuronCore-v2 architecture, on-chip SRAM topology, and FP32, BF16, FP8, and INT8 data types. Experience benchmarking machine-learning training workloads on Trn1 or Trn2 instances. Work Terms Remote role, open to candidates located in the United States. Hourly engagement. Compensation $70 to $90 per hour.","datePosted":"2026-08-31T11:28:06.298Z","dateModified":"2026-08-31T11:28:06.298Z","hiringOrganization":{"@type":"Organization","name":"Confidential","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f3fd3abb25bc1ffac9a14119"},"url":"https://jobsearcher.com/jobs/f3fd3abb25bc1ffac9a14119"}}