AIPayList

Jobs / Mercor

Trainium (NKI) Kernel Expert

up to $90/hr

source wording: “70–90 USD HOUR

Platform
Mercor
Category
Other AI work
Eligibility
Worldwide
Freshness
posted 17h ago · seen live 1h ago

What this role asks for

  • 2+ yrs experience
Read from the posting's own words: show the exact sentences
  • 2+ yrs experience: “…clear, rubric-based written feedback. Basic Qualifications • 2+ years of hands-on experience developing or optimizing kernels using the Neuron Kernel Interface (NKI) targeting AWS…

The role, as Mercor describes it

Evaluate the quality, correctness, and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks used to train and evaluate a frontier AI lab's models. You'll assess CUDA→NKI migration fidelity, Trainium-specific performance-optimization quality, and cross-platform numerical-correctness standards — and provide clear, rubric-based written feedback.

Basic Qualifications

• 2+ years of hands-on experience developing or optimizing kernels using the Neuron Kernel Interface (NKI) targeting AWS Trainium/Inferentia2 hardware

• Strong understanding of NKI-specific development patterns: tile-based computation, SBUF/PSUM/HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration

• Demonstrated experience assessing CUDA→NKI migration quality

Read the rest of the description (7 more paragraphs)

• Familiarity with Trainium-specific performance profiling (NeuronCore pipeline utilization, tensor-engine throughput, memory-bandwidth bottlenecks)

• Experience defining or evaluating cross-platform numerical-correctness standards (GPU vs Trainium accumulation order, rounding behavior, mixed-precision semantics)

Preferred Qualifications

• Direct experience with AWS Neuron SDK, Neuron Compiler internals, or contributions to NKI kernel libraries

• Prior CUDA or Triton kernel development

• Familiarity with Trainium hardware specifications (NeuronCore-v2 architecture, on-chip SRAM topology, supported data types: FP32/BF16/FP8/INT8)

• Experience benchmarking ML training workloads on Trn1/Trn2 instances

Posted by Mercor, reproduced here so you can judge the role before clicking. Original posting ↗

Source: platform job feed · first seen 1h ago · ID list_AAABoDrL07CqmjVBGadLDIZ6