AIPayList

Jobs / Mercor

GPU Kernel Expert

up to $90/hr

source wording: “70–90 USD HOUR

Platform
Mercor
Category
Other AI work
Eligibility
Worldwide
Freshness
posted 17h ago · seen live 1h ago

What this role asks for

  • 3+ yrs experience
  • Translation
Read from the posting's own words: show the exact sentences
  • 3+ yrs experience: “…clear, rubric-based written feedback. Basic Qualifications • 3+ years of hands-on experience developing, optimizing, or verifying GPU/accelerator kernels in at least two of: CUDA,…
  • Translation: “…at least three kernel task types: generation from specification, translation/lowering across frameworks, migration between hardware targets, debugging, performance…

The role, as Mercor describes it

Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks used to train and evaluate a frontier AI lab's models. You'll assess numerical correctness, performance-benchmarking fairness, task scoping, and compilation/runtime validity across diverse kernel task types — and provide clear, rubric-based written feedback.

Basic Qualifications

• 3+ years of hands-on experience developing, optimizing, or verifying GPU/accelerator kernels in at least two of: CUDA, Triton, NKI, or Pallas (JAX)

• Strong understanding of numerical-correctness criteria for kernels (absolute/relative/ULP tolerances, reference-implementation selection)

Read the rest of the description (8 more paragraphs)

• Demonstrated experience with performance profiling and benchmarking (nsight, ncu, roofline analysis, or framework-native profilers)

• Familiarity with common compilation and runtime failure modes (driver mismatches, OOM, launch-configuration errors, shape/stride mismatches, autotuning failures)

• Experience with at least three kernel task types: generation from specification, translation/lowering across frameworks, migration between hardware targets, debugging, performance optimization, or operator fusion

Preferred Qualifications

• Experience across both NVIDIA GPU (CUDA/Triton) and custom-accelerator (NKI/Pallas/TPU) ecosystems

• Background in compiler engineering, MLIR, or intermediate-representation lowering

• Understanding of memory-hierarchy optimization (shared-memory tiling, register pressure, bank conflicts, coalescing patterns)

• Contributions to kernel libraries (cuBLAS, cuDNN, Triton community kernels, JAX/XLA custom calls)

Posted by Mercor, reproduced here so you can judge the role before clicking. Original posting ↗

Source: platform job feed · first seen 1h ago · ID list_AAABoDrLqkDg31vugrpECa5L