NKI Kernel Optimization Specialist

Mercor
San Francisco, CA
Job Description
Role Overview

Evaluate the quality, correctness, and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks for training and evaluating AI models. Collaborate with AI research teams to ensure consistency and coverage across datasets. Work independently and asynchronously to improve AI model performance.

What You Will Do

Evaluate NKI development tasks, assess CUDA→NKI migration fidelity, review numerical-correctness standards, and provide feedback. Collaborate with AI research teams to ensure consistency and coverage across datasets.

Why It Might Be a Fit

Must have 2+ years of experience developing or optimizing kernels using NKI targeting AWS Trainium/Inferentia2 hardware. Strong understanding of NKI-specific development patterns and memory-hierarchy management.

Requirements

  • 2+ years of experience developing or optimizing kernels using NKI targeting AWS Trainium/Inferentia2 hardware
  • Strong understanding of NKI-specific development patterns and memory-hierarchy management
  • Experience assessing CUDA→NKI migration quality
  • Familiarity with Trainium-specific performance profiling
  • Experience defining or evaluating cross-platform numerical-correctness standards
]]>