Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.
Evaluate the quality, correctness, and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks for training and evaluating AI models.
Assess CUDANKI migration fidelity and Trainium-specific performance-optimization quality.
Review cross-platform numerical-correctness standards and provide clear, rubric-based feedback.
Collaborate with AI research teams to ensure consistency and coverage across datasets.
Work independently and asynchronously to improve AI model performance.
Qualifications
Must-Have
2+ years of experience developing or optimizing kernels using NKI targeting AWS Trainium/Inferentia2 hardware.
Strong understanding of NKI-specific development patterns and memory-hierarchy management.
Experience assessing CUDANKI migration quality.
Familiarity with Trainium-specific performance profiling.
Experience defining or evaluating cross-platform numerical-correctness standards.
Preferred
Experience with AWS Neuron SDK, Neuron Compiler internals, or contributions to NKI kernel libraries.
Prior CUDA or Triton kernel development.
Familiarity with Trainium hardware specifications.
Experience benchmarking ML training workloads on Trn1/Trn2 instances.
Application Process (Takes 20–30 mins to complete)
Upload resume
AI interview based on your resume
Submit form
Resources & Support
For details about the interview process and platform information, please check: https://talent.docs.mercor.com/welcome
For any help or support, reach out to: support@mercor.com
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
#hiringmercor
Numbers & Facts
Location
San Francisco, California (Remote)
Salary
$70–$90 Per Hour
Skills
Amazon Web Services (AWS)unmatched
Artificial Intelligence (AI)unmatched
Benchmarkingunmatched
CUDA (Compute Unified Device Architecture)unmatched
Data Setsunmatched
Hardware Specificationunmatched
Kernel Programmingunmatched
Memory Managementunmatched
Multiplatform/Cross-Platformunmatched
Performance Managementunmatched
Performance Modelingunmatched
Performance Tuning/Optimizationunmatched
Research Laboratoryunmatched
Training/Teachingunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.