
51 - 200 employees
Founded 2023
đĽ Funding within the last year
đ° $350M Series C - Mercor on 2025-10
Mercor is a company for which no descriptive text was provided in the input. Additional information (products, services, target customers, or industry specifics) is needed to create an accurate summary and select appropriate industries.
đĽ 0 minutes ago
đşđ¸ United States â Remote
đľ $500 / year
âł Contract/Temporary
đ˘ Junior
đˇđťââď¸ Engineer
đŤđ¨âđ No degree required
đť Ghost score 0%
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
Founded 2023
đĽ Funding within the last year
đ° $350M Series C - Mercor on 2025-10
Mercor is a company for which no descriptive text was provided in the input. Additional information (products, services, target customers, or industry specifics) is needed to create an accurate summary and select appropriate industries.
⢠Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization ⢠Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements ⢠Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms ⢠Write, modify, and reason about C++17, Python, and GPU programming code ⢠Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes ⢠Document optimization decisions clearly, including when specific profiler metrics are or are not useful ⢠Submit a resume or relevant technical background ⢠Complete a brief technical assessment or submit additional information if requested
⢠Available to work at least 20 hrs/wk ⢠Fluent in core C++ features through C++17 ⢠Working knowledge of Python and Git ⢠Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming ⢠At least 1 year of professional or graduate-level research experience working with GPUs ⢠Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels ⢠Ability to optimize GPU kernels without needing deep prior context on every algorithm ⢠Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus ⢠Experience optimizing kernels for NVIDIA Blackwell hardware is a plus ⢠Familiarity with NSight Compute is a plus ⢠Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus ⢠Open-source contributions related to GPU kernel optimization are a plus
⢠Earn up to $480 for each successful referral ⢠Flexible freelance contract opportunity ⢠Remote work
Apply Nowđ June 24
CUDA engineering expert optimizing GPU kernels for a leading AI lab. Requires strong C++ skills and GPU programming experience in a remote contract role.