GPU Kernel Engineer

🔥 20 hours ago

🗽 New York – Remote

infoinfo

💵 $60 - $80 / hour

⏱ Part Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 24-MAG

24-MAG

2 - 10 employees

🤝 B2B

💼 Consulting

B2B • Consulting

24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.

📋 Description

• Evaluate GPU and accelerator kernels for technical correctness and completeness • Review implementations developed from specifications or reference operators • Assess mathematical behavior, implementation defects, unsupported assumptions, and incomplete solutions • Review kernel development across CUDA, Triton, NKI, and Pallas (JAX) • Evaluate framework-specific implementation choices, execution constraints, translations, and migrations • Compare alternative kernel implementations for correctness and technical quality • Assess outputs against reference implementations using absolute, relative, and ULP-based tolerances • Review floating-point behavior and precision-related edge cases • Evaluate performance using Nsight, Nsight Compute, roofline analysis, and framework-native profilers • Assess benchmarking methodology, latency, throughput, utilization, memory behavior, and claimed improvements • Review compute- and memory-efficiency optimization strategies, including tiling, vectorization, parallelization, and workload decomposition • Review registers, shared memory, caches, memory access, data locality, bandwidth utilization, bank conflicts, coalescing, and register pressure • Diagnose compilation and runtime failures involving drivers, out-of-memory conditions, launch configurations, shape or stride mismatches, and autotuning • Validate task reliability in intended environments and evaluate debugging approaches • Review kernel translation, lowering, hardware migration, compiler transformations, and intermediate-representation decisions • Diagnose kernel implementations using outputs, profiler data, runtime behavior, and source code • Evaluate operator fusion and whether fused kernels preserve intended semantics • Assess assigned tasks against structured technical criteria and provide evidence-based written evaluations • Distinguish valid implementation alternatives from technically flawed approaches

🎯 Requirements

• 3+ years of hands-on experience developing, optimising, or verifying GPU or accelerator kernels • Practical experience with at least two of CUDA, Triton, NKI, or Pallas (JAX) • Strong understanding of numerical correctness, including absolute, relative, and ULP tolerances • Experience selecting and validating appropriate reference implementations • Strong performance profiling and benchmarking experience • Familiarity with Nsight, Nsight Compute, roofline analysis, or comparable profiling tools • Strong understanding of common kernel compilation and runtime failure modes • Experience with at least three of: kernel generation from specification, framework translation or lowering, hardware-target migration, kernel debugging, performance optimisation, operator fusion • Experience across both NVIDIA GPU and custom-accelerator ecosystems preferred • Background in compiler engineering, MLIR, or intermediate-representation lowering advantageous • Strong understanding of memory-hierarchy optimisation preferred • Contributions to kernel or accelerator libraries such as cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls advantageous • Strong written communication and ability to provide precise technical feedback • Must complete work without using confidential or proprietary information belonging to any employer, client, institution, or other third party • H1-B and STEM OPT support unavailable

🏖️ Benefits

• Part-time independent contractor engagement • Fully remote within the United States • Flexible scheduling based on project requirements • Projects may be extended, shortened, or concluded based on project needs and performance • H1-B and STEM OPT support is unavailable for this engagement

Apply Now

Similar Jobs

🕒 August 21

Leidos

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

Part-time transmission engineer designing high-voltage overhead and underground utility systems for Leidos. Applying engineering standards and preparing analyses, specifications, procurement, and construction documents.

🕒 August 11

HeiTech

11 - 50

🤖 Artificial Intelligence

📱 Media

☁️ SaaS

Part-time OCR Engineer developing classification, extraction, validation, and automation for federal document records. Supporting reliable, traceable information processing with minimal manual review.

🕒 July 30

Weekday (YC W21)

11 - 50

💼 Consulting

👥 HR Tech

☁️ SaaS

Chemical Engineering Expert reviewing and validating AI-generated chemical engineering content. Collaborating with experts to enhance AI systems in chemical engineering applications.

🕒 July 27

Welo Global

1001 - 5000

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Prompt Engineer responsible for end-to-end technical migration workflow for transitioning templates to LLM autoraters. Leveraging prompt engineering techniques for maximum model performance using internal tools.

🕒 July 27

Welo Global

1001 - 5000

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Prompt Engineer managing end-to-end technical migration workflow for transitioning templates to LLM autoraters. Leveraging prompt engineering techniques with client’s internal tools to maximize model performance.