GPU Kernel Engineer – CUDA, Triton, Accelerator Performance

🔥 18 hours ago

🇦🇷 Argentina – Remote

💵 $65 / hour

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Anyone AI

Anyone AI

11 - 50 employees

Founded 2022

📚 Education

🤖 Artificial Intelligence

🎯 Recruiter

💰 $1.5M Pre Seed Round - Anyone AI on 2022-06

Education • Artificial Intelligence • Recruitment

Anyone AI is a Latin America–focused career accelerator and training academy that develops software engineers into AI and machine learning professionals through intensive, remote, mentor-led programs. It combines technical coursework (e. g. , Machine Learning Developer track), professional preparation (resume/LinkedIn optimization, interview coaching, English practice), and a placement-oriented model (income-share payments and access to employer partnerships) to help graduates secure higher-paying AI roles globally. Anyone AI also operates a talent marketplace connecting its community to companies seeking AI talent.

📋 Description

• Review GPU and accelerator kernel implementations for correctness • Compare outputs against reference implementations • Evaluate numerical tolerance thresholds • Review kernel benchmarks and determine whether comparisons are fair • Identify performance bottlenecks and optimization opportunities • Assess whether performance targets are realistic given hardware limits • Review kernel translations and hardware migrations • Identify compilation, driver, memory, shape, and runtime issues • Determine whether technical tasks are genuinely difficult or incorrectly configured • Provide clear, actionable technical feedback • Implement and debug kernels • Optimize CUDA and Triton kernels • Translate between kernel frameworks • Perform hardware migration and operator fusion • Profile and benchmark performance • Verify numerical correctness • Debug compilation and runtime issues • Optimize memory hierarchy and kernel-level AI workload performance

🎯 Requirements

• 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels • Strong experience with at least two of: CUDA; Triton; NKI / AWS Neuron; Pallas / JAX • Strong understanding of GPU performance optimization • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers • Understanding of memory bandwidth, compute throughput, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts • Strong understanding of floating-point numerical correctness and tolerance thresholds • Experience debugging kernel compilation and runtime issues • Ability to distinguish software defects, environment problems, and genuine optimization challenges • Experience writing kernels from technical specifications, translating kernels between frameworks, migrating kernels across hardware platforms, debugging incorrect implementations, optimizing kernel performance, and fusing multiple operations into optimized kernels • Experience across both NVIDIA GPU and custom accelerator ecosystems (nice to have) • Experience with AWS Trainium, TPU, JAX, or other accelerators (nice to have) • Compiler engineering experience (nice to have) • Familiarity with MLIR, XLA, or intermediate representation lowering (nice to have) • Contributions to GPU or ML kernel libraries (nice to have) • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls (nice to have) • Experience with AI model evaluation, RLHF, or technical benchmark development (nice to have)

🏖️ Benefits

• $65 per hour compensation • Part-time, project-based consulting engagement • Remote work

Apply Now

Similar Jobs

🕒 September 4

Interview Pen

1 - 10

📚 Education

Interview Engineer facilitating coding interviews for Karat's developer hiring platform. Evaluating candidates and helping improve inclusive, objective hiring processes.

🕒 August 27

Coderio

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Senior AI Engineer building shared GenAI foundations, LLM gateways, and agentic workflows. Coderio delivers scalable digital solutions for global companies.

AWS

Cloud

Docker

Google Cloud Platform

Kubernetes

Microservices

Python

Terraform

🕒 August 18

Coderio

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Senior AI Engineer building GenAI tooling, copilots, and automation for Coderio’s cloud platforms. Researching and deploying AI solutions across developer platforms, CI/CD pipelines, and DevOps workflows.

Ansible

AWS

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

SDLC

Terraform

Go

🕒 August 13

BPCS, Comprehensive marketing solutions, ltd.

1 - 10

💼 Consulting

📣 Marketing

Substrate PAVC Engineer supporting Microsoft through Blueprint Technologies. Automating Windows patching, troubleshooting deployment failures, and monitoring vulnerability compliance across large server fleets.

Azure

.NET

🕒 August 13

BPCS, Comprehensive marketing solutions, ltd.

1 - 10

💼 Consulting

📣 Marketing

Substrate PAVC Engineer supporting Blueprint’s Microsoft engagement. Automating Windows patching, troubleshooting deployments, and monitoring vulnerability compliance across distributed server environments.

Azure

.NET