
1 - 10 employees
Founded 2024
🤖 Artificial Intelligence
☁️ SaaS
Artificial Intelligence • Cloud Computing • SaaS
Yotta Labs is building the DeOS for AI optimization and orchestration at planet scale. The company provides a high-performance framework to aggregate geo-distributed GPUs, delivering high throughput for heterogeneous compute resources. Yotta Labs focuses on affordability and accessibility for AI training and inference across a spectrum of GPUs, including commodity to high-end options. Their platform supports large language models (LLMs) and enables users to fine-tune AI applications seamlessly, deploying secure AI agents in the cloud with optimized API endpoints.
🔥 1 minute ago
🌐 United States, Hong Kong, +2 more countries – Remote
👨🎓 Internship
⚪️ Entry-level
⚙️ Systems Engineer
Improve your chances of getting an interview by checking your resume score before you apply.

1 - 10 employees
Founded 2024
🤖 Artificial Intelligence
☁️ SaaS
Artificial Intelligence • Cloud Computing • SaaS
Yotta Labs is building the DeOS for AI optimization and orchestration at planet scale. The company provides a high-performance framework to aggregate geo-distributed GPUs, delivering high throughput for heterogeneous compute resources. Yotta Labs focuses on affordability and accessibility for AI training and inference across a spectrum of GPUs, including commodity to high-end options. Their platform supports large language models (LLMs) and enables users to fine-tune AI applications seamlessly, deploying secure AI agents in the cloud with optimized API endpoints.
• Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium. • Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA. • Profile and improve inference performance in vLLM, SGLang, and our custom runtimes — kernel fusion, scheduling, KV-cache and memory optimizations. • Build benchmarks, chase down performance regressions, and turn profiler traces into concrete speedups. • Ship code upstream to open-source AI infrastructure projects, with tests and documentation.
• Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field • Solid programming skills in Python and familiarity with C++ • Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy) from coursework, research, or projects • Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels — class projects and personal projects count • Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler) • Strong problem-solving skills and the ability to work independently in a collaborative, remote environment.
• Competitive internship compensation • Flexible remote work environment • Direct mentorship from engineers from leading institutions and tech companies • Fast path to a full-time return offer for top performers
Apply Now