
2 - 10 employees
🤝 B2B
💼 Consulting
B2B • Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
🔥 20 hours ago
Improve your chances of getting an interview by checking your resume score before you apply.

2 - 10 employees
🤝 B2B
💼 Consulting
B2B • Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
• Evaluate GPU and accelerator kernels for technical correctness and completeness • Review implementations developed from specifications or reference operators • Assess mathematical behavior, implementation defects, unsupported assumptions, and incomplete solutions • Review kernel development across CUDA, Triton, NKI, and Pallas (JAX) • Evaluate framework-specific implementation choices, execution constraints, translations, and migrations • Compare alternative kernel implementations for correctness and technical quality • Assess outputs against reference implementations using absolute, relative, and ULP-based tolerances • Review floating-point behavior and precision-related edge cases • Evaluate performance using Nsight, Nsight Compute, roofline analysis, and framework-native profilers • Assess benchmarking methodology, latency, throughput, utilization, memory behavior, and claimed improvements • Review compute- and memory-efficiency optimization strategies, including tiling, vectorization, parallelization, and workload decomposition • Review registers, shared memory, caches, memory access, data locality, bandwidth utilization, bank conflicts, coalescing, and register pressure • Diagnose compilation and runtime failures involving drivers, out-of-memory conditions, launch configurations, shape or stride mismatches, and autotuning • Validate task reliability in intended environments and evaluate debugging approaches • Review kernel translation, lowering, hardware migration, compiler transformations, and intermediate-representation decisions • Diagnose kernel implementations using outputs, profiler data, runtime behavior, and source code • Evaluate operator fusion and whether fused kernels preserve intended semantics • Assess assigned tasks against structured technical criteria and provide evidence-based written evaluations • Distinguish valid implementation alternatives from technically flawed approaches
• 3+ years of hands-on experience developing, optimising, or verifying GPU or accelerator kernels • Practical experience with at least two of CUDA, Triton, NKI, or Pallas (JAX) • Strong understanding of numerical correctness, including absolute, relative, and ULP tolerances • Experience selecting and validating appropriate reference implementations • Strong performance profiling and benchmarking experience • Familiarity with Nsight, Nsight Compute, roofline analysis, or comparable profiling tools • Strong understanding of common kernel compilation and runtime failure modes • Experience with at least three of: kernel generation from specification, framework translation or lowering, hardware-target migration, kernel debugging, performance optimisation, operator fusion • Experience across both NVIDIA GPU and custom-accelerator ecosystems preferred • Background in compiler engineering, MLIR, or intermediate-representation lowering advantageous • Strong understanding of memory-hierarchy optimisation preferred • Contributions to kernel or accelerator libraries such as cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls advantageous • Strong written communication and ability to provide precise technical feedback • Must complete work without using confidential or proprietary information belonging to any employer, client, institution, or other third party • H1-B and STEM OPT support unavailable
• Part-time independent contractor engagement • Fully remote within the United States • Flexible scheduling based on project requirements • Projects may be extended, shortened, or concluded based on project needs and performance • H1-B and STEM OPT support is unavailable for this engagement
Apply Now🕒 August 21
Part-time transmission engineer designing high-voltage overhead and underground utility systems for Leidos. Applying engineering standards and preparing analyses, specifications, procurement, and construction documents.
🇺🇸 United States – Remote
💵 $73.5k - $132.8k / year
⏱ Part Time
🟠 Senior
👷🏻♀️ Engineer
🦅 H1B Visa Sponsor
🕒 August 11
Part-time OCR Engineer developing classification, extraction, validation, and automation for federal document records. Supporting reliable, traceable information processing with minimal manual review.
🕒 July 30
Chemical Engineering Expert reviewing and validating AI-generated chemical engineering content. Collaborating with experts to enhance AI systems in chemical engineering applications.
🕒 July 27
Prompt Engineer responsible for end-to-end technical migration workflow for transitioning templates to LLM autoraters. Leveraging prompt engineering techniques for maximum model performance using internal tools.
🕒 July 27
Prompt Engineer managing end-to-end technical migration workflow for transitioning templates to LLM autoraters. Leveraging prompt engineering techniques with client’s internal tools to maximize model performance.