
51 - 200 employees
Founded 2022
π€ Artificial Intelligence
βοΈ SaaS
π€ B2B
π° $20M Seed on 2024-06
Artificial Intelligence β’ SaaS β’ B2B
Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.
π₯ 55 minutes ago
πΊπΈ United States β Remote
π΅ $150k - $240k / year
β° Full Time
π‘ Mid-level
π Senior
π Backend Engineer
Linux
NFS
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
Founded 2022
π€ Artificial Intelligence
βοΈ SaaS
π€ B2B
π° $20M Seed on 2024-06
Artificial Intelligence β’ SaaS β’ B2B
Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.
β’ Own Distributed Storage Architecture: Define, evolve, and operate Runpodβs global storage platforms, supporting training, inference, checkpointing, and dataset access at scale. β’ Build the Storage Engineering Team: Manage and grow a team of storage and systems engineers. Set clear ownership, technical direction, and operational standards across regions. β’ High-Performance Shared Filesystems: Design and operate large-scale SAN and NFS deployments, including performance-sensitive shared storage for GPU clusters. β’ Advanced Filesystems & Platforms: Lead deployments and operations of VAST Data and experience with Lustre or similar parallel filesystems used in HPC and AI environments. β’ End-to-End Performance Ownership: Drive performance optimization from NAND and NVMe media through controllers, networking, and client access patterns. β’ Next-Generation Storage Technologies: Evaluate and deploy cutting-edge capabilities such as NFS over RDMA, GPU Direct Storage (GDS), and low-latency data paths for accelerated workloads. β’ Reliability & Scale: Establish best practices for replication, data tiering, data protection, failure recovery, capacity planning, and lifecycle management. β’ Automation & Observability: Build automation for provisioning, expansion, upgrades, and monitoring. Ensure deep observability into throughput, latency, and error characteristics. β’ Cross-Functional Collaboration: Partner with Datacenter Networking, GPU Platform, SRE, and Product teams to ensure storage systems meet evolving workload and customer needs. β’ Vendor & Partner Management: Own technical relationships with storage vendors, hardware partners, and colocation providers; drive roadmap alignment and issue resolution.
β’ Engineering Leadership Experience: 3+ years managing storage, systems, or infrastructure engineering teams in production environments. β’ Distributed Storage Expertise: 8+ years designing and operating large-scale storage systems, including SAN and NFS architectures at multi-petabyte scale. β’ VAST Data Experience: Hands-on experience deploying, operating, or deeply integrating VAST Data in production environments is required. β’ Parallel Filesystems: Experience with Lustre or comparable HPC filesystems (e.g., GPFS, BeeGFS) supporting high-concurrency workloads. β’ Low-Level Storage Knowledge: Deep understanding of NAND, NVMe, PCIe, storage controllers, and performance characteristics across the stack. β’ High-Performance Data Paths: Proven experience with NFS over RDMA, RDMA-capable transports, or similar technologies. Familiarity with GPU Direct Storage strongly preferred. β’ Linux Systems Expertise: Strong Linux internals knowledge, including filesystems, I/O scheduling, memory management, and tuning for performance workloads. β’ Operational Excellence: Experience running 24/7 storage platforms with strong incident response, change management, and post-mortem discipline. β’ Communication & Leadership: Ability to clearly communicate complex technical tradeoffs and lead teams through high-stakes infrastructure decisions. β’ Successful completion of a background check.
β’ Meaningful equity in a fast-growing company- everyone on the team receives stock options β your impact drives our growth, and you share in the upside. β’ Generous medical, dental & vision plans β’ Flexible PTO- take the time you need to recharge β’ Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication β’ Join a passionate team on the cutting edge of AI infrastructure β where culture, learning, and ownership are at the heart of how we scale.
Apply Nowπ₯ 1 hour ago
Backend Engineer owning core systems for Nebulock's threat hunting solutions. Focused on scalable data ingestion and processing for security telemetry.
πΊπΈ United States β Remote
π₯ Funding within the last year
π° $6M Seed on 2025-08
β° Full Time
π‘ Mid-level
π Senior
π Backend Engineer
π₯ 1 hour ago
Backend Developer designing and maintaining backend services and APIs for CelebriOS products. Collaborating with engineering and business stakeholders for scalable solutions.
π₯ 1 hour ago
Software Engineer developing core database features for VillageSQL, a community-driven database technology. Join a passionate team building innovative solutions and engaging with open-source communities.
πΊπΈ United States β Remote
π΅ $118k - $250k / year
β° Full Time
π‘ Mid-level
π Senior
π Backend Engineer
π₯ 1 hour ago
Senior Software Engineer developing scalable backend solutions for TRM's AI-powered crime investigation platform. Collaborate on building APIs and features for internal and external stakeholders in a remote role.
πΊπΈ United States β Remote
π΅ $210k - $240k / year
β° Full Time
π Senior
π Backend Engineer
π₯ 1 hour ago
Software Engineer focusing on building third-party API integrations for TRMβs AI-native products. Collaborate with teams to improve integration processes and systems.
πΊπΈ United States β Remote
π΅ $180k - $240k / year
β° Full Time
π‘ Mid-level
π Senior
π Backend Engineer