Software Engineer, Distributed Systems

🔥 9 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of fal

fal

51 - 200 employees

🤖 Artificial Intelligence

🔌 API

☁️ SaaS

Artificial Intelligence • API • SaaS

fal is a generative media platform for developers that provides access to a large gallery of production-ready image, video, audio and 3D generative models alongside serverless GPU inference and on-demand compute clusters for training and fine-tuning. The platform offers unified APIs and SDKs to call hundreds of open models or private weights, a high-performance inference engine, managed serverless GPU deployments, and dedicated clusters with modern NVIDIA hardware for large-scale training. fal targets developer and enterprise customers with features like SOC 2 compliance, private endpoints, usage analytics, and enterprise support, and is positioned for building, deploying, and scaling generative AI-powered products.

📋 Description

• Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etc • Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world • Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems • Profile and tune low level CPU and memory performance

🎯 Requirements

• 5+ years experience building distributed compute and orchestration platforms in Python or Rust • Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning • Deep understanding of computational complexity and memory allocation • Track record of designing systems that scale under real production load • Experience building and using observability to drive performance and reliability decisions • Excellent communication and ability to drive technical decisions across teams • Self-starter who executes quickly, takes ownership, and constantly seeks improvement • Nice to have: Experience with AI/ML inference or training infrastructure • Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency) • Background in building multi-tenant compute platforms • Understanding of networking fundamentals and performance characteristics • Familiarity with GPU workload characteristics and scheduling constraints

🏖️ Benefits

• Interesting and challenging work • A lot of learning and growth opportunities • Regular team events and offsites

Apply Now