Member of Technical Staff – Storage Infrastructure

🕒 vor 19 Tagen

🏄 California – Remote

infoinfo

💵 $150.000 - $300.000 / Jahr

⏰ Vollzeit

🔴 Experte

🖥 Softwareentwickler

👻 Geisterscore 0%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Prime Intellect

Prime Intellect

1 - 10 Mitarbeiter

🤖 Künstliche Intelligenz

☁️ SaaS

Artificial Intelligence • SaaS • Cloud Computing

Prime Intellect ist ein Unternehmen, das sich darauf konzentriert, die KI-Entwicklung zu demokratisieren, indem es skalierbare und dezentrale Computerressourcen für das Training von Modellen bereitstellt. Ihre Plattform ermöglicht es Nutzern, weltweit verfügbare Computerressourcen zu finden und zu teilen, um den Training von hochmodernen Modellen durch verteilte Cluster zu ermöglichen. Sie fördern das kollektive Eigentum an KI-Innovationen, einschließlich Sprach- und wissenschaftlicher Modelle. Prime Intellect bietet zudem eine Reihe von GPU-Optionen, um kostengünstiges und effizientes Modelltraining zu erleichtern. Sie streben danach, die Forschung im Bereich des dezentralen Trainings und die Entwicklung von Open-Source-KI weltweit voranzutreiben.

Beschreibung

• Design and operate storage architectures for training datasets, checkpointing, inference artifacts, and shared research workflows • Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads • Benchmark throughput, latency, metadata performance, and concurrent access with representative training and checkpoint workloads • Build provisioning, capacity planning, lifecycle management, and operational automation for storage services • Design and test replication, recovery, backup, and failure-handling procedures with explicit durability and availability targets • Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices • Implement access controls, tenant separation, quotas, monitoring, and runbooks • Collaborate with compute and networking teams • Work directly with customers pushing the boundaries of AI • Collaborate with the engineering team on systems powering next-generation AI breakthroughs

🎯 Anforderungen

• 3+ years building or operating production distributed storage systems • Hands-on experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS • Strong Linux administration and performance troubleshooting skills • Experience automating infrastructure operations in Python, Go, Bash, or similar languages • Understanding of storage failure modes, data integrity, consistency, replication, and recovery • Knowledge of block, file, and object storage semantics and performance tradeoffs • Experience with NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking • Knowledge of high-throughput storage networking and distributed client behavior • Experience with capacity forecasting, observability, alerting, and safe maintenance procedures • Knowledge of authentication, authorization, encryption, and secure data lifecycle management • Experience supporting large GPU training clusters and high-volume checkpoint workloads • Experience with S3-compatible object storage, data tiering, or distributed caching • Experience with RDMA-enabled storage or GPUDirect Storage • Experience with Kubernetes storage integrations or SLURM environments • Experience with storage cost optimization and contributions to open-source storage systems

🏖️ Vorteile

• Equity incentives

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 19 Tagen

Juniper Square

201 - 500

💸 Finanzen

🏠 Immobilien

☁️ SaaS

Engineering Director modernizing Juniper Square’s private-markets operations platform. Leading monolith modularization, AI-native development, application security, and engineering organization growth.

🇺🇸 Vereinigte Staaten – Remote

💵 $230.000 - $285.000 / Jahr

💰 €75.000.000 Series C im 2019-11

⏰ Vollzeit

🔴 Experte

🖥 Softwareentwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 19 Tagen

Comcast

10.000+ Mitarbeiter

📡 Telekommunikation

🏛️ Regierung

SAP ABAP and Fiori Developer building and maintaining applications for federal clients. Developing ABAP programs and SAP Fiori experiences using SAPUI5 and OData.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 19 Tagen

Re:Build Manufacturing

501 - 1000

💼 Beratung

🏥 Gesundheitswesen

📦 Logistik

Director leading Forward Deployed Engineers delivering full-stack digital solutions for Re:Build Manufacturing's industrial platform. Driving software delivery, AI adoption, and measurable operational impact across manufacturing sites.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🔴 Experte

🖥 Softwareentwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 19 Tagen

CACI International Inc

10.000+ Mitarbeiter

🎖️ Verteidigung

🏛️ Regierung

🔒 Cybersecurity

Qlik Developer transforming Oracle EBS financial data into Advana dashboards for federal clients. Supporting federal reporting, analytics, ETL, validation, and secure government-cloud deployments.

🇺🇸 Vereinigte Staaten – Remote

💵 $75.200 - $158.100 / Jahr

🔥 Finanzierung im letzten Jahr

💰 €500.000.000 Post-IPO Debt im 2026-02

⏰ Vollzeit

🟠 Senior

🔴 Experte

🖥 Softwareentwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 20 Tagen

First American

10.000+ Mitarbeiter

🏠 Immobilien

💸 Finanzen

🏢 Unternehmen

Engineering Director scaling First American’s AI-driven title automation platform. Owning document intelligence, data architecture, distributed teams, and production engineering outcomes.

🗣️🇺🇸🇬🇧 Englisch erforderlich

Cloud

Unity