Infrastructure Engineer – Storage Platform

🔥 1 minute ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of TensorWave

TensorWave

11 - 50 employees

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

TensorWave is a cutting-edge AI compute company offering advanced solutions for enterprises focusing on training, fine-tuning, and inference. Powered by AMD's Instinct MI300X accelerators, TensorWave delivers a cost-effective and high-performance alternative to Nvidia's H100 chip. The company provides options for both bare-metal nodes and fully-managed Kubernetes clusters, ensuring ease of use with native support for PyTorch and TensorFlow. TensorWave is committed to data security with SOC2 Type II and HIPAA compliance, making it a reliable partner for businesses requiring dedicated and secure environments.

📋 Description

• Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) • Manage storage lifecycle operations - cluster expansion, upgrades and migrations • Monitor and maintain storage health, including capacity utilization, data distribution and balance, cluster state and recovery operations • Analyze and troubleshoot storage performance across IOPS, throughput, and latency (including tail latency) • Identify and remediate bottlenecks across disk subsystems, network paths (including RDMA where applicable), client access patterns • Support incident response and root cause analysis for storage-related issues • Ensure storage platforms meet performance expectations for GPU and Kubernetes workloads • Operate and support Kubernetes-integrated storage - CSI drivers, StorageClasses, PersistentVolumes / PersistentVolumeClaims • Troubleshoot storage-related issues in Kubernetes environments, including stateful workloads, performance inconsistencies, scheduling and provisioning failures • Execute and improve automation for storage deployment and operations using Ansible, Terraform, Kubernetes manifests / Helm • Contribute to improving monitoring and alerting, operational workflows, runbooks and documentation • Partner with DevOps and Platform Engineering (automation and orchestration), Network Engineering (high-throughput and RDMA networking), Compute / Virtualization teams • Help ensure end-to-end performance across compute, network, and storage layers

🎯 Requirements

• 4–7+ years of experience in infrastructure, systems, or storage operations • Strong hands-on experience operating distributed storage systems in production • Experience with Ceph (RBD, CephFS, or RGW) • Strong Linux systems knowledge • Experience with modern storage platforms such as: • Weka, VAST Data, or similar high-performance systems • Solid understanding of: • Storage performance characteristics (IOPS, throughput, latency) • Data replication and failure domains • Ability to troubleshoot across: • Storage systems • Network paths • Compute clients

🏖️ Benefits

• Stock Options • 100% paid Medical, Dental, and Vision insurance for Employees • Company Health Savings Account Contributions • 100% paid Short Term and Long Term Disability Insurance for Employees • Life and Voluntary Supplemental Insurance Options • Other Insurance Options, such as Pet & Legal Insurance • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support • Flexible Spending Account • 401(k) • Employee Assistance Program • Flexible PTO • Paid Holidays • Parental Leave • Other In-Office Perks

Apply Now

Similar Jobs

🔥 7 hours ago

Harrington Process Solutions

501 - 1000

🤝 B2B

🏭 Manufacturing

🍽️ Food & Beverage

Hands-on Infrastructure Engineer supporting IT infrastructure across on-premise and cloud environments. Key role in managing operations, leading initiatives, and collaborating across teams.

🔥 8 hours ago

Dayforce

5001 - 10000

👥 HR Tech

☁️ SaaS

🏢 Enterprise

IT-OT Infrastructure & Cyber Architect defining secure, resilient architectures for enterprise systems. Collaborating across IT, OT, and cybersecurity to meet business objectives.

🇺🇸 United States – Remote

💵 $130.7k - $196.1k / year

💰 $1G Post-IPO Debt - Dayforce on 2024-03

⏰ Full Time

🟠 Senior

🔴 Lead

👷 Infrastructure Engineer

🔥 9 hours ago

Centralize

2 - 10

☁️ SaaS

🤝 B2B

🏢 Enterprise

Infrastructure Engineer focusing on scalability for Centralize's backend systems and integration pipelines handling millions of events daily. Working with Postgres, AWS, and Redis in a fast-growing startup environment.

🔥 10 hours ago

The Voleon Group

201 - 500

💸 Finance

🤖 Artificial Intelligence

Software Engineer contributing to data infrastructure at Voleon, focusing on scalable systems and collaboration with teams. Engaging in mentorship and technical guidance within the engineering culture.

🔥 10 hours ago

The Voleon Group

201 - 500

💸 Finance

🤖 Artificial Intelligence

Senior/Staff Software Engineer in Data Infrastructure at Voleon, building scalable data infrastructure and collaborating with teams to enhance their data use.