AI Infrastructure Engineer

🕒 July 28

🇺🇸 United States – Remote

💵 $150k - $200k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 7%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of vCluster

vCluster

51 - 200 employees

☁️ SaaS

🏢 Enterprise

💰 $24M Series A on 2024-05

SaaS • Enterprise

vCluster is a Kubernetes virtualization product (from loft-sh) that creates virtual Kubernetes clusters on top of existing Kubernetes infrastructure. It enables dedicated, multi-tenant or private-node clusters for use cases like internal platform standardization, hybrid Kubernetes deployments, sovereign or dedicated customer environments, AI cloud providers and distributed inference. vCluster is positioned as a tooling/product solution for enterprises operating Kubernetes at scale, offering managed features, releases and integrations for internal K8s platforms and cloud-native workflows.

📋 Description

• Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment. • Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand. • Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s. • Knowledge Transfer: Work alongside customer teams to build self-sufficiency, ensuring they can operate and grow the platform independently. • Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer's head start. • Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap. • Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value.

🎯 Requirements

• Production K8s Mastery: 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments. • GPU Fluency: Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes. • Networking Fundamentals: Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments. • Storage Expertise: Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn. • Operational Agility: Comfort operating in ambiguous, fast-moving environments where you are often writing the playbook in real time. • Modern Tech Mindset: You thrive in environments that reject legacy tech and prefer a modern stack where you can solve a variety of problems from pipelines to internal services. • Bonus points for: Automation Skills: Experience writing automation scripts with Bash, Python, or Go. • Kubernetes Depth: Relevant certifications such as CKA (Certified Kubernetes Administrator) or experience writing Kubernetes Operators. • AI/ML Familiarity: Experience with inference serving, GPU scheduling, and the tooling around LLM deployment. • Documentation: Experience building AI Automation in documentation to contribute to a shared knowledge base.

🏖️ Benefits

• Competitive Salary: We offer a competitive compensation package, including equity. • Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country). • Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day. • Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.

Apply Now

Similar Jobs

🕒 July 28

Pilot.com

51 - 200

💸 Finance

🤝 B2B

☁️ SaaS

Senior Software Engineer focused on building and maintaining infrastructure at Pilot. Collaborating with teams to ensure efficient and scalable development processes.

🕒 July 28

George Jon

51 - 200

💼 Consulting

⚖️ Legal

☁️ SaaS

Senior Infrastructure Support Engineer for a tech advisory firm specializing in eDiscovery. Responsibilities include supporting secure and available eDiscovery platforms, servers, and networks.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

👷 Infrastructure Engineer

MS SQL Server

SQL

Switching

VMware

🕒 July 27

Atomic Industries

51 - 200

🏭 Manufacturing

🔧 Hardware

☁️ SaaS

Infrastructure Engineer managing and improving CI/CD pipelines, orchestrating workloads, and designing secure systems for simulation at Atomic Industries. Collaborating in a fast-paced engineering environment to scale infrastructure.

🇺🇸 United States – Remote

💰 $25M Series A - Atomic Industries on 2025-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

🕒 July 27

Ondo Finance

51 - 200

₿ Crypto

💳 Fintech

💸 Finance

Senior Security Engineer responsible for cloud security posture across AWS and GCPs. Leading security initiatives and mentoring engineers at Ondo Finance's remote team.

🇺🇸 United States – Remote

💰 Initial Coin Offering - Ondo Finance on 2024-01

⏰ Full Time

🟠 Senior

👷 Infrastructure Engineer

🕒 July 27

Law School Admission Council (LSAC)

201 - 500

📚 Education

🤝 Non-profit

👥 B2C

Infrastructure Engineer responsible for the maintenance of enterprise infrastructure platforms. Supporting network environment and collaborating with teams at LSAC.

🇺🇸 United States – Remote

💵 $87k - $94k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer