Senior Solutions Architect – AI Infrastructure Networking

🔥 1 hour ago

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 21%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mirantis

Mirantis

501 - 1000 employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Mirantis is a company that specializes in container management and cloud infrastructure solutions. It offers a range of products, including Mirantis Kubernetes Engine (MKE), Mirantis OpenStack for Kubernetes (MOSK), and Mirantis Container Cloud (MCC), which provide enterprise-level Kubernetes and container management platforms. Mirantis also develops tools for secure software supply chains, such as the Mirantis Container Runtime (MCR) and Mirantis Secure Registry (MSR). As an advocate for open source technologies, Mirantis supports various projects and provides resources like Lens Desktop, a popular Kubernetes IDE, and technical support for enterprises adopting cloud-native technologies. Their solutions cater to sectors such as public services, financial services, and broader SaaS and technology services industries.

📋 Description

• Own the network solutioning of k0rdent AI, Mirantis's platform for building and operating GPU clouds and AI factories • Design large GPU cluster interconnects, tenant isolation, and workload connectivity from containers, virtual machines and bare-metal nodes • Design and publish network reference architectures and solution designs from single-rack to multi-thousand-GPU clusters • Define InfiniBand and RoCEv2 Ethernet compute fabrics, including rail-optimised and fat-tree/Clos designs, oversubscription and failure domains • Define front-end, storage and management networks and connectivity to customer data centers, public clouds and hybrid environments • Specify multi-tenant isolation using InfiniBand partitions, VRFs, EVPN-VXLAN and Kubernetes network policy • Document scale limits and trade-offs involving cost, performance, operability and vendor lock-in • Define Linux host networking and expose NICs, DPUs and SuperNICs to workloads • Design Kubernetes networking for AI workloads, including CNIs, Multus, SR-IOV, RDMA device plugins, NVIDIA Network Operator and DRA • Design KubeVirt networking for virtual machines, including passthrough and SR-IOV for GPU and RDMA traffic • Research and evaluate emerging AI networking technologies and standards • Prototype alternative designs in the lab and measure them against current practice • Publish internal research notes, design proposals and selected external papers, blog posts or talks • Build and run proofs of concept with customers, partners and on Mirantis or customer hardware • Write automation and tooling using Python, Go, Bash, Ansible, Helm, Kubernetes manifests and Terraform • Validate and benchmark fabrics and host configurations using tools such as NCCL tests, perftest and ib_write_bw • Act as the network subject matter expert in customer discovery, design reviews and architecture workshops • Work with hardware and networking partners on joint designs and validations • Feed research results back to Product and Engineering and help shape the k0rdent AI roadmap • Present at industry events, webinars and partner summits • Run hands-on technical workshops and write reference architectures, solution briefs, blog posts and enablement content

🎯 Requirements

• Bachelor's degree in Computer Science, Electrical Engineering, Telecommunications or a related field, or equivalent practical experience • 8+ years in network engineering or network architecture • At least 3 years in data center, HPC or cloud infrastructure networking • Customer-facing experience as a solutions architect, pre-sales engineer, consultant or technical lead • Expertise in data center network design, including spine-leaf and Clos topologies, rail-optimised GPU fabrics, oversubscription and ECMP • Expertise in InfiniBand subnet management, partitioning, adaptive routing, and NCCL/RDMA traffic • Expertise in RoCEv2 on Ethernet, including PFC, ECN, DCQCN, QoS, MTU and buffer tuning • Expertise in BGP, EVPN-VXLAN and VRFs for multi-tenancy • Experience with hybrid and multi-cloud connectivity, interconnects, VPN, cloud networking, IP addressing and DNS planning • Linux networking knowledge, including NIC drivers, PCI passthrough, SR-IOV, IOMMU, NUMA affinity, iproute2, ethtool and devlink • Kubernetes networking experience with CNI plugins, Multus, SR-IOV, RDMA device plugins, network policy and service exposure • Experience with KubeVirt and VM networking, passthrough and SR-IOV • Programming or scripting in at least one language; Python or Go preferred • Comfortable using Git, CI and infrastructure-as-code • Excellent written and spoken English • Comfortable presenting to large audiences and running hands-on workshops • Able to work across time zones and different work cultures • Remote candidates must be based in the United States East Coast • Ability to work some meetings outside standard local hours • Nice-to-have experience with NVIDIA networking, GPU cloud/neocloud/HPC operators, bare-metal provisioning, network automation, SDN controllers, global load balancing, NVMe-oF, Mirantis products, open-source contributions, industry standards, published research, patents or white papers, and additional languages

🏖️ Benefits

• Remote work arrangement • Travel of up to 25% for customer engagements, partner meetings, lab work and industry events • Collaborate with a world-class, distributed team • Opportunity to work directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises • Opportunity to shape the product narrative and influence go-to-market success

Apply Now

Similar Jobs

🔥 1 hour ago

Pure Storage

1001 - 5000

💼 Consulting

📦 Logistics

🏥 Healthcare

Cyber Resilience Solutions Architect leading technical sales enablement for Everpure’s data platform. Supporting public-sector customers, partners, demos, and go-to-market strategy.

🔥 1 hour ago

Tanium

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Solution Engineer guiding Tanium’s autonomous IT platform evaluations for U.S. state, local, and higher-education customers. Leading demos, Proof of Value engagements, and technical sales strategy.

🔥 2 hours ago

Broadcom

10,000+ employees

🏭 Manufacturing

📦 Logistics

💼 Consulting

ValueOps Solution Engineer guiding enterprise customers on Broadcom’s portfolio planning and agile management software. Delivering demos, adoption roadmaps, and strategic value engagements.

🔥 2 hours ago

HIKE2

51 - 200

💼 Consulting

🤖 Artificial Intelligence

📋 Compliance

Salesforce consultancy architect designing scalable OmniStudio solutions for complex client transformations. Advising stakeholders, leading integrations, and shaping Salesforce practice standards.

🔥 2 hours ago

U.S. Financial Technology

201 - 500

💳 Fintech

☁️ SaaS

🤝 B2B

Lead AI engineer designing secure AWS, Snowflake, LLM, and agentic AI solutions. Supporting U.S. FinTech’s mortgage securitization platform with scalable data pipelines and automation.