Staff Network Engineer – AI Fabric, Datacenter and Edge Networking

Job not on LinkedIn

🔥 3 hours ago

🇪🇺 Europe – Remote

⏰ Full Time

🔴 Lead

🤖 AI Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Submer

Submer

51 - 200 employees

Founded 2015

🔧 Hardware

⚡ Energy

🤝 B2B

Hardware • Energy • B2B

Submer is a provider of connected intelligence for AI infrastructure, specializing in advanced liquid/immersion cooling and modular data-center solutions. They design, build and operate AI-ready environments (from power and land to cloud and edge), manufacture immersion cooling pods (SmartPod EVO/EXO), and offer GPUaaS/AIaaS and modular deployment services. Submer emphasizes energy and water savings, high-density thermal architectures for demanding AI workloads, and sovereign-ready, globally deployable solutions.

📋 Description

• Design, implement, and operate network infrastructure powering a GPU cloud platform for cloud gaming, AI, and machine learning applications in telecom carrier networks • Design and operate high-performance GPU networking fabrics and large-scale RoCE fabrics • Optimize network performance for GPU communication patterns and east-west traffic • Define reusable AI-fabric reference architectures and design principles • Design and operate Layer-2 and Layer-3 datacenter networks using scalable routing architectures • Maintain tenant isolation, ingress/egress routing, traffic management, and overlay networking • Deploy and maintain north-south security infrastructure, WAF protections, and platform security controls • Design private interconnects, dark fiber rings, high-capacity WAN connectivity, and global backbone integration • Lead end-to-end engineering delivery from design and lab validation through production deployment • Validate network bills of materials, contribute to datacenter layouts and rack elevations, and drive capacity planning • Establish deployment standards, validation criteria, rollback approaches, and reusable acceptance patterns • Own networking operational performance and reliability • Automate provisioning, configuration management, monitoring, and lifecycle management • Lead incident response and root-cause analysis for major network events • Define and track SLAs, SLOs, reliability metrics, and operational benchmarks • Collaborate with infrastructure, platform, SRE, compute, storage, observability, and datacenter operations teams • Act as the primary networking design authority and influence the platform networking roadmap • Mentor engineers and help build the networking function

🎯 Requirements

• Strong hands-on experience designing and operating large-scale datacenter networks • Expert knowledge of BGP, OSPF, ECMP, and EVPN/VXLAN • Proven experience operating high-speed Ethernet networks in production environments • Experience operating NVIDIA/Mellanox networking platforms • Deep expertise designing and operating networking fabrics for large-scale GPU clusters and distributed AI workloads • Deep understanding of NCCL communication patterns • Experience tuning RoCE fabrics • Strong knowledge of RDMA transport behavior and failure modes • Practical experience implementing and tuning PFC and ECN • Understanding of GPU collective communication patterns • Experience designing rail-optimized GPU networking fabrics • Understanding of networking performance impacts on PyTorch and TensorFlow • Ability to debug cross-layer issues involving hardware, firmware, kernel networking, and distributed application communication layers • Strong knowledge of networking hardware, optics, and high-speed interconnects • Experience designing network observability systems • Strong automation skills using Python and/or Bash • Experience applying software engineering practices to infrastructure automation • Proven ability to lead complex technical initiatives across teams • Demonstrated ability to set architectural direction and drive adoption of engineering standards • Strong mentoring capability • Experience owning architecture and direct implementation in lean or fast-scaling environments is strongly preferred

🏖️ Benefits

• Attractive compensation package reflecting your expertise and experience • Flexible work environment • Hybrid-friendly approach • International and diverse work environment • Career evolution opportunities in a fast-growing scale-up

Apply Now

Similar Jobs

🕒 6 days ago

Virtusa

10,000+ employees

💼 Consulting

🏢 Enterprise

🤖 Artificial Intelligence

Cloud AI Engineer building Agentic AI workflows and cloud machine learning solutions. Applying Vertex AI, MLOps, and Generative AI for IT clients across Europe.

🇪🇺 Europe – Remote

💰 $108M Post-IPO Equity - Virtusa on 2017-05

⏰ Full Time

🟠 Senior

🔴 Lead

🤖 AI Engineer

Apache

Cloud

ETL

Hadoop

Java

Perl

Python

PyTorch

Spark

Tensorflow

C++

🕒 July 28

Wand AI

51 - 200

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Staff Software Engineer building AI systems for governance at Wand. Analyzing organizational data and creating agentic workflows with autonomy.

Neo4j