Senior Cloud Infrastructure and Network Operations Solutions Architect

🔥 0 minutes ago

🌐 United Kingdom, Spain, +2 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

💻 Solutions Engineer

🇬🇧 UK Skilled Worker Visa Sponsor

infoinfo

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Engage directly with customers, partners and cross-functional teams to assess, architect and operate compute, storage and scale-up fabrics supporting large GPU estates • Own the fabric side of NVIDIA’s Cloud Partner operating model across the full Day 1 to Day 2 lifecycle • Manage NVLink/NVSwitch partition operations, maintenance-partition isolation and safe partition-change workflows for multi-tenant estates • Own switch software and firmware lifecycle, including Cumulus Linux, SONiC and switch-OS upgrades, CPLD and field-notice rollout campaigns • Support Day 1 fabric validation and acceptance, including InfiniBand/UFM bring-up, Spectrum-X/RoCE Ethernet configuration, cabling and link-health verification, routing and congestion-control validation, and multi-day burn-in • Minimise time from cluster handover to first production workload by removing duplicated fabric validation • Drive fabric reliability at fleet scale through telemetry, fault detection, remediation, root-cause analysis and improvement of MTBI and job goodput • Provide consultative guidance and hands-on troubleshooting across NICs, DPUs, switch OS, routing, congestion control, host networking and Kubernetes integration • Act as technical leader for assigned accounts, run structured knowledge transfer and produce runbooks for partner teams

🎯 Requirements

• BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience • 5+ years of professional experience in data centre networking, fabric engineering or large-scale network operations roles • Deep understanding of data centre network architectures and RDMA fabrics—InfiniBand and RoCE/Ethernet—including topology and routing design, congestion control, lossless/QoS configuration, and troubleshooting across NICs, switches and high-speed interconnects • Hands-on experience with InfiniBand and UFM, Spectrum-X Ethernet, Cumulus Linux and/or SONiC, ConnectX and BlueField NICs/DPUs, and NVLink/NVSwitch systems • Deep knowledge of Linux (RedHat, Ubuntu), switch operating systems, network OS internals, host networking and OS-level security • Understanding of storage and compute traffic patterns of HPC/AI clusters • Proficiency in Python and Bash scripting, configuration management, Infrastructure-as-Code tools such as Ansible and Terraform, GitOps-based network configuration and firmware/upgrade management, and observability stacks including Grafana, Loki and Prometheus • Demonstrated ability to measure and improve MTBI and job goodput on large GPU clusters, including fault detection, drain and remediation workflows, SLO/error-budget definition and post-incident review • Strong consultative background leading architectural reviews and presenting to executive stakeholders • Experience with Kubernetes networking in GPU clusters, including CNI plugins, multi-network attachment, SR-IOV and RDMA device plugins • Experience with NVIDIA Network and GPU Operators • Expertise with DOCA and DPU infrastructure services such as DOCA DPF • Familiarity with fabric and GPU health telemetry, including DCGM and XID diagnostics, link-level and switch counters, node-level health agents, and fleet-wide reliability intelligence • Knowledge of large-scale training traffic behaviour, including NCCL tuning, rail alignment, and diagnosing network-bound performance regressions • Experience delivering multi-tenant network isolation using VRF/VLAN/EVPN segmentation, tenant partitioning and secure fabric handover

Apply Now

Similar Jobs

🔥 22 hours ago

OX Security

51 - 200

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

Solutions Engineer driving technical sales, demos, and proof-of-value engagements for OX’s AI-driven cybersecurity platform. Partnering with Sales, Product, Engineering, and enterprise customers.

🕒 Yesterday

Salesforce

10,000+ employees

💼 Consulting

📣 Marketing

☁️ SaaS

Presales Data and Integration Solution Engineer presenting Salesforce data, integration, analytics, and governance solutions. Supporting UK public-sector and not-for-profit customers through discovery, demonstrations, and technical sales engagements.

🕒 2 days ago

Palo Alto Networks

10,000+ employees

🔒 Cybersecurity

🏢 Enterprise

Solutions Consultant guiding UK higher-education and research customers through Palo Alto Networks security transformations. Designing, presenting, and validating cybersecurity solutions while driving adoption and business value.

🕒 2 days ago

Daisy Group™

1001 - 5000

📡 Telecommunications

🤝 B2B

🔒 Cybersecurity

Cyber Presales Solution Architect designing secure IT and cybersecurity solutions for Wavenet’s managed services customers. Shaping bids, proposals, and resilient architectures with sales, engineering, and executives.

🕒 2 days ago

Codec Ireland

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Solution Architect shaping enterprise Dynamics 365 and Power Platform solutions for Codec, a Microsoft Solutions Partner. Leading architecture, integration, security, Azure, governance and technical strategy.