Technical Lead – GPU Infrastructure

🕒 vor 1 Tag

🇮🇹 Italien – Remote

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

👻 Geisterscore 25%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Tether.to

Tether.to

11 - 50 Mitarbeiter

Gegründet 2014

₿ Crypto

💳 Fintech

💸 Finanzen

Crypto • Fintech • Finance

Tether. to ist ein führendes Unternehmen für digitale Vermögenswerte, das die Nutzung von Stablecoins im Blockchain-Bereich vorantreibt. Als der am weitesten verbreitete Stablecoin sind Tether-Token darauf ausgelegt, im Verhältnis 1:1 an Fiatwährungen gekoppelt zu sein, und bieten Nutzern eine stabile digitale Vermögensoption. Die Plattform erleichtert diese Token-Transaktionen über mehrere Blockchains hinweg und fördert grenzüberschreitende Transaktionen, während durch tägliche Aufzeichnungen der Gesamtvermögenswerte und Reserven Transparenz gewahrt wird. Zu den Initiativen von Tether gehören Bildungsprogramme zur Förderung der Nutzung digitaler Vermögenswerte, insbesondere in Regionen wie dem Nahen Osten, der Türkei und den Philippinen. Tether positioniert sich somit als Disruptor des traditionellen Finanzsystems, indem es eine stabile, effiziente Methode für Transaktionen in der digitalen Währungswelt ermöglicht.

Beschreibung

• Own the end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and maintenance of the baseline • Lead and line-manage a distributed team of approximately twelve engineers across backend, frontend, DevOps, QA, and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, draining, autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Manage NVIDIA GPU Operator and Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement • Own managed inference architecture, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and design a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and advise on capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and set the technical bar for new engineers • Own architecture, implementation, and delivery plans within a fixed initial six-month delivery window

🎯 Anforderungen

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running • Experience operating an HPC or GPU training cluster for a research population ideally • Bare-metal NVIDIA GPU fleet operation, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Strong Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy • HPC storage and data movement experience with VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Observability and operations experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions • Experience delivering a multi-tenant IaaS, PaaS, or research computing service with isolation, quotas, usage metering, and user-facing API and CLI surfaces • People management across time zones, cross-track review, written architecture decisions, and ability to challenge partners or executives with reasons • Excellent written and spoken English • Must be based between UTC and UTC+5:30 • Desirable: Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable: modern serving stacks such as vLLM, SGLang, and TensorRT-LLM • Desirable: VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and hardware-provider partnership experience

🏖️ Vorteile

• Fully remote work arrangement • Occasional travel to partner sites and team events

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 8 Tagen

WeRoad

51 - 200

🏨 Gastgewerbe

📦 Logistik

✈️ Reisen

Senior full-stack engineer building scalable APIs, web interfaces, and internal tools for WeRoad, a global group-adventure travel company. Leading product architecture across Vue, Nuxt, NestJS, and TypeScript.

🇮🇹 Italien – Remote

💵 €75.000 - €90.000 / Jahr

💰 Series B im 2023-11

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 8 Tagen

WeRoad

51 - 200

🏨 Gastgewerbe

📦 Logistik

✈️ Reisen

Senior Full Stack Product Engineer building scalable Vue/NestJS travel platforms for WeRoad’s group adventures. Delivering product features, architecture, and performance across remote European markets.

🇮🇹 Italien – Remote

💵 €55.000 - €65.000 / Jahr

💰 Series B im 2023-11

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Chiesi Group

5001 - 10000

🏥 Gesundheitswesen

🏭 Fertigung

💊 Pharmazie

Senior CMC leader directing development and IMP supply for Chiesi’s rare-disease medicines. Managing regulatory strategy, budgets, CROs/CDMOs, and technical teams across Europe.

🇮🇹 Italien – Remote

💵 €68.480 / Jahr

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Chiesi Group

5001 - 10000

🏥 Gesundheitswesen

🏭 Fertigung

💊 Pharmazie

Senior CMC leader guiding recombinant protein drug-substance development for Chiesi’s rare-disease biopharmaceuticals. Shaping expression, process, quality, regulatory, and lifecycle strategies.

🇮🇹 Italien – Remote

💵 €68.480 / Jahr

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Chiesi Group

5001 - 10000

🏥 Gesundheitswesen

🏭 Fertigung

💊 Pharmazie

Senior analytical CMC leader guiding validated methods and control strategies for Chiesi’s rare-disease biologics. Leading technical teams, partners, and regulatory-ready development from candidate selection through lifecycle management.

🇮🇹 Italien – Remote

💵 €68.480 / Jahr

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich