Technical Lead – GPU Infrastructure

🕒 vor 6 Tagen

🇷🇴 Rumänien – Remote

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

👻 Geisterscore 25%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Tether.to

Tether.to

11 - 50 Mitarbeiter

Gegründet 2014

₿ Crypto

💳 Fintech

💸 Finanzen

Crypto • Fintech • Finance

Tether. to ist ein führendes Unternehmen für digitale Vermögenswerte, das die Nutzung von Stablecoins im Blockchain-Bereich vorantreibt. Als der am weitesten verbreitete Stablecoin sind Tether-Token darauf ausgelegt, im Verhältnis 1:1 an Fiatwährungen gekoppelt zu sein, und bieten Nutzern eine stabile digitale Vermögensoption. Die Plattform erleichtert diese Token-Transaktionen über mehrere Blockchains hinweg und fördert grenzüberschreitende Transaktionen, während durch tägliche Aufzeichnungen der Gesamtvermögenswerte und Reserven Transparenz gewahrt wird. Zu den Initiativen von Tether gehören Bildungsprogramme zur Förderung der Nutzung digitaler Vermögenswerte, insbesondere in Regionen wie dem Nahen Osten, der Türkei und den Philippinen. Tether positioniert sich somit als Disruptor des traditionellen Finanzsystems, indem es eine stabile, effiziente Methode für Transaktionen in der digitalen Währungswelt ermöglicht.

Beschreibung

• Own Cosmic AC platform architecture end to end, including architecture proposals and high- and low-level designs • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build and operate a managed Slurm service for research users • Own controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Implement and operate NVIDIA GPU Operator, Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup, recovery and node replacement • Own managed inference architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability and confidential-compute-capable capacity • Establish metrics, logging, alerting and SLOs across the control plane, GPU fleet and application tiers • Lead incident response, post-incident reviews and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Convert requirements into written specifications and acceptance tests, run escalations to closure, and contribute to capacity planning and hardware sourcing • Translate research, model-training and product workloads into platform requirements and broker capacity when limited • Complete the platform team and set the technical bar for new engineers

🎯 Anforderungen

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running • Experience operating HPC or GPU training clusters for research users is ideal • Experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in and acceptance • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues • Deep Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operations experience, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident reviews • Working fluency in JavaScript and Node.js sufficient to review and make architecture decisions on control-plane, CLI and worker services • Experience shipping a platform with real users, such as multi-tenant IaaS/PaaS or research computing service • People management across time zones, cross-track review, written architecture decisions, and ability to challenge partners or executives with reasons • Excellent written and spoken English • Ability to work from UTC through UTC+5:30 to overlap with Europe and India • Occasional travel to partner sites and team events

🏖️ Vorteile

• Fully remote work arrangement • Occasional travel to partner sites and team events • Global, distributed team environment • Opportunity to work on innovative digital finance, GPU infrastructure, AI and blockchain platforms

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 7 Tagen

accesa.eu

1001 - 5000

💼 Beratung

🏭 Fertigung

🏥 Gesundheitswesen

Microsoft 365 Software Engineer building SharePoint and Microsoft 365 solutions for Germany’s financial-banking sector at Accesa. Integrating React, Azure, Graph API, and AI tools to improve digital workplaces.

🇷🇴 Rumänien – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 7 Tagen

Software Mind

1001 - 5000

🤖 Künstliche Intelligenz

☁️ SaaS

📡 Telekommunikation

Senior Full-Stack Engineer developing and enhancing .NET and React web solutions for global clients. Collaborating on architecture, delivery, documentation, and business-user requirements.

🇷🇴 Rumänien – Remote

💰 Private Equity Round im 2020-12

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 7 Tagen

Brainlab

1001 - 5000

💼 Beratung

📦 Logistik

🏭 Fertigung

Softwareentwickler für die Entwicklung kranialer Navigationsanwendungen in C++ und React für die Medizintechnik von Brainlab. Entwicklung bildgestützter Operationssoftware in multidisziplinären Forschungs- und Entwicklungsteams.

🇷🇴 Rumänien – Remote

💰 Private Equity Round im 2018-09

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 8 Tagen

accesa.eu

1001 - 5000

💼 Beratung

🏭 Fertigung

🏥 Gesundheitswesen

Microsoft 365 engineer building SharePoint, React, and Azure solutions for Accesa’s banking-sector clients. Integrating Microsoft 365 services and shaping secure digital workplaces.

🇷🇴 Rumänien – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 8 Tagen

Kraken Digital Asset Exchange

1001 - 5000

₿ Crypto

💸 Finanzen

💳 Fintech

Senior Rust Engineer building secure, scalable core services for Kraken’s global crypto-financial platform. Designing distributed systems and foundational libraries for trading infrastructure.

🇷🇴 Rumänien – Remote

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich