Principal Software Engineer – Rack Scale Systems Infrastructure

Vaga não está no LinkedIn

🕒 Maio 16

🏄 California, North Carolina, +2 estados a mais – Remoto

infoinfo

💵 $272.000 - $431.250 / ano

⏰ Tempo Integral

🔴 Especialista

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 36%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of NVIDIA

NVIDIA

10.000+ funcionários

Fundada em 1993

🏥 Saúde

🏭 Manufatura

🤖 Inteligência Artificial

Healthcare • Manufacturing • Artificial Intelligence

A NVIDIA é uma empresa de tecnologia líder, especializada em computação acelerada e inteligência artificial. A companhia é pioneira em avanços em unidades de processamento gráfico (GPUs), computação em nuvem, data centers e realidade virtual, com foco nos setores de games, automotivo, saúde e robótica. As inovações da empresa, como o NVIDIA Omniverse, transformam processos digitais tradicionais ao viabilizar simulações de alta fidelidade e tarefas de renderização. Suas aplicações abrangem diversos setores, desde veículos autônomos com o NVIDIA DRIVE até soluções de saúde com o NVIDIA Clara, além de análises e fluxos de trabalho impulsionados por IA.

Descrição

• Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software. • Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. • Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments. • Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. • Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization. • Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. • Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments. • Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development. • Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.

🎯 Requisitos

• BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. • Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering. • Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs. • Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. • Experience with Rust is highly valued. • Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. • Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows. • Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. • Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems. • Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. • Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation. • Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations. • Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation. • Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. • Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.

🏖️ Benefícios

• equity • benefits

Candidatar-se

Vagas Similares

🕒 Maio 16

Fieldguide

11 - 50

💼 Consultoria

🏥 Saúde

🤖 Inteligência Artificial

Staff Platform Engineer designing and building foundational platform services for Fieldguide, a fintech company automating assurance and audit work. Leading technical architecture and mentoring engineers across teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $210.000 - $265.000 / ano

⏰ Tempo Integral

🔴 Especialista

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 16

Weave

1 - 10

🧬 Biotecnologia

🤖 Inteligência Artificial

🏥 Saúde

Principal Engineer leading the Machine Learning Team at Weave, enabling AI feature development and infrastructure design. Focus on building scalable ML systems for healthcare communication.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 15

Confluent

1001 - 5000

🤖 Inteligência Artificial

☁️ SaaS

Strategic technical leader defining and driving AI capabilities for Confluent’s productivity. Collaborating across teams to integrate AI and smart automation solutions into the development lifecycle.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $285.000 - $342.000 / ano

💰 Secondary Market em 2021-06

⏰ Tempo Integral

🔴 Especialista

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 15

ParentSquare

201 - 500

📚 Educação

☁️ SaaS

🤝 B2B

Staff Software Engineer focusing on AI in communication tools for schools and parents. Lead technical initiatives, collaborate with cross-functional teams, and drive architectural evolution.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $180.000 - $230.000 / ano

⏰ Tempo Integral

🔴 Especialista

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 14

Celestica

10.000+ funcionários

💼 Consultoria

🎖️ Defesa

🏥 Saúde

Principal Engineer / Networking System Architect at Celestica, defining architecture for network products and leading design teams. Engaging with suppliers and customers on technical aspects.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $171.000 - $268.000 / ano

💰 $660.400.000 Post-IPO Debt em 2021-09

⏰ Tempo Integral

🔴 Especialista

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Assembly