Senior Engineer

🕒 Agosto 24

🏄 California, Washington – Remoto

infoinfo

💵 $184.000 - $356.500 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 1%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of NVIDIA

NVIDIA

10.000+ funcionários

Fundada em 1993

🏥 Saúde

🏭 Manufatura

🤖 Inteligência Artificial

Healthcare • Manufacturing • Artificial Intelligence

A NVIDIA é uma empresa de tecnologia líder, especializada em computação acelerada e inteligência artificial. A companhia é pioneira em avanços em unidades de processamento gráfico (GPUs), computação em nuvem, data centers e realidade virtual, com foco nos setores de games, automotivo, saúde e robótica. As inovações da empresa, como o NVIDIA Omniverse, transformam processos digitais tradicionais ao viabilizar simulações de alta fidelidade e tarefas de renderização. Suas aplicações abrangem diversos setores, desde veículos autônomos com o NVIDIA DRIVE até soluções de saúde com o NVIDIA Clara, além de análises e fluxos de trabalho impulsionados por IA.

Descrição

• Lead NVIDIA Cloud Partner Day 2 operational readiness efforts following initial deployment and activation • Collaborate with NVIDIA Cloud Partners to establish systems, procedures, automation, and operational methods for accelerated infrastructure • Build continuous validation for GPU, CPU, storage, and network health across large-scale AI clusters • Establish telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, networking, storage, Kubernetes, and AI workloads • Develop automated workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service • Manage GPU fleet lifecycle, including drivers, firmware, Kubernetes nodes, OS patching, configuration management, upgrades, and configuration drift • Translate NVIDIA NCP requirements and reference architectures into production operating practices, validation criteria, runbooks, automation, and measurable standards • Define health signals, SLOs, metrics, acceptance criteria, and infrastructure readiness validation • Build reusable tooling, automation, implementation guides, runbooks, playbooks, and reference implementations across NCP environments

🎯 Requisitos

• BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience • 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments • Strong experience operating Linux-based distributed systems and cloud infrastructure in production • Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments • Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service level agreements • Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management • Strong networking fundamentals and experience troubleshooting complex distributed systems across compute, network, and storage layers • Programming and automation experience using Python, Go, shell scripting, or similar languages • Experience managing extensive GPU or accelerated computing infrastructure that supports AI training and inference workloads • Experience with NVIDIA technologies including DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software • Proven experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure and operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency • Extensive knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines • Knowledge of failure modes related to large distributed AI workloads and the infrastructure features necessary to consistently support extended training and production inference

🏖️ Benefícios

• Equity • Benefits

Candidatar-se

Vagas Similares

🕒 Agosto 22

ShippyPro

51 - 200

📦 Logística

☁️ SaaS

🛍️ Comércio Eletrônico

Senior backend engineer building scalable PHP Laravel services for ShippyPro’s shipping and fulfillment platform. Architecting microservices, distributed workflows, and AI automation for global merchants.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 €42.000 - €56.000 / ano

💰 $15.000.000 Series B - ShippyPro em 2023-11

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 22

Cisco

10.000+ funcionários

🔧 Hardware

🔐 Segurança

🏢 Corporativo

Technical Leader shaping Cisco’s C++ testing ecosystem and developer infrastructure. Building frameworks, CI/CD quality gates, and engineering metrics platforms.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 22

Cisco

10.000+ funcionários

🔧 Hardware

🔐 Segurança

🏢 Corporativo

Technical Leader advancing Cisco’s C++/Python testing ecosystem and developer infrastructure. Designing frameworks, CI/CD quality gates, and engineering metrics platforms for teams across Cisco.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 22

Gartner

10.000+ funcionários

📣 Marketing

📦 Logística

🏥 Saúde

Senior Gartner analyst shaping AI strategy, SDLC adoption, and ROI measurement for software engineering leaders. Delivering research, client guidance, presentations, and sales support.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 22

RR Digital

51 - 200

📣 Marketing

⚖️ Jurídico

🤝 B2B

Full Stack Developer building React, TypeScript, Python, and AWS document automation platforms for RR Digital, a legal marketing agency. Developing scalable production software for law-firm marketing operations.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório