Service Reliability Engineer

🔥 9 horas atrás

🤠 Texas – Remoto

info

💵 $168.000 - $333.500 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of NVIDIA

NVIDIA

10.000+ funcionários

Fundada em 1993

🏥 Saúde

🏭 Manufatura

🤖 Inteligência Artificial

Healthcare • Manufacturing • Artificial Intelligence

A NVIDIA é uma empresa de tecnologia líder, especializada em computação acelerada e inteligência artificial. A companhia é pioneira em avanços em unidades de processamento gráfico (GPUs), computação em nuvem, data centers e realidade virtual, com foco nos setores de games, automotivo, saúde e robótica. As inovações da empresa, como o NVIDIA Omniverse, transformam processos digitais tradicionais ao viabilizar simulações de alta fidelidade e tarefas de renderização. Suas aplicações abrangem diversos setores, desde veículos autônomos com o NVIDIA DRIVE até soluções de saúde com o NVIDIA Clara, além de análises e fluxos de trabalho impulsionados por IA.

Descrição

• Operate within a 24/7 follow-the-sun support model across multiple continents • Manage a 4-day, 10-hour schedule, including either Saturday or Sunday, with flexible early or late shifts • Monitor and manage extensive production GPU and Kubernetes environments • Detect, prevent, and respond to incidents proactively • Analyze logs, metrics, and system behavior to diagnose issues and implement resolutions • Develop predictive automated support routines • Improve automation through auto-healing and automated break-fix solutions • Perform systems administration, network administration, and security monitoring • Coordinate with domain experts and service owners to resolve complex issues • Continuously improve service quality and operational processes based on incident feedback • Coordinate effectively across teams during incident resolution • Deliver customer-focused support throughout client interactions

🎯 Requisitos

• Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management • Familiarity with GPU hardware and high-performance computing environments • Proficiency with Grafana, OpenTelemetry, PagerDuty, and JIRA • Experience with AWS, Azure, GCP, or OCI is a plus; strong preference for on-prem expertise • Ability to work effectively with multifunctional teams • 8+ years of experience coordinating large-scale production systems • More than 3 years of experience in high-availability Internet, Cloud, or Data Center settings • BS in Computer Science, Engineering, Physics, Mathematics, or equivalent experience • Expert-level Linux system administration • Automation using Ansible and/or Python • Strong expertise in shell scripting, DNS, DHCP, storage systems, and core networking • Proven experience maintaining large-scale bare-metal infrastructure • Excellent partnership, documentation, and mentoring skills • Experience with scripting languages, particularly Python • Experience running virtual machines under community-supported or commercial hypervisors • Knowledge of application containers and container orchestration systems • Basic understanding of Git • Ability to master and maintain complicated environments

🏖️ Benefícios

• Equity • Benefits

Candidatar-se

Vagas Similares

🔥 9 horas atrás

Entarian

1001 - 5000

🚀 Aeroespacial

🎖️ Defesa

🏛️ Governo

DevSecOps Engineer automating secure cloud infrastructure and CI/CD platforms for Entarian’s U.S. Navy missions. Supporting Kubernetes, Terraform, cloud networking, compliance validation, and infrastructure testing.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🔥 11 horas atrás

Endava

10.000+ funcionários

🏥 Saúde

📣 Marketing

📦 Logística

Senior DevOps Engineer designing Dynatrace observability across Endava’s cloud-native enterprise platforms. Improving monitoring, incident response, reliability, and resilience for web, mobile, API, and microservices environments.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Post-IPO Debt em 2023-02

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🔥 11 horas atrás

The Mind Company

51 - 200

🏥 Saúde

📣 Marketing

💼 Consultoria

Senior DevOps Engineer shaping AI tooling, infrastructure, CI/CD, and mobile releases for a company delivering mental fitness apps. Driving platform reliability and developer productivity across fully remote engineering teams in the Americas.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🔥 14 horas atrás

CyberSheath

51 - 200

💼 Consultoria

🎖️ Defesa

🔒 Cibersegurança

Cloud Operations Engineer delivering managed cybersecurity infrastructure services for Defense Industrial Base clients. Supporting Azure, AWS, Linux, Office 365, migrations, and secure technology implementations.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $110.000 - $127.000 / ano

💰 Private Equity Round em 2021-12

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🔥 14 horas atrás

JFrog

1001 - 5000

💼 Consultoria

🏥 Saúde

📦 Logística

Professional Services DevOps Engineer building CI/CD platforms for JFrog, which manages and secures software delivery from code to production. Designing pipelines, cloud infrastructure and DevOps solutions for enterprise customers.

🗣️🇺🇸🇬🇧 Inglês obrigatório