Senior Kubernetes Operations Engineer

March 15

Apply Now

Loading...

Lambda

Designing the world's most advanced GPU systems for Deep Learning.

Deep Learning • Machine Learning • Artificial Intelligence

51 - 200

💰 $39.7M Venture Round on 2022-11

Description

• Remotely install, upgrade, operate and maintain bare-metal Kubernetes clusters (up to thousands of nodes each) • Handle cluster degradation, recovery and resizing using our fleet management tooling • Perform out-of-hours on-call response for critical incidents as part of a well-balanced on-call rotation • Work on improving our tooling, automation, and processes, for both daily operations, alerting, and incident response • Dive into systems at a low level to solve unique cluster problems and write up your findings • Assist customers with high-level Kubernetes questions and integration with applications, storage and authentication • Assist with initial cluster build-outs and validation to help identify failed hardware before customer delivery • Work closely with our HPC Ops and Datacenter Ops teams on issues that require lower-level expertise or cross-functional solutions • Mentor and assist less-experienced team members • Have a voice in our product direction and help us think about how to minimize operational costs and complexity

Requirements

• An experienced operations engineer, SRE, sysadmin or similar with a deep knowledge of running Linux clusters and systems • Very familiar with running on bare-metal (including knowledge of BMCs, kernel drivers, PXE, RAID, VLANs, hypervisors) • A good understanding of containers, virtualisation, and the mechanisms underpinning them • A good understanding of daily operation, bug-fixing and maintenance of Kubernetes • Experience in an on-call environment and with incident response • Ability to perform incident post-mortems and develop procedures and tooling to prevent root causes from reoccurring • An excellent ability to learn on-the-fly and adapt to solve problems • Able to work either independently with limited direction, or as part of a team • Able to work with customers during incidents either via tickets, live messaging, or as part of a larger call.

Benefits

• Health, dental, and vision coverage for you and your dependents • Commuter/Work from home stipends • 401k Plan with 2% company match • Flexible Paid Time Off Plan that we all actually use

Apply Now
Built by Lior Neu-ner. I'd love to hear your feedback — Get in touch via DM or lior@remoterocketship.com
Jobs by Title
Remote Account Executive jobsRemote Accounting, Payroll & Financial Planning jobsRemote Administration jobsRemote Android Engineer jobsRemote Backend Engineer jobsRemote Business Operations & Strategy jobsRemote Chief of Staff jobsRemote Compliance jobsRemote Content Marketing jobsRemote Content Writer jobsRemote Copywriter jobsRemote Customer Success jobsRemote Customer Support jobsRemote Data Analyst jobsRemote Data Engineer jobsRemote Data Scientist jobsRemote DevOps jobsRemote Engineering Manager jobsRemote Executive Assistant jobsRemote Full-stack Engineer jobsRemote Frontend Engineer jobsRemote Game Engineer jobsRemote Graphics Designer jobsRemote Growth Marketing jobsRemote Hardware Engineer jobsRemote Human Resources jobsRemote iOS Engineer jobsRemote Infrastructure Engineer jobsRemote IT Support jobsRemote Legal jobsRemote Machine Learning Engineer jobsRemote Marketing jobsRemote Operations jobsRemote Performance Marketing jobsRemote Product Analyst jobsRemote Product Designer jobsRemote Product Manager jobsRemote Project & Program Management jobsRemote Product Marketing jobsRemote QA Engineer jobsRemote SDET jobsRemote Recruitment jobsRemote Risk jobsRemote Sales jobsRemote Scrum Master + Agile Coach jobsRemote Security Engineer jobsRemote SEO Marketing jobsRemote Social Media & Community jobsRemote Software Engineer jobsRemote Solutions Engineer jobsRemote Support Engineer jobsRemote Technical Writer jobsRemote Technical Product Manager jobsRemote User Researcher jobs