Site Reliability Engineer L4/L5 – Live Cloud Platform SRE

🕒 June 27

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Netflix

Netflix

10,000+ employees

Founded 1997

📱 Media

👥 B2C

Media • B2C

Netflix is a global streaming entertainment company and content producer whose stated mission is "to entertain the world. " It operates a consumer-facing subscription platform offering on-demand TV shows, films, and original programming, and also runs a public careers site emphasizing culture, inclusion, and hiring accommodations. The provided text highlights Netflix’s focus on recruiting talent worldwide, its work-life and culture pages, and its public-facing employer materials.

📋 Description

• Drive continual improvement in observability, monitoring, and scalability with the primary goal to solve the thundering herd problem with cloud traffic (API gateway, IPC between microservices) for live streaming. • Implement, automate, execute, and analyze the results from a broad range of live streaming delivery focused functional, performance, resilience, and fault injection testing. • Write and review code, develop documentation and capacity plans, and debug the hardest problems on some of the largest and most complex systems in the world • Coordination, collaboration, and partnership across multiple stakeholders for the smooth execution of live-streaming events • Participate in an on-call rotation and be able to work with flexible hours based on the live events schedule

🎯 Requirements

• 5+ years of service reliability/operational experience running large-scale, high-performance systems & internet services with a focus on traffic at scale • Knowledge of and proven experience with L4 Load Balancer, HTTP cache, and reverse proxy technologies • Expert-level knowledge of Unix or Linux systems and TCP/IP network fundamentals • Proficient understanding of networking principles, transport, and application protocols, especially DNS, TLS, and HTTP(s) etc. • Proficient in a programming language such as Go, Python, Rust etc. • Experience with using real-time and Big Data analytic processing technologies (Kafka, time series database and Presto/Trino, Spark SQL, etc) • Ability to work in a highly collaborative environment and to communicate effectively with internal and external partners • Preferred - B.S. in Computer Science, Electrical or Computer Engineering (or equivalent professional experience)

🏖️ Benefits

• Health Plans • Mental Health support • 401(k) Retirement Plan with employer match • Stock Option Program • Disability Programs • Health Savings and Flexible Spending Accounts • Family-forming benefits • Life and Serious Injury Benefits • Paid leave of absence programs • Full-time hourly employees accrue 35 days annually for paid time off • Flexible time off for full-time salaried employees

Apply Now

Similar Jobs

🕒 June 27

Berkeley Research Group (BRG)

1001 - 5000

🏗️ Construction

📦 Logistics

🎖️ Defense

Site Reliability Engineer designing, building, and maintaining highly available systems for health technology company. Collaborating with software developers to improve reliability and automate processes.

🇺🇸 United States – Remote

💵 $130k - $160k / year

💰 Venture Round on 2020-07

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 June 26

CACI International Inc

10,000+ employees

💼 Consulting

🎖️ Defense

Cloud DevOps Engineer managing CI/CD pipelines and applications in AWS cloud. Collaborating on security initiatives and providing DevSecOps training with Agile teams.

🕒 June 26

Alkami Technology

501 - 1000

💼 Consulting

📣 Marketing

🏦 Banking

Site Reliability Engineer at Alkami developing and testing code for application releases. Collaborating with teams to improve delivery and participate in on-call rotations.

🇺🇸 United States – Remote

💵 $110k - $137.5k / year

💰 $300M Post-IPO Debt - Alkami Technology on 2025-03

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 June 26

Hearst Health

1001 - 5000

🏥 Healthcare

⚕️ Healthcare Insurance

☁️ SaaS

Senior DevOps Engineer helping build and improve healthcare technology platforms at Zynx Health. Operating cloud infrastructure, deployment pipelines, and security practices for clinical decision support solutions.

🕒 June 26

THEMIS Waste Recovery Technology

11 - 50

🏥 Healthcare

📦 Logistics

🍽️ Food & Beverage

DevOps Engineer managing cloud infrastructure and CI/CD pipelines for secure, reliable operations at Themis. Ensuring system health and performance while automating processes and maintaining security controls.