Principal Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $167.3k - $242.6k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Expel

Expel

201 - 500 employees

Founded 2016

🔒 Cybersecurity

☁️ SaaS

Cybersecurity • SaaS • Technology

Expel is a leading cybersecurity company specializing in Managed Detection and Response (MDR) services. They offer a range of solutions, including phishing investigation, threat hunting, and vulnerability prioritization, tailored for organizations of all sizes with 24x7 protection. Expel's Security Operations Platform, Expel Workbench™, integrates with existing tech to enhance security operations. Their expert team and advanced technology help reduce alert noise, respond swiftly to incidents, and improve overall security posture, enabling organizations to focus on core business activities without worrying about cybersecurity threats.

📋 Description

• Lead project work building and maintaining platform features across product reliability, networking, and cloud infrastructure • Push infrastructure-as-code commits daily • Occasionally write and test application code in Python, Golang, and JavaScript • Mentor and motivate service owners on deploying, measuring, monitoring, and operating services at scale • Participate in a weekly support rotation, including on-call pager duties and working-hours support for platform users • Lead incident response, triage, and root cause analysis support • Collaborate with architects and product stakeholders on reliability initiatives • Pair-program with and mentor junior SREs • Work from a shared backlog, pair-program weekly, peer-review work, and participate in blame-free retrospectives

🎯 Requirements

• Significant experience operating Kubernetes in highly distributed environments • Experience running systems in GCP or AWS • Exposure to monitoring and observability infrastructure and standard methodologies • Understanding of infrastructure-as-code practices, tools, and patterns • Some software development experience in Linux environments, preferably Python and/or Golang • Six years of systems experience in operations or development • Customer-minded approach supporting platform users and building organizational trust • Collaborative disposition for working across teams

🏖️ Benefits

• Bonus eligibility • Equity • Unlimited PTO • Work location flexibility • Up to 24 weeks of parental leave • Excellent health benefits • Reasonable accommodation for disabilities during hiring, work, and access to benefits

Apply Now

Similar Jobs

🔥 12 hours ago

Accellor

201 - 500

💼 Consulting

🏥 Healthcare

🏨 Hospitality

Principal FDE designing and deploying governed AI architectures for Accellor’s enterprise clients. Leading customer engagements, production implementations, and the forward-deployed engineering team.

🔥 17 hours ago

Skydio

501 - 1000

🎖️ Defense

🏭 Manufacturing

📦 Logistics

Staff Site Reliability Engineer operating Kubernetes, AWS, and Terraform infrastructure for Skydio’s autonomous drone platform. Ensuring reliable, scalable cloud services through observability, automation, and incident response.

🔥 19 hours ago

Renesas Electronics

10,000+ employees

🏭 Manufacturing

🏥 Healthcare

📦 Logistics

Quality and reliability engineer qualifying AI server power modules at Renesas, a global semiconductor solutions company. Developing reliability tests, component qualification processes, validation plans, and manufacturing prototypes for high-performance computing products.

🔥 23 hours ago

General Dynamics Information Technology

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

GDIT DevSecOps Engineer maintaining secure AWS, Kubernetes, and GitLab pipelines. Supporting defense and intelligence technology, video encoding, automated security scanning, and system performance validation.

🕒 Yesterday

CareSource

1001 - 5000

🏥 Healthcare

🛡️ Insurance

⚕️ Healthcare Insurance

Cloud, DevOps, and SRE architect designing infrastructure, deployment automation, and reliability strategies for CareSource’s digital healthcare platform. Scaling enterprise technology for national delivery.