Director, Site Reliability Engineering

🔥 12 hours ago

🇺🇸 United States – Remote

💵 $243.8k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of DuckDuckGo

DuckDuckGo

51 - 200 employees

Founded 2008

💼 Consulting

📣 Marketing

🔒 Cybersecurity

Consulting • Marketing • Cybersecurity

DuckDuckGo is an independent Internet privacy company dedicated to providing users with tools that enhance their online privacy. Best known for its search engine, DuckDuckGo offers a free web browser and browser extensions for iOS, Android, Mac, and Windows that include built-in privacy features such as tracker blocking, encryption, and email protection. Unlike traditional search engines and browsers that track search and browsing history, DuckDuckGo prioritizes user privacy by preventing tracking and targeted advertising. The company operates on the premise that privacy should be simple and accessible, offering a private alternative to major tech companies while remaining financially sustainable through non-invasive ads based on search results rather than personal data.

📋 Description

• Build and maintain world-class infrastructure for millions of privacy-focused users • Work with high-level languages including Perl, Go, TypeScript, and Python • Ensure Duck.ai meets reliability standards and minimize user friction during failures • Scale DuckDuckGo’s index infrastructure to handle billions of documents • Create privacy-respecting anti-fraud verifications • Address complex operational challenges involving software, systems, automation, and process analysis • Read, write, troubleshoot, and deploy software for large-scale reliability challenges • Lead and collaborate on complex projects from proposal through postmortem • Identify and address reliability risks using tools, services, alerts, and response processes • Investigate and root-cause instability in high-traffic distributed systems • Help define the future technical direction of deployment to improve reliability and performance • Partner with software engineers to triage production issues and determine remediation, including code changes and performance considerations • Use cloud-native services and architectures to improve reliability and scalability

🎯 Requirements

• 10+ years relevant professional experience in reliability, platform, infrastructure, or software engineering • 4+ years leading SRE teams • Experience participating in a 24x7 on-call rotation for a large-scale deployment • Ability to lead and collaborate on high-impact, complex projects from proposal through postmortem • Proficiency in AI-driven development, including designing and implementing agentic workflows • Ability to solve vague problems, propose innovative solutions, and execute with a strong focus on metrics • Experience developing tools, services, alerts, and responses to identify and address reliability risks • Ability to root-cause instability in high-traffic, distributed systems • Deep experience administering and troubleshooting Linux and web technologies • Ability to automate infrastructure provisioning and configuration management • Advanced programming skills • Experience with cloud-native services and architectures • Hands-on experience packaging and deploying applications using Docker and Docker Compose • Must attend meetings on camera via video conferencing • Must pass a background check as a condition of joining • Must be legally authorized to work in the country of residence; DuckDuckGo does not sponsor or assist with individual immigration needs

🏖️ Benefits

• Stock options • Company-sponsored health benefits for team members based in the United States • Paid parental leave • Office setup allowance • Co-working allowance • Flexible work arrangement with no core hours • Remote-first work environment • Company all-hands meetup and team retreat travel opportunities

Apply Now

Similar Jobs

🔥 17 hours ago

ExpertVoice

51 - 200

📣 Marketing

🛒 Retail

🛍️ eCommerce

Site Reliability Engineer IV setting technical direction for ExpertVoice’s platform serving leading consumer brands. Owning reliability, automation, CI/CD, architecture, and incident response at scale.

🕒 Yesterday

Cadwell

51 - 200

🏥 Healthcare

🏭 Manufacturing

🔧 Hardware

Cloud Site Reliability Engineer operating AWS infrastructure for Cadwell’s neurodiagnostic healthcare software. Automating deployments, observability, security, disaster recovery, and incident response for clinical customers.

🕒 Yesterday

Cast & Crew

501 - 1000

☁️ SaaS

📱 Media

👥 HR Tech

Staff DevOps Engineer owning AWS EKS, Azure DevOps CI/CD, and platform reliability. Supporting Cast & Crew’s entertainment technology and services business.

🇺🇸 United States – Remote

💵 $190k - $235k / year

💰 Private equity on 2013-03

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 Yesterday

Cast & Crew

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Staff DevOps Engineer owning AWS EKS, Terraform, and Azure DevOps platforms for Cast & Crew’s entertainment technology business. Improving reliability, developer experience, and infrastructure standards.

🕒 3 days ago

Clinician Nexus

51 - 200

🏥 Healthcare

⚕️ Healthcare Insurance

📚 Education

DevOps Manager leading secure, reliable platform engineering for Clinician Nexus, a healthcare workforce technology company. Balancing team leadership with hands-on AWS, Kubernetes, Terraform, and CI/CD engineering.