Senior Site Reliability Engineer

🔥 14 hours ago

🇺🇸 United States – Remote

💵 $180k - $220k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Zocdoc

Zocdoc

501 - 1000 employees

Founded 2007

🏥 Healthcare

⚕️ Healthcare Insurance

🏪 Marketplace

💰 $150M Private Equity Round on 2021-02

Healthcare • Healthcare Insurance • Marketplace

Zocdoc is a digital health platform that connects patients with healthcare providers, allowing users to book appointments online with doctors and dentists who accept their insurance. The service offers a wide range of specialties, including primary care, dentistry, ob-gyn, dermatology, and psychiatry, among others. Zocdoc aims to make healthcare access easier by providing an app for appointment scheduling and management, enhancing the visibility and reputation of healthcare providers through verified reviews, and serving as a partner for health systems and private practices. The company does not provide medical advice but facilitates the connections and processes necessary for patients to access healthcare services efficiently.

📋 Description

• Develop, monitor, and maintain distributed production systems • Build frameworks and processes for ensuring uptime for patients and providers • Monitor and maintain complex cloud-based infrastructure, systems, and services • Automate and develop tooling, processes, and infrastructure to make development faster, repeatable, and error-proof • Support product engineering teams with scaling, performance, and uptime needs • Diagnose and debug production-related issues • Analyze and performance-tune systems, code, and networking for scaling and optimal operation • Work with cutting-edge GenAI tools and technology • Enforce a culture of strong DevOps and shared product-team responsibility for site reliability and first response

🎯 Requirements

• 5+ years of supporting consumer facing web application production environments and systems in a Site Reliability Engineering or Production Engineering role • 2+ years of on-call experience in a 24/7 cloud-based production environment • 2+ years of experience in managing and supporting modern cloud-based environments and infrastructure like AWS/GCP, Docker, Kubernetes, etc. • Experience with edge technologies such as load balancers, reverse proxies, web application firewalls, routing, etc. • Deep understanding of protocols such as TCP/IP, HTTP/HTTPS, TLS, DNS, NTP • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent engineering experience is a plus, but not required • Comfortable in an outage situation and believe in blameless post-mortems • Highly collaborative, autonomous, individually accountable, and committed to diverse and inclusive teams

🏖️ Benefits

• Flexible, hybrid work environment at our convenient Soho location (If based in NYC) • Unlimited Vacation • 100% paid employee health benefit options (including medical, dental, and vision) • Commuter Benefits • 401(k) with employer funded match • Corporate wellness program with Wellhub • Sabbatical leave (for employees with 5+ years of service) • Competitive paid parental leave and fertility/family planning reimbursement • Cell phone reimbursement • Catered lunch everyday along with beverages and snacks • Employee Resource Groups and ZocClubs to promote shared community and belonging • Great Place to Work Certified

Apply Now

Similar Jobs

🔥 16 hours ago

Tradeify

51 - 200

💳 Fintech

💸 Finance

DevSecOps Engineer owning AWS infrastructure, CI/CD, and security operations for Tradeify’s high-performance futures and crypto trading platform. Improving reliability, observability, vulnerability management, and incident response.

🔥 17 hours ago

Solventum

10,000+ employees

🏥 Healthcare

📦 Logistics

💼 Consulting

Site Reliability Engineer supporting Solventum’s healthcare speech products. Maintaining production systems, monitoring, alerting, and cloud infrastructure with AWS and Kubernetes.

🔥 17 hours ago

Megaport

201 - 500

📡 Telecommunications

Senior Site Reliability Engineer improving reliability, resilience, and observability for Megaport’s global Network as a Service infrastructure. Automating production systems across cloud and Kubernetes environments.

🔥 18 hours ago

Capstone Integrated Solutions

51 - 200

💼 Consulting

🛒 Retail

AWS DevOps Engineer building AWS cloud, automation, and MLOps infrastructure for CapNexus, a software development and systems integration services provider. Supporting SageMaker, Bedrock, CI/CD, security, and Azure-to-AWS migration.

🔥 18 hours ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Senior SRE maintaining NVIDIA’s managed DGX Cloud AI clusters across major cloud providers. Improving Kubernetes reliability, observability, GPU workloads, and production incident response.