Site Reliability Engineer – US Central/Eastern Time

Job not on LinkedIn

🕒 4 days ago

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of PostHog

PostHog

11 - 50 employees

Founded 2020

💼 Consulting

📣 Marketing

📦 Logistics

Consulting • Marketing • Logistics

PostHog is a comprehensive platform that empowers developers to build successful products by providing tools for product analytics, web analytics, session replay, feature flags, experiments, and surveys. It integrates seamlessly into existing workflows, offering data pipelines and warehousing solutions that synchronize with popular platforms like Stripe, Hubspot, Zendesk, and more. With PostHog, teams can safely roll out new features, run experiments with statistical significance, and gather in-depth insights with AI and LLM products. The platform is built with full API access, enabling complete control over customer data. PostHog scales with businesses from startups to growth stages, making it a versatile tool for engineering teams seeking to streamline their data operations while focusing on product development.

📋 Description

• Turn a fast-growing, stateful system into a predictable, well-automated platform through provisioning, scaling, rebalancing, and recovery • Operate EKS clusters across several environments using Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments • Manage and evolve a multi-AWS-account organization, including provisioning, networking, access control, and cross-account connectivity • Maintain the Terraform/Terragrunt infrastructure-as-code platform, including modules, plan-on-PR/apply-on-merge pipelines, and safe shared-infrastructure patterns • Improve operational tooling for deployments, schema changes, backups, restores, and incident response • Identify recurring operational pain points and eliminate them through code and self-healing automation • Optimize cloud spend • Participate in on-call and incident response while reducing incident frequency over time • Design and automate the platform layer supporting the organization’s services

🎯 Requirements

• Deep hands-on experience with Kubernetes in production; EKS preferred • Experience debugging node pressure, networking issues, and deployment failures at scale, including thousands of nodes • Strong experience operating production infrastructure on AWS across multiple accounts • Understanding of AWS organizational boundaries, IAM, and inter-account networking • Experience automating infrastructure with Terraform or Terragrunt at scale, including module design and state management • Solid understanding of Linux systems, including disk, memory, networking, and failure modes • Experience supporting stateful systems such as databases, queues, and storage systems • Ability to debug and reason about production performance and reliability issues • Comfortable owning systems end-to-end, including on-call responsibilities • US Central or Eastern timezone

🏖️ Benefits

• Remote work arrangement • Meeting-free days on Tuesdays and Thursdays • Heads-down building time prioritized over perfect coordination • Async communication by default • Fair and accessible interview process with accommodations or adjustments available

Apply Now

Similar Jobs

🕒 4 days ago

ArangoDB

51 - 200

🏥 Healthcare

📦 Logistics

💼 Consulting

Site Reliability Engineer scaling Arango’s cloud-native contextual AI data platform. Automating Kubernetes, AWS, and Google Cloud infrastructure with Golang, observability, and resilient production operations.

🇺🇸 United States – Remote

💰 $27.8M Series B on 2021-10

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 4 days ago

CampMinder

51 - 200

💼 Consulting

🏥 Healthcare

🏨 Hospitality

Senior DevSecOps Engineer securing Campminder’s software for summer camps. Hardening cloud infrastructure, embedding security in CI/CD, and leading compliance and threat-remediation work.

🇺🇸 United States – Remote

💵 $180k - $200k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 4 days ago

MeridianLink

501 - 1000

💳 Fintech

🏦 Banking

☁️ SaaS

Senior SRE owning reliability, observability, and scalability for MeridianLink’s financial SaaS applications. Designing resilient AWS/Azure infrastructure, automation, incident response, and security practices.

🇺🇸 United States – Remote

💵 $104.1k - $177.6k / year

💰 $485M Post-IPO Debt on 2021-11

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 4 days ago

eFinancial

201 - 500

🏥 Healthcare

💼 Consulting

🛡️ Insurance

DevOps Team Lead building AWS infrastructure, CI/CD pipelines, and shared engineering platforms for a life insurance provider. Coaching engineers and improving deployment reliability and automation.

🕒 5 days ago

NEC Software Solutions

5001 - 10000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior DevOps Engineer managing AWS cloud infrastructure, Kubernetes, Terraform, and CI/CD for NEC SWS public-sector systems. Hybrid role requiring 50% office attendance and SC eligibility.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)