Senior Site Reliability Engineer

Job not on LinkedIn

🕒 July 16

🇺🇸 United States – Remote

💵 $160k - $200k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Epic Kids

Epic Kids

11 - 50 employees

📚 Education

👥 B2C

📱 Media

Education • B2C • Media

Epic Kids is a digital subscription service offering a curated library of children's eBooks, audiobooks, and learning videos. The platform provides families and educators access to thousands of age-appropriate titles from major publishers to encourage reading, discovery, and digital literacy. Epic Kids offers plans for parents and free access for educators, positioning itself as an educational media resource for children.

📋 Description

• Drive the reliability of Epic's infrastructure—set and track SLOs/SLIs, reduce toil, and engineer out recurring instability. • Build and operate the cloud infrastructure and container platform for high availability, scalability, and cost efficiency—including workload scheduling, autoscaling, networking, and graceful failure handling. • Maintain and improve CI/CD pipelines for fast, safe delivery across engineering teams. • Own and evolve the observability stack—metrics, logs, traces, dashboards, and alerts. • Manage infrastructure as code across the organization, with a focus on consistency, change safety, and reproducibility. • Own platform security practices—including secrets management, IAM policies, and network segmentation. • Support compliance-aware infrastructure practices—including vulnerability management, access reviews, audit-evidence flows, and incident-response readiness. • Participate in a frequent on-call rotation; drive incident response, blameless post-mortems, and follow-through on systemic fixes. • Partner with product and data engineering teams to troubleshoot platform issues and guide developers on infrastructure best practices.

🎯 Requirements

• Bachelor's degree or higher in Computer Science, Software Engineering, or a related field. • 5+ years of experience in infrastructure, platform, DevOps, or a related engineering role, with a track record of measurably improving production reliability—including defining SLOs, reducing incident frequency or MTTR, and eliminating recurring failure modes. • Hands-on experience with Google Cloud Platform (GCP), including GCE, GCS, VPC, IAM, Cloud Monitoring, and related services. • Experience with Docker and Kubernetes (GKE), including containerizing workloads, Helm, and cluster fundamentals. • Experience with CI/CD pipelines such as GitHub Actions, ArgoCD, Jenkins, or similar tools. • Experience with an observability platform such as New Relic, including metrics, logging, alerting, and dashboards. • Proficiency with Terraform for managing infrastructure as code. • Scripting or programming experience with Python, Bash, or similar languages.

🏖️ Benefits

• Join a mission-driven company making a meaningful impact on children's literacy and education. • Work alongside talented teammates in a collaborative, supportive, and global environment. • Enjoy the flexibility of a fully remote, U.S.-based position. • Help build and scale the infrastructure powering millions of young readers around the world.

Apply Now

Similar Jobs

🕒 July 16

Global Alliant Inc

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Fullstack DevSecOps Engineer developing enterprise applications on AWS. Collaborating in Agile environments with strong focus on secure coding practices.

🕒 July 16

CACI International Inc

10,000+ employees

💼 Consulting

🎖️ Defense

CACI Cloud DevSecOps Engineer developing scalable software solutions for a cloud-native ecosystem. Collaborate with data scientists and DevSecOps engineers for analytics environment enhancements.

🇺🇸 United States – Remote

💵 $105.1k - $231.1k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 16

NetImpact Strategies Inc.

201 - 500

💼 Consulting

🏥 Healthcare

🎖️ Defense

AI/DevSecOps Installation Architect at NetImpact delivering innovative solutions for Federal Government. Collaborating with teams to integrate AI technologies for enhancing digital transformation.

🇺🇸 United States – Remote

💵 $145k - $175k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 16

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior Site Reliability Engineer II at Akamai collaborating across software development, operations, and network teams. Responsible for developing standards and tooling for global platform stability.

🇺🇸 United States – Remote

💵 $146.4k - $263.6k / year

💰 Post-IPO Equity on 2001-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 July 16

Horizon3.ai

51 - 200

🔒 Cybersecurity

🤖 Artificial Intelligence

☁️ SaaS

Senior Engineering Manager responsible for building Site Reliability function at Horizon3.ai. Leading hiring and development of a team while ensuring operational excellence and cultural change across engineering.

🇺🇸 United States – Remote

💵 $260k - $280k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)