Search Remote Jobs

Senior Site Reliability Engineer

🕒 6 days ago

🍂 Massachusetts – Remote

info

đŸ’” $121.4k - $218.6k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🩅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 employees

🔒 Cybersecurity

💰 Post-IPO Equity on 2001-07

Cloud Computing ‱ Cybersecurity ‱ Content Delivery

Akamai Technologies is a leading cloud services provider that specializes in delivering security, cloud computing, and content delivery solutions. It offers a range of services such as API security, DDoS protection, and performance optimization for web applications, ensuring secure and reliable user experiences. With a robust global infrastructure, Akamai empowers businesses to streamline their digital presence while safeguarding against various cyber threats and enhancing application performance.

📋 Description

‱ Oversee, scale, and optimize next-generation dedicated AI hardware infrastructure ‱ Ensure uptime and reliability of AI hardware infrastructure offerings ‱ Collaborate with product teams from early development through deployment to ensure reliability, scalability, and performance ‱ Define key performance indicators and defend them when breached ‱ Develop Python programmatic tooling and infrastructure-as-code utilities to automate fleet-wide provisioning and reduce operational toil ‱ Integrate automated workflows across corporate ticketing systems for hardware and network break-fix events ‱ Use AI utilities and LLM-assisted development to accelerate technical execution and system analysis ‱ Improve availability, latency, and systemic health of high-density hardware environments using private cloud and compute technologies ‱ Design telemetry pipelines, Prometheus/Grafana dashboards, and AI-based anomaly detection for bare-metal and virtualized environments ‱ Participate in 24x7x365 on-call rotations and lead real-time incident management ‱ Manage high-severity service disruption protocols through PagerDuty and Slack workflows ‱ Partner with third-party infrastructure vendors and coordinate on-site field technicians

🎯 Requirements

‱ 5 years of relevant experience ‱ Bachelor's degree in Computer Engineering, Computer Science or equivalent ‱ Tooling and coding ability in Python for scalable operational tools, API integrations, and automation frameworks ‱ Hands-on experience with Prometheus, Grafana, OpenTelemetry, and Loki ‱ Working understanding of advanced networking topologies and high-bandwidth routing/switching infrastructure ‱ Knowledge of BGP and dual-stack IPv4/IPv6 networks ‱ Experience designing new service rollouts, including operational readiness criteria, telemetry baselines, and alerting thresholds ‱ Extensive experience building technical runbooks, leading complex incident response bridges, and driving blameless post-mortems ‱ Ability to own ambiguous technical problems, coordinate cross-functional teams, and deliver production-grade solutions ‱ Participation in 24x7x365 on-call rotations

đŸ–ïž Benefits

‱ Flexible work through Akamai's FlexBase program: at home, in an office, or a combination of both ‱ Annual bonus or incentives ‱ Equity awards ‱ Employee Stock Purchase Plan (ESPP) ‱ Healthcare ‱ 401K savings plan ‱ Company holidays ‱ Vacation/PTO ‱ Sick time ‱ Parental leave ‱ Employee assistance program ‱ Mental and financial wellness support

Apply Now

Similar Jobs

🕒 6 days ago

PrizePicks

201 - 500

🎼 Gaming

⚜ Sports

Senior SRE ensuring reliable, scalable infrastructure for PrizePicks’ daily fantasy sports platform. Leading incident response, observability, Kubernetes operations, and reliability improvements.

đŸ‡ș🇾 United States – Remote

đŸ’” $120k - $175k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 6 days ago

Veeam Software

1001 - 5000

đŸ’Œ Consulting

📩 Logistics

☁ SaaS

Senior SRE building reliability engineering for Veeam Data Cloud’s Government and Sovereign Cloud SaaS platform. Designing Azure infrastructure, observability, incident response, and compliance-ready delivery practices.

đŸ‡ș🇾 United States – Remote

đŸ’” $158.4k - $294.1k / year

💰 $500M Private Equity Round on 2019-01

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🩅 H1B Visa Sponsor

info

🕒 6 days ago

Veeam Software

1001 - 5000

đŸ’Œ Consulting

📩 Logistics

☁ SaaS

Site Reliability Engineer building reliability practices for Veeam’s Government and Sovereign Cloud SaaS platform. Designing Azure infrastructure, observability, automation, and incident-response systems in regulated environments.

đŸ‡ș🇾 United States – Remote

đŸ’” $138.9k - $231.4k / year

💰 $500M Private Equity Round on 2019-01

⏰ Full Time

🟠 Senior

🔮 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🩅 H1B Visa Sponsor

info

🕒 6 days ago

Karat

201 - 500

đŸ‘„ HR Tech

🏱 Enterprise

☁ SaaS

Senior Deployment Engineer helping Karat, a technical interviewing company, implement and optimize enterprise interview frameworks. Advising clients, analyzing hiring performance, and delivering executive training.

đŸ‡ș🇾 United States – Remote

đŸ’” $116.9k - $148.2k / year

💰 Funding Round on 2022-04

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🩅 H1B Visa Sponsor

info

🕒 6 days ago

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏱 Enterprise

Site Reliability Engineer architecting Cisco’s developer platform and infrastructure for cloud application delivery. Consolidating legacy tools and improving scalable, resilient engineering workflows.

đŸ‡ș🇾 United States – Remote

đŸ’” $138.1k - $198.2k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🩅 H1B Visa Sponsor

info