Senior Site Reliability Engineer

🔥 17 hours ago

🌐 United States, Canada – Remote

infoinfo

🤠 Texas – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 13%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Megaport

Megaport

201 - 500 employees

Founded 2013

📡 Telecommunications

Networking • Cloud Computing • Telecommunications

Megaport is a leading provider of global private connectivity solutions that enable simplified network interconnection. The company offers a platform for deploying secure, scalable, and agile networks that interconnect data centers, clouds, and virtual points of presence. Megaport's services allow users to create secure and dynamic network connections on-demand, without hardware or long-term contracts, offering flexibility and speed to businesses. By partnering with global service providers, data center operators, and systems integrators, Megaport ensures robust and widespread network access across 930+ locations in 25 countries. Its smart software tools and APIs allow for easy network management, making it a trusted choice for cloud networking and hybrid cloud solutions.

📋 Description

• Improve production reliability and system resilience within an SRE-scoped team • Champion high standards and industry best practices • Communicate with teams and stakeholders throughout projects • Bring fresh ideas and encourage others • Investigate complex technical problems • Work across numerous technologies in a fast-changing industry • Participate in on-call rotation, incident response, and blameless post-incident reviews • Write code, handle alerts, improve solutions, and support others • Engage stakeholders in requirements analysis and demonstrations • Ensure systems are secure, maintainable, and available • Support customer success and company goals

🎯 Requirements

• 5+ years administering Linux systems and related infrastructure in production environments • Familiarity with SLIs, SLOs, SLAs, error budgets, blast radius, and blameless postmortems • Focus on automation, reducing toil, and preventing problem recurrence • Track record of writing runbooks for broader teams • Strong Kubernetes and broader ecosystem fundamentals • Cloud infrastructure experience; AWS strongly preferred • Bare-metal experience is a bonus • Strong tool development using Bash, plus Python or Go preferred, or similar • Infrastructure-as-code tooling experience; Terraform preferred • CI/CD and version control experience; GitHub preferred • Database experience with Postgres, Cassandra, or ClickHouse preferred • Experience operating production observability stacks covering metrics, logs, and traces • Strong troubleshooting instincts and ownership of incident response • History of continual professional development • Self-directed style suited to an async, globally distributed team • Comfortable picking up adjacent work when needed

🏖️ Benefits

• Flexible working environment – a remote-first culture with coworking options available. • 4 weeks of paid annual leave • Parental leave • Birthday leave • Purchased annual leave program • Wellness allowance • Employee wellbeing initiatives • Generous study and training allowance • 5 days of paid study leave • Creative, modern workspaces • Recognition programs, including Legend and Kudos awards

Apply Now

Similar Jobs

🔥 18 hours ago

Capstone Integrated Solutions

51 - 200

💼 Consulting

🛒 Retail

AWS DevOps Engineer building AWS cloud, automation, and MLOps infrastructure for CapNexus, a software development and systems integration services provider. Supporting SageMaker, Bedrock, CI/CD, security, and Azure-to-AWS migration.

🔥 18 hours ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Senior SRE maintaining NVIDIA’s managed DGX Cloud AI clusters across major cloud providers. Improving Kubernetes reliability, observability, GPU workloads, and production incident response.

🔥 19 hours ago

Pluribus Digital

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevOps Engineer designing secure, scalable Azure architectures for government agencies. Leading cloud migration, IaC, CI/CD, governance, and federal compliance strategies.

🔥 20 hours ago

FindErnest

11 - 50

💼 Consulting

🏢 Enterprise

🤝 B2B

DevOps Lead managing production Azure AKS environments and advanced Terraform IaC for IT services. Driving Flux CD, Helm, Dynatrace observability, CI/CD, security, and team performance.

🔥 20 hours ago

Guidehouse

10,000+ employees

🏥 Healthcare

🎖️ Defense

📦 Logistics

Cloud infrastructure engineer automating secure AWS and hybrid environments for Guidehouse’s federal clients. Building Terraform, CI/CD, security, monitoring, and disaster recovery capabilities.