Site Reliability Engineer

🕒 July 2

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 33%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Offchain Labs

Offchain Labs

11 - 50 employees

Founded 2018

₿ Crypto

🌐 Web 3

Crypto • Web 3 • Blockchain Technology

Offchain Labs is a venture-backed company founded at Princeton that specializes in developing blockchain technologies, particularly for scaling the Ethereum network. Known for its groundbreaking Arbitrum technology, Offchain Labs offers a suite of products aimed at enhancing Ethereum's capabilities, including Arbitrum One, Arbitrum Nova, and Arbitrum Orbit. The company also provides tools like Prysm, a leading consensus client for Ethereum, and Stylus, a programming environment for deploying EVM-compatible smart contracts. Offchain Labs focuses on creating innovative and accessible solutions for developers, businesses, and individuals to utilize Ethereum effectively.

📋 Description

• Operate production Kubernetes clusters and build scalable, declarative infrastructure using Terraform or similar tools • Deploy and maintain Kubernetes environments, manage system components, and troubleshoot applications running on the platform • Design CI/CD workflows with ArgoCD, GitHub Actions, CodeBuild, or similar tools for infrastructure and application deployments • Design and operate observability systems using time-series metrics, logs, and dashboards with Prometheus, Loki, Mimir, Grafana, and CloudWatch • Diagnose networking and storage issues across complex, distributed systems • Implement secure-by-default infrastructure and contribute to architecture reviews and threat models • Automate operational workflows using Python, Go, or Bash • Participate in on-call rotations, respond to incidents, troubleshoot under pressure, and drive postmortems to improve system reliability

🎯 Requirements

• Eager to dive into blockchain technology, even if it’s new territory • Enjoy solving infrastructure problems in unconventional ways and thinking beyond standard patterns • Use tools like k9s or ArgoCD for speed and abstraction, but comfortable dropping into YAML, logs, or low-level debugging when things go sideways • Experienced with GitOps-style systems and treating both infrastructure and application delivery as code • Have scaled deployment automation using patterns like ArgoCD ApplicationSets or similar tooling • Curious about how things work under the hood and not satisfied with surface-level fixes • Comfortable in Linux, fluent in shell scripting, and productive in languages like Python or Go • Comfortable operating within a cloud platform (e.g., AWS, GCP, Azure), with a strong understanding of the underlying components making it easy to adapt to or migrate across providers • Participated in an on-call rotation, responding to incidents, troubleshooting under pressure, and driving postmortems to improve system reliability over time • Design systems with security in mind, applying principles like least privilege and threat modeling • Bring a strong technical foundation, excellent problem-solving skills, and a genuine commitment to high-quality work • Take ownership, collaborate openly, and contribute to a culture of clarity, curiosity, and continuous improvement • Operated production Kubernetes clusters and built scalable, declarative infrastructure using Terraform or similar tools • Deployed and maintained Kubernetes environments, managed system components, and troubleshot applications running on the platform • Designed CI/CD workflows with ArgoCD, GitHub Actions, CodeBuild, or similar tools, covering both infra and app deployments • Designed and operated observability systems using time-series metrics, logs, and dashboards with tools like Prometheus, Loki, Mimir, Grafana, and CloudWatch • Diagnosed tough networking and storage issues across complex, distributed systems • Implemented secure-by-default infrastructure and contributed to architecture reviews and threat models • Automated operational workflows using scripting or programming in Python, Go, or Bash

🏖️ Benefits

• Remote-first global workforce + NY office • Professional reimbursement program (facilitates industry conference attendance, certifications, and more) • Medical, dental & vision coverage (US + some other countries) • 401k retirement plan + company match (US only) • Wellness stipend • Home office set up / ergonomic equipment program

Apply Now

Similar Jobs

🕒 July 2

Sanity.io

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

SRE managing scalable content operations infrastructure for AI-powered platform. Collaborating with dev teams and ensuring reliability for high request volume systems.

🕒 June 30

Casa Inc.

11 - 50

💼 Consulting

💸 Finance

₿ Crypto

DevSecOps Engineer securing infrastructure, software, and employee access at Casa, a Bitcoin self-custody provider. Automating controls, addressing vulnerabilities, and strengthening its security program.

🕒 June 30

Zapata Technology

51 - 200

🎖️ Defense

🔒 Cybersecurity

🤖 Artificial Intelligence

DevSecOps Engineer developing and resolving technical issues for logistics middleware applications in IT team. Collaborating with engineers and customers to ensure data quality and support development.

🕒 June 30

CloudBees

501 - 1000

🤝 B2B

Consultant guiding customers on DevSecOps strategies and software delivery transformations. Collaborating with engineering leaders and executives to improve outcomes across various industries.

🕒 June 29

Motorola Solutions

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

AWS DevOps Full-Stack Engineer responsible for maintaining cloud infrastructure for Motorola's safety systems. Collaborating with teams to automate deployment and ensure security compliance within the U.S.