Site Reliability Engineer II

Job not on LinkedIn

🕒 June 18

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Pinterest

Pinterest

1001 - 5000 employees

Founded 2010

📱 Media

👥 B2C

💰 Post IPO equity on 2022-08

Media • B2C

Pinterest is a visual discovery and bookmarking platform that helps users discover, save, and organize images and ideas (Pins) for interests, projects, and shopping. It provides personalized recommendations and visual search so people can plan and explore topics like home design, fashion, recipes, and events, and it supports commerce features and ads for creators and retailers.

📋 Description

• Ensuring the reliability, availability, and performance of production infrastructure and platform services • Operating and scaling Kubernetes platforms, including governance and support for multi-tenant workloads • Managing GitOps-based deployment workflows using ArgoCD and Helm • Supporting infrastructure provisioning and change management through Terraform/Terragrunt • Building and supporting CI/CD automation and deployment workflows using GitHub Actions • Participating in incident response, root cause analysis, and post-incident improvement initiatives • Reducing operational toil through scripting, tooling, and process automation • Advancing observability practices across logs, metrics, traces, dashboards, and alerting • Supporting secure secrets integration, IAM-aware operations, and platform guardrails • Partnering closely with application, security, and platform teams to improve reliability and delivery outcomes

🎯 Requirements

• 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure • Strong hands-on experience operating AWS in production environments • Good expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration • Experience with Kubernetes multi-tenancy, including namespaces, RBAC, quotas, policies, and tenant isolation patterns • Experience implementing and operating ArgoCD within a GitOps delivery model • Strong hands-on experience with Helm • Experience with Terraform/Terragrunt for infrastructure provisioning and environment management • Solid scripting and automation skills using Bash and/or Python • Experience building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions • Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems • Experience with monitoring, alerting, and observability in production environments • Demonstrated ownership mindset with experience handling incidents and resolving production issues • Strong collaboration and communication skills, with the ability to work effectively across engineering, security, and platform teams • Bachelor’s degree in computer science, engineering, a related field or equivalent experience • Demonstrated ability to use AI to improve speed and quality in your day-to-day workflow for relevant outputs • Strong track record of critical evaluation and verification of AI-assisted work (e.g., testing, source-checking, data validation, peer review) • High integrity and ownership: you protect sensitive data, avoid over-reliance on AI, and remain accountable for final decisions and deliverables.

🏖️ Benefits

• Information regarding the culture at Pinterest and benefits available for this position can be found [here](https://www.pinterestcareers.com/pinterest-life/)

Apply Now

Similar Jobs

🕒 June 18

Aledade, Inc.

501 - 1000

🏥 Healthcare

⚕️ Healthcare Insurance

🏢 Enterprise

Salesforce DevOps Analyst ensuring quality and reliability of Salesforce solutions. Collaborating with teams for test strategies, automation, and CI/CD management.

Cloud

Cyber Security

Jenkins

Selenium

🕒 June 18

Guidehouse

10,000+ employees

🏥 Healthcare

🎖️ Defense

📦 Logistics

Site Reliability Engineer collaborating with teams to establish SRE practices and participate in system design reviews at Guidehouse. Focused on AWS cloud infrastructure and promoting automation.

Ansible

AWS

Azure

Cloud

Linux

Packer

Python

SDLC

Terraform

🕒 June 17

Intermedia Cloud Communications

1001 - 5000

💼 Consulting

🏥 Healthcare

⚖️ Legal

DevOps Engineer managing GCP infrastructure for cloud communications. Collaborating with development teams to maintain application deployment and infrastructure.

Ansible

Cloud

Docker

ElasticSearch

Google Cloud Platform

Jenkins

Kubernetes

Linux

Postgres

Python

RabbitMQ

Redis

Terraform

Go

🕒 June 17

Accela

201 - 500

💼 Consulting

📦 Logistics

🏗️ Construction

Lead Site Reliability Engineer ensuring scalability, performance, and operational excellence of Accela's Civic Platform through technical leadership and cloud modernization. Collaborating with cross-functional teams to enhance SaaS offerings for government software solutions.

Azure

Cloud

Distributed Systems

Kubernetes

Python

🕒 June 17

Clinician Nexus

51 - 200

🏥 Healthcare

⚕️ Healthcare Insurance

📚 Education

Senior Manager, DevOps enabling health organizations with technology while leading a team of engineers. Overseeing CI/CD, infrastructure automation, and collaborating with multiple stakeholders.

AWS

Cloud

DNS

Flux

Grafana

Kubernetes

Prometheus

Python

Terraform

Vault