Search Remote Jobs

Staff Site Reliability Engineer, Core AI Infrastructure

đź•’ June 9

🏄 California – Remote

infoinfo

đź’µ $218k - $256.5k / year

⏰ Full Time

đź”´ Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

đź‘» Ghost score 16%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Coinbase

Coinbase

1001 - 5000 employees

Founded 2012

đź’Ľ Consulting

₿ Crypto

đź’¸ Finance

đź’° $21.4M Post-IPO Equity on 2022-11

Consulting • Crypto • Finance

Coinbase is a leading cryptocurrency exchange platform that allows individuals and institutions to buy, sell, and trade various crypto assets such as Bitcoin and Ethereum. The company offers advanced trading tools, institutional solutions, and a self-hosted wallet for storing and managing cryptocurrencies. With a strong focus on security and transparency, Coinbase provides a trusted platform used by millions globally. It supports various features including staking, earning rewards, and spending crypto through their cards. Additionally, Coinbase provides developer tools and APIs for building onchain applications, making it a comprehensive hub for engaging in the crypto economy.

đź“‹ Description

• Own the reliability, monitoring, and incident response lifecycle for AI infrastructure services, including on-call support for AWS deployment pipelines, root cause analysis, and blameless retros. • Build automation and tooling to streamline operational IT workflows, eliminate manual tasks, and improve deployment velocity across CI/CD frameworks and Kubernetes environments. • Partner with the Coinbase Infrastructure team to extend CI/CD frameworks supporting IT services and enterprise network platforms, and with Security and Compliance to integrate surveillance tooling into deployment pipelines. • Strengthen observability and documentation standards across IT engineering by defining metrics, implementing monitoring solutions, and maintaining technical documentation that sets a standard of excellence. • Develop full-stack applications that power internal AI products and infrastructure with Go or Python.

🎯 Requirements

• 8+ years of experience automating and supporting cloud infrastructure (AWS) and network environments, with hands-on use of infrastructure-as-code tools (Terraform, Ansible, Chef, Puppet, or Salt). • Proven experience deploying, managing, and troubleshooting containerized workloads using Docker and Kubernetes in production environments. • Proficiency in at least one scripting or programming language (Python, Bash, Ruby, or Go) and version control workflows using Git-based CI/CD pipelines. • Track record of leading incident response in environments with strict SLAs, including root cause analysis, blameless retros, and measurable reliability improvements. • Utilizes generative AI responsibly, maintaining human oversight to deliver business-ready outputs and drive measurable improvements in workflow efficiency, cost, and quality.

🏖️ Benefits

• medical • dental • vision • 401(k)

Apply Now

Similar Jobs

đź•’ May 30

Ad Hoc LLC

501 - 1000

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

Staff DevOps Engineer at Ad Hoc, leading technical solutions and improving software engineering processes. Expert in CI/CD and infrastructure, with a focus on federal service delivery.

đź•’ May 29

Capgemini

10,000+ employees

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

Mainframe DevOps Migration Consultant at Capgemini Engineering supporting application migration projects utilizing client’s DBB/Git/IDD Solutions.

đź•’ May 29

Capgemini

10,000+ employees

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

Software Change Management Consultant supporting application migration projects utilizing IBM DBB/Git/IDD solutions. Leading technical training and troubleshooting in a remote capacity across North America.

đź•’ May 27

Titan AI

1 - 10

🎮 Gaming

🥽 AR/VR

🤖 Artificial Intelligence

VP of Site Reliability managing SRE and operational functions for banking AI software company. Leading engineering practices and ensuring reliable platform deployment for financial institutions.

đź•’ May 20

Crusoe

51 - 200

đź’Ľ Consulting

🏭 Manufacturing

📦 Logistics

Staff Instrumentation & Controls Engineer for Crusoe, focusing on deployment of automation solutions in hyperscale data centers. Overseeing complex projects, ensuring seamless integrations, and driving efficiencies.