Site Reliability Engineer – Observability Platform

🔥 5 minutes ago

☕ Washington – Remote

infoinfo

💵 $85.4k - $192.9k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Ford Motor Company

Ford Motor Company

10,000+ employees

Founded 1903

📦 Logistics

💼 Consulting

📣 Marketing

💰 Post-IPO Debt on 2023-08

Logistics • Consulting • Marketing

Ford Motor Company is a globally renowned automotive company based in the United States, established by Henry Ford. The company is committed to building a better world where every individual has the freedom to move and follow their dreams. Ford is dedicated to innovation, with a focus on services, experiences, and software alongside its traditional vehicle manufacturing. The company is actively involved in sustainability initiatives and aims to meet ambitious environmental targets. Ford values service, community impact, and strives to combine business success with social and environmental responsibility. With a rich history of over 121 years, Ford continues to adapt and lead in the evolving automotive landscape.

📋 Description

• Design and implement scalable observability pipelines spanning metrics, logging, tracing, and alerting • Define and operationalize SLIs and SLOs and establish error budgets • Build reusable infrastructure-as-code templates and frameworks for observability instrumentation and onboarding • Architect, design, and develop automation to improve application resilience, recoverability, availability, and scalability • Perform destructive testing to discover vulnerabilities • Develop tooling to improve reliability, quality, and time-to-market • Reduce or eliminate toil through automation • Collaborate with development teams to build and operate scalable, resilient, cloud-native systems • Identify stability risks and establish mitigation plans with engineering leadership • Review technical metrics including errors, response times, caching, capacity, and resource utilization • Conduct performance analysis and optimization of new and production systems • Solve complex architecture, design, and business problems by simplifying processes and removing bottlenecks • Evaluate and integrate emerging technologies and architectures • Troubleshoot distributed production systems and drive root-cause analysis for platform incidents • Participate in incident response, support, recovery, and postmortem analysis • Provide technical guidance and mentorship • Integrate AI/ML capabilities to enhance anomaly detection, alerting precision, and performance insights • Embed observability best practices into system design and deployment workflows

🎯 Requirements

• Bachelor’s Degree in Computer Science or equivalent experience • 3+ years of experience in an SRE role • 5+ years of programming experience with one or more of: Python, Go, Java/Scala, C, or C++ • 3+ years of experience building reusable infrastructure-as-code templates and frameworks in Terraform or ToFu • 3+ years of experience with APM and monitoring tools such as Dynatrace, New Relic, ELK, Splunk, Prometheus, Sensu, Nagios, Kafka, or DataDog • 3+ years of experience with J2EE, NoSQL/SQL datastores, Spring Boot, GCP/AWS/Azure, and Docker/Kubernetes in developing multi-tier applications • Experience with RESTful APIs and microservices platforms • Working knowledge of the TCP/IP stack, internet routing, and load balancing • Strong proficiency with Google Cloud Platform and its library of services • Experience with automated, test-driven development in CI/CD pipelines • Thorough understanding of software development and agile methodologies • Understanding of, and ability to implement, effective observability strategies to improve MTTD/MTTR • Must be legally authorized to work in the United States • Visa sponsorship is not available for this position

🏖️ Benefits

• Immediate medical, dental, vision and prescription drug coverage • Flexible family care days • Paid parental leave • New parent ramp-up programs • Subsidized back-up childcare • Family building benefits including adoption and surrogacy expense reimbursement and fertility treatments • Vehicle discount program for employees and family members and management leases • Tuition assistance • Established and active employee resource groups • Paid time off for individual and team community service • Generous schedule of paid holidays, including the week between Christmas and New Year’s Day • Paid time off and the option to purchase additional vacation time

Apply Now

Similar Jobs

🔥 3 hours ago

Entarian

1001 - 5000

🚀 Aerospace

🎖️ Defense

🏛️ Government

Senior Site-Reliability Engineer maintaining reliable Windows production infrastructure for Entarian’s mission-critical engineering solutions. Automating operations, monitoring services, and improving deployment reliability.

🔥 3 hours ago

URBN (Urban Outfitters, Anthropologie Group, Free People & Nuuly)

10,000+ employees

👥 B2C

🛒 Retail

👗 Fashion

DevOps Engineer scaling Nuuly’s GCP, Kubernetes, and Kafka infrastructure for its fashion rental platform. Automating deployments, improving reliability, and optimizing event-driven systems.

🔥 4 hours ago

Innosphere

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior Site Reliability Engineer building reliable AWS infrastructure and CI/CD systems. Supporting Innosphere’s distributed technology staffing and software development teams.

🔥 4 hours ago

Harris Computer

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

DevSecOps Engineer securing CI/CD pipelines, cloud infrastructure, containers, and vulnerability management. Supporting STChealth’s technology platform for immunization data exchange and public health.

🔥 4 hours ago

Harris Computer

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

Platform & DevSecOps Architect building automated CI/CD, security, and AI infrastructure for STChealth’s public-health software. Enabling compliant delivery across Kubernetes and legacy systems.