Senior Manager, Site Reliability Engineering

🕒 vor 1 Monat

🇺🇸 Vereinigte Staaten – Remote

💵 $155.000 - $170.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

👻 Geisterscore 2%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Claritas Rx

Claritas Rx

51 - 200 Mitarbeiter

🏥 Gesundheitswesen

💼 Beratung

📦 Logistik

💰 Private Equity Round im 2021-03

Healthcare • Consulting • Logistics

Claritas Rx ist ein Unternehmen für Gesundheitsanalytik, das KI nutzt, um den Patientenzugang zu lebensverändernden Behandlungen zu verbessern. Das Unternehmen bietet eine Reihe von Lösungen, darunter die Ascend AI Platform, Patient Watchtower, Patient Services CRM und Cell and Gene Performance Benchmarking. Die Angebote konzentrieren sich darauf, die Sichtbarkeit der Patienten zu verbessern, frühe Risiken vorherzusagen, Partner effizient einzubinden und Strategien zur Patientenunterstützung zu optimieren. Durch den Einsatz von KI-gesteuerter Analytik hilft Claritas Rx Gesundheitsdienstleistern, die Erfüllungsraten und die Therapietreue ihrer Marken zu steigern und sicherzustellen, dass Patienten nicht durchs Raster fallen. Das Unternehmen ist auf die Bereitstellung von Einblicken in Zugangsprobleme von Patienten und Performance-Benchmarking in der Pharmaindustrie spezialisiert.

Beschreibung

• Own the reliability, availability, performance, and scalability of Claritas Rx's AWS-hosted SaaS platform, with accountability for SLA/SLO commitments made to customers. • Define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all production services; use error budgets to drive engineering prioritization conversations. • Lead incident response and on-call operations: triage, coordinate resolution, communicate to stakeholders, and conduct thorough post-incident reviews with actionable corrective actions. • Drive a proactive reliability culture — identifying risks before they become incidents through load testing, chaos engineering, and systematic failure mode analysis. • Architect, build, and maintain AWS cloud-native infrastructure using infrastructure-as-code (AWS CDK, Terraform, or equivalent), ensuring environments are reproducible, auditable, and secure. • Oversee and continuously improve CI/CD pipelines (GitHub Actions) to enable rapid, safe, and consistent delivery of application and infrastructure changes across environments. • Manage and optimize core AWS services including ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, and CloudFront. • Ensure robust observability across the stack — centralizing logs, metrics, traces, and alerts using CloudWatch, Sentry, and related tooling — so the team can detect and respond to issues quickly. • Manage platform capacity planning, cost optimization, and cloud spend governance. • Ensure all infrastructure design and operational practices meet HIPAA, SOC 2, and HITRUST requirements, given the PHI our platform processes. • Partner with the Security function on vulnerability management, infrastructure hardening, secrets management, and access control. • Maintain and regularly test disaster recovery (DR) and business continuity plans, including defined RTO/RPO targets for all production systems. • Support audit readiness and evidence collection for compliance certifications. • Lead, mentor, and grow a blended team of full-time SREs/DevOps engineers and offshore contractors, fostering a culture of ownership, continuous improvement, and operational excellence. • Manage distributed team dynamics effectively — establishing clear communication rhythms, documentation standards, and handoff protocols to ensure offshore resources are productive and well-integrated. • Conduct regular 1:1s, set clear goals and development plans for direct reports, and advocate for your team's growth and recognition. • Build and maintain a healthy on-call rotation with appropriate tooling, runbooks, and escalation paths to protect team sustainability. • Collaborate with Software Engineering teams to embed reliability practices into the SDLC — including production readiness reviews, deployment standards, and shared observability tooling. • Work with Product Management and Engineering leadership to balance feature delivery velocity against operational risk and technical debt. • Contribute to architecture decisions across the platform, providing an operational and reliability perspective on new services and major technical changes. • Champion automation-first thinking: eliminate toil through tooling, scripting, and process improvement wherever possible.

🎯 Anforderungen

• 7+ years of experience in SRE, DevOps, or infrastructure engineering, with at least 3 years in a people management or team lead capacity. • Deep, hands-on AWS expertise — you understand how to architect, operate, and optimize cloud-native workloads at the service level, not just at a conceptual level. • Strong infrastructure-as-code skills (AWS CDK, Terraform, or equivalent); you treat infrastructure like software. • Demonstrated experience owning SLOs, incident management, and on-call operations in a commercial SaaS environment. • Experience managing CI/CD pipelines and developer productivity tooling, with a strong understanding of deployment safety (canary releases, feature flags, rollback strategies). • Solid working knowledge of security and compliance requirements relevant to regulated data environments (HIPAA, SOC 2, HITRUST). • Proven ability to lead and develop a team, including working effectively with offshore or distributed contractors across time zones. • Strong written and verbal communication skills — you can explain infrastructure risk and trade-offs clearly to both technical and non-technical audiences. • Comfort operating in a fast-paced, high-growth startup environment where priorities evolve and initiative is expected.

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

Paramount

10.000+ Mitarbeiter

💼 Beratung

📣 Marketing

📱 Medien

Lead DevOps Engineer building and maintaining scalable infrastructure for personalization services in Paramount's streaming platforms. Involves CI/CD pipeline management, cloud infrastructure, and Kubernetes expertise.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Snapsheet Inc

501 - 1000

💼 Beratung

📦 Logistik

📣 Marketing

Senior DevOps Engineer managing AWS cloud infrastructure for Snapsheet. Owns architecture, cloud networking, and database infrastructure while collaborating with Full Stack engineers.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Dyson

10.000+ Mitarbeiter

🔧 Hardware

🏭 Fertigung

🛒 Einzelhandel

Software Engineer II at Robert Half developing and supporting GenAI-enabled applications. Analyzing user needs and contributing to production deployment and documentation across all SDLC phases.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

ICF

5001 - 10000

🏥 Gesundheitswesen

📦 Logistik

📣 Marketing

Senior DevOps Engineer driving AWS cloud solutions within ICF's Health Engineering Solutions team. Implementing CI/CD pipelines and collaborating on best practices to improve cloud infrastructure.

🇺🇸 Vereinigte Staaten – Remote

💵 $108.476 - $184.409 / Jahr

💰 €30.000.000 Grant im 2021-03

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

HealthCatalyst Norway

2 - 10

🏥 Gesundheitswesen

💼 Beratung

🤝 B2B

Site Reliability Engineer on the Central AI team facilitating responsible AI adoption. Guiding teams with AI systems and governance under a healthcare performance improvement company.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich