Senior Site Reliability Engineer – Fedramp

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $85k - $141k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Coalfire

Coalfire

1001 - 5000 employees

Founded 2001

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Coalfire is a cybersecurity services provider that helps businesses improve their security resilience and streamline regulatory compliance. The company offers expert-led services, including threat-focused cybersecurity programs, compliance automation, risk management, and security advisory services across various industries such as financial services, healthcare, retail, and technology. Coalfire is known for its hacker and defender expertise, and its platforms are designed to fortify clients' cyber resilience, reduce attack surfaces, and accelerate the achievement of compliance objectives like FedRAMP and HITRUST.

📋 Description

• Own an operational capability for the managed estate, including automation, runbooks, and service standards • Design observability for regulated cloud environments, including telemetry and log pipelines, service-level objectives, alert quality, and escalation paths • Build and maintain continuous-monitoring evidence pipelines • Own backup and recovery engineering, including tested recovery procedures, measurable recovery objectives, and outage automation • Serve as the senior escalation point in client environments and resolve complex operational incidents • Lead incident and problem management, including major-event incident command, blameless post-incident reviews, and corrective actions • Automate operational toil using infrastructure-as-code, pipelines, and scripting • Partner with Engagement Architects and Build teams on transition into managed operations • Hold on-call responsibility and improve rotation coverage, alert actionability, and team load • Represent operational posture to clients and support renewals and expansions • Mentor and lead Site Reliability Engineers and junior staff • Author and peer review code, runbooks, operational design documentation, and compliance artifacts

🎯 Requirements

• BS or above in a related Information Technology field or equivalent combination of education and experience • Bachelor’s degree or equivalent combination of education and work experience • Professional- or specialty-level certification in AWS, Azure, or GCP; associate-level certification considered with equivalent demonstrated depth • 5+ years in site reliability engineering, cloud operations, platform engineering, or managed services • 5+ years operating production cloud environments in AWS, Azure, or GCP, including monitoring, incident response, and automation • Automation-first mindset with deep Infrastructure-as-Code, CI/CD, scripting, and policy-as-code • Deep operational command of at least one major cloud platform and working knowledge of a second • Observability engineering experience with metrics, logging, log pipelines, distributed tracing, SLI/SLO definition, and alert design • Demonstrated incident response and incident command capability • Backup, recovery, and resilience engineering experience • Working command of NIST 800-53, FedRAMP, or comparable security control frameworks • Ability to lead technical client conversations about operational posture, risk, and trade-offs • Demonstrated ability to mentor engineers and improve team output • Excellent communication, organizational, and problem-solving skills • Effective documentation skills, including technical diagrams, runbooks, and written descriptions • Ability to work independently and as part of a team • Critical thinking and ability to balance security and availability requirements against mission needs • Demonstrated experience owning an operational capability, monitoring platform, or reusable automation used by multiple teams or clients • Experience as the senior operational escalation point on client-facing managed services, including incident command on major events • Advanced experience with Infrastructure-as-Code and orchestration/automation tools such as Terraform and Ansible • Experience transitioning environments from build into steady-state operations

🏖️ Benefits

• Flexible work model allowing employees to choose when and where they work • Paid parental leave • Flexible time off • Certification and training reimbursement • Digital mental health and wellbeing support membership • Comprehensive insurance options • Employee resource groups • In-person and virtual events • Annual incentive, commission, and/or recognition programs may be available

Apply Now

Similar Jobs

🔥 7 minutes ago

OnePay

501 - 1000

💳 Fintech

🏦 Banking

₿ Crypto

SRE Lead building reliable infrastructure for OnePay’s consumer fintech platform. Leading senior engineers while coding, automating operations, and improving incident response.

🔥 43 minutes ago

Ad Hoc LLC

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevOps Engineer building AWS infrastructure and CI/CD pipelines for Ad Hoc’s Veterans Affairs digital services. Improving security, reliability, developer experience, and software delivery speed.

🔥 2 hours ago

Octus

501 - 1000

💼 Consulting

⚖️ Legal

📚 Education

Lead DevOps Engineer leading cloud infrastructure, CI/CD, and security for Octus, a global credit intelligence and analytics provider. Mentoring DevOps engineers and ensuring reliable, scalable systems.

🔥 2 hours ago

PhoenixTeam

51 - 200

💳 Fintech

🏠 Real Estate

🤖 Artificial Intelligence

DevOps Manager modernizing Jenkins-based CI/CD and Fortify quality controls for PhoenixTeam's federal FHA mortgage program. Coordinating releases, documentation, and delivery across development teams.

🔥 5 hours ago

Sprezzatura

51 - 200

🏛️ Government

💼 Consulting

🏥 Healthcare

DevSecOps Engineer building AWS, Kubernetes, and Terraform infrastructure for VA.gov. Automating secure deployments and platform services supporting millions of Veterans.