Search Remote Jobs

Data Site Reliability Engineer, SRE

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $111.2k - $150.4k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of General Dynamics Information Technology

General Dynamics Information Technology

10,000+ employees

Founded 1954

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

General Dynamics Information Technology is a company at the forefront of technological innovation, offering a wide range of services including consulting, digital modernization, and application services. The company is heavily involved in implementing solutions related to artificial intelligence, cloud computing, cybersecurity, high-performance computing, and quantum technologies. GDIT is committed to supporting government and defense sectors, providing mission-critical services such as logistics and supply chain management, intelligence, and homeland security. The company also focuses on diverse and inclusive hiring practices and actively promotes employee well-being. Through its digital accelerator solutions and pioneering use of emerging technologies, GDIT aims to propel agencies' missions forward and address complex technological challenges.

📋 Description

• Provide technical leadership for day-to-day operational support, reliability, performance, and continuous improvement of CMM data platforms, pipelines, applications, and analytics services • Provide real-time monitoring, incident and event management, capacity planning, and operational reporting • Maintain and audit cloud user roles and responsibilities • Integrate SSO, MFA, and group identity management through JENIE to enforce least-privilege access • Assess and improve credential management processes • Provide disaster recovery and continuity of operations options, including fault-tolerant and automated failover designs • Integrate DevSecOps tools and processes with enterprise systems • Manage centralized secrets management with automated rotation, access logging, and policy enforcement • Integrate SAST, DAST, SCA, and CSPM security tools into pipelines • Implement continuous 24/7/365 monitoring for security, performance, and compliance with dashboards and alerting • Provide supplemental monitoring of event response activities beyond normal business hours • Automate generation and management of SBOMs for deployed artifacts • Provide diagnostics, metrics gathering, and performance tuning • Provide canary releases for end-user and beta testing • Configure alerts for unusual behavior • Operate incident and event management processes integrated with enterprise SIEM solutions • Detect, log, diagnose, escalate, and resolve incidents; perform root cause analysis • Identify and eliminate recurring incident root causes • Recommend improvements to incident and problem management • Maintain a knowledge base of known issues, resolutions, and best practices • Perform automated full-stack health checks across operating systems, applications, databases, and PaaS services • Produce monthly issues management reports • Develop thresholds, rules, and response procedures • Monitor cloud resource utilization and manage threshold-breach resolution • Improve reliability, observability, automation, scalability, and operational resilience • Monitor, maintain, and optimize cloud infrastructure, databases, and platform services • Act as FinOps Analyst and perform cost optimization • Support the CMM program, a cloud-based solution for the Administrative Office of the US Courts serving 204+ federal courts

🎯 Requirements

• Bachelor's degree in Computer Science, Software Engineering, or related field, or equivalent experience • 5+ years of experience in IT systems engineering, systems development, systems coding, and programming • Must be a US Person: Green Card Holder, US Permanent Resident Alien, Refugee, Asylee, or US Citizen • Must be able to pass a background check to obtain a position of Public Trust • Deep expertise with AWS services, including monitoring, logging, compute, storage, and networking • Proficiency with Infrastructure as Code tools such as Terraform, AWS CloudFormation, or Azure Bicep • Hands-on experience with monitoring and APM tools such as CloudWatch, Azure Monitor, Datadog, Prometheus, Grafana, or New Relic • Understanding of incident response, change management, and ITIL-based operational support • Familiarity with CI/CD toolchains and automation platforms including Jenkins, GitHub Actions, GitLab, and ArgoCD • Strong scripting skills in Python, PowerShell, and Bash • Advanced DevSecOps implementation experience using GitOps or similar tools • Experience developing, testing, and maintaining containerized applications • Expert knowledge of source version control, build/release tools, CI/CD pipelines, and software build processes • Experience building and maintaining CI/CD pipelines for large enterprises with complex applications • Experience with FinOps practices, cost modeling, forecasting, and cloud optimization tools • Understanding of federal compliance and security frameworks such as FedRAMP, NIST, and JISF Rev 5 • Ability to analyze logs and metrics and conduct performance tuning for cloud services and applications • Experience working across multiple product teams to assess overall product/program health • ITIL, AWS SysOps, or Google Professional Cloud DevOps Engineer certifications are a plus • Excellent presentation and communication skills • Consultant mindset and ability to work with high-level customer stakeholders • Strong analytical and problem-solving skills • Experience with process design and documentation methodologies, deliverables, process/use case modeling, and business case development • Ability to work effectively independently and as part of a team • Ability to work flexibly across several products while supporting multiple teams

🏖️ Benefits

• Comprehensive medical plan options, some with Health Savings Accounts • Dental plan options • Vision plan options • 401(k) plan with pre-tax and post-tax contributions and company match • Full-flex work weeks where possible • Paid time off, including vacation, sick, personal time, holidays, paid parental, military, bereavement, and jury duty leave • Typically 15 days of paid leave per calendar year • 10 paid holidays per year • Up to 160 hours of paid family leave in a rolling 12-month period for eligible employees • Short- and long-term disability benefits • Life insurance • Accidental death and dismemberment insurance • Personal accident insurance • Critical illness insurance • Business travel and accident insurance • AI-powered career tool identifying career steps and learning opportunities • Internal mobility team • Wellness packages • Competitive pay • Award-winning culture of innovation and military-friendly workplace

Apply Now

Similar Jobs

🔥 54 minutes ago

Modivcare

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

🚗 Transport

Senior Manager, DevOps supporting Modivcare’s service-oriented healthcare transportation operations remotely. Performing essential DevOps functions with extensive computer and telephone use.

🔥 4 hours ago

PointClickCare

1001 - 5000

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Senior SRE operating PointClickCare’s cloud healthcare platform serving 30,000+ provider organizations. Automating resilient, observable mission-critical application operations.

🔥 10 hours ago

Astronomer

201 - 500

🏥 Healthcare

🏭 Manufacturing

💼 Consulting

Customer Reliability Engineer maintaining Astronomer's managed Airflow platform, cloud infrastructure, and Kubernetes clusters. Troubleshooting customer environments, operating observability systems, and resolving incidents.

🔥 16 hours ago

Ninety

51 - 200

☁️ SaaS

🤝 B2B

⚡ Productivity

SRE Delivery Manager leading SRE and Delivery teams for Ninety’s EOS business-management software. Improving AWS infrastructure, deployments, observability, security, and incident response.

🇺🇸 United States – Remote

💵 $200k - $220k / year

💰 $35M Series B - Ninety on 2023-11

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 18 hours ago

Prompt Therapy Solutions Inc

11 - 50

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior Database Reliability Engineer maintaining Aurora MySQL reliability, performance, and cost efficiency. Scaling Prompt’s automated healthcare software platform for rehab therapy businesses.