
501 - 1000 employees
Founded 2015
đ¤ Artificial Intelligence
đ Aerospace
đď¸ Defense
Artificial Intelligence ⢠Aerospace ⢠Defense
Shield AI is a leading developer of AI-driven military solutions, focusing on enhancing mission autonomy and battlefield awareness. Their platform, Hivemind, enables rapid deployment of intelligent systems for various defense applications, including drone operation and surveillance. With a commitment to utilizing advanced technology, Shield AI aims to protect service members and civilians by revolutionizing defense technologies through autonomous systems.
đ August 12
đşđ¸ United States â Remote
đľ $180k - $270k / year
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 1%
Improve your chances of getting an interview by checking your resume score before you apply.

501 - 1000 employees
Founded 2015
đ¤ Artificial Intelligence
đ Aerospace
đď¸ Defense
Artificial Intelligence ⢠Aerospace ⢠Defense
Shield AI is a leading developer of AI-driven military solutions, focusing on enhancing mission autonomy and battlefield awareness. Their platform, Hivemind, enables rapid deployment of intelligent systems for various defense applications, including drone operation and surveillance. With a commitment to utilizing advanced technology, Shield AI aims to protect service members and civilians by revolutionizing defense technologies through autonomous systems.
⢠Own operational excellence for the Databricks platform, including monitoring, alerting, observability, incident response support, and production runbook patterns ⢠Define and maintain CI/CD and promotion standards for Databricks assets and environments from development to production ⢠Design and maintain standards for job orchestration, cluster and compute policies, service principal usage, environment isolation, and reliable production execution ⢠Establish reusable operational templates and enablement patterns for onboarding new domains to Databricks ⢠Partner with the Senior Data Engineer on observable, recoverable, cost-aware, and secure ingestion and medallion patterns ⢠Align Databricks configuration and usage with enterprise cloud standards across commercial and future government-hosted environments ⢠Enforce technical controls for data segregation, access boundaries, and operational compliance ⢠Track and improve platform health metrics, including job success rates, incident trends, pipeline reliability, cost efficiency, and environment drift ⢠Document platform standards, operational expectations, and support models ⢠Mentor internal engineers developing platform responsibilities ⢠Own the Databricks operational layer covering reliability, observability, deployment standards, compute and job policies, and platform enablement
⢠12+ years of relevant experience in data platform engineering, platform operations, site reliability engineering, or modern cloud data infrastructure ⢠Hands-on production experience with Databricks or a closely related cloud data platform ⢠Experience designing or operating CI/CD, environment promotion, version control, and deployment automation for data platforms and pipelines ⢠Strong understanding of observability, monitoring, alerting, incident management, and reliability engineering ⢠Experience with compute policy design, workload isolation, service principals, and secure production execution patterns on cloud data platforms ⢠Ability to work in regulated or security-sensitive environments with access control, auditability, and operational discipline ⢠Collaboration with cloud/infrastructure, security, data engineering, and analytics stakeholders ⢠Preferred: Databricks certification and/or expertise with Delta Lake, Unity Catalog, Workflows, and Databricks Asset Bundles ⢠Preferred: Infrastructure-as-code and platform automation experience ⢠Preferred: Experience supporting commercial and government or segregated environments ⢠Preferred: Experience in defense, aerospace, federal, or another regulated industry
⢠Bonus ⢠Benefits for full-time regular employees ⢠Equity ⢠Temporary benefits package applicable after 60 days of employment for temporary employees ⢠Benefits eligibility excluded for military fellows and part-time employees ⢠Equal employment opportunity and disability or special-needs accommodation
Apply Nowđ August 12
Site Reliability Architect designing observability and reliability platforms for Summitâs regulated-industry application hosting and cloud services. Improving resilience, automation, and incident response across teams.
đşđ¸ United States â Remote
đľ $136k - $175k / year
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor
đ August 12
Senior Site Reliability Engineer operating MeridianLinkâs serverless AWS platform. Managing production reliability, databases, backups, monitoring, incident response, and infrastructure automation.
đşđ¸ United States â Remote
đľ $104.1k - $140k / year
đ° $485M Post-IPO Debt on 2021-11
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor
đ August 11
Senior Backend DevOps Engineer operating AWS containerized microservices for the VAâs JLV clinical data viewer. Building CI/CD, observability, security, and disaster recovery capabilities.
đşđ¸ United States â Remote
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ August 11
Senior Network Reliability Engineer automating insurance company network platforms at Group 1001. Applying SRE, cloud, Kubernetes, security, and observability practices to improve reliability and reduce operational toil.
đşđ¸ United States â Remote
đľ $135k - $190k / year
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor
đ August 11
Reliability Engineer improving asset performance across Hexionâs North American manufacturing plants. Leading failure elimination, maintenance optimization, and cross-site reliability standardization.
đşđ¸ United States â Remote
â° Full Time
đ Senior
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)