Search Remote Jobs

Senior Staff Platform, Data Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Shield AI

Shield AI

501 - 1000 employees

Founded 2015

🤖 Artificial Intelligence

🚀 Aerospace

🎖️ Defense

Artificial Intelligence • Aerospace • Defense

Shield AI is a leading developer of AI-driven military solutions, focusing on enhancing mission autonomy and battlefield awareness. Their platform, Hivemind, enables rapid deployment of intelligent systems for various defense applications, including drone operation and surveillance. With a commitment to utilizing advanced technology, Shield AI aims to protect service members and civilians by revolutionizing defense technologies through autonomous systems.

📋 Description

• Own operational excellence for the Databricks platform, including monitoring, alerting, observability, incident response support, and production runbook patterns • Define and maintain CI/CD and promotion standards for Databricks assets and environments from development to production • Design and maintain standards for job orchestration, cluster and compute policies, service principal usage, environment isolation, and reliable production execution • Establish reusable operational templates and enablement patterns for onboarding new domains to Databricks • Partner with the Senior Data Engineer on observable, recoverable, cost-aware, and secure ingestion and medallion patterns • Align Databricks configuration and usage with enterprise cloud standards across commercial and future government-hosted environments • Enforce technical controls for data segregation, access boundaries, and operational compliance • Track and improve platform health metrics, including job success rates, incident trends, pipeline reliability, cost efficiency, and environment drift • Document platform standards, operational expectations, and support models • Mentor internal engineers developing platform responsibilities • Own the Databricks operational layer covering reliability, observability, deployment standards, compute and job policies, and platform enablement

🎯 Requirements

• 12+ years of relevant experience in data platform engineering, platform operations, site reliability engineering, or modern cloud data infrastructure • Hands-on production experience with Databricks or a closely related cloud data platform • Experience designing or operating CI/CD, environment promotion, version control, and deployment automation for data platforms and pipelines • Strong understanding of observability, monitoring, alerting, incident management, and reliability engineering • Experience with compute policy design, workload isolation, service principals, and secure production execution patterns on cloud data platforms • Ability to work in regulated or security-sensitive environments with access control, auditability, and operational discipline • Collaboration with cloud/infrastructure, security, data engineering, and analytics stakeholders • Preferred: Databricks certification and/or expertise with Delta Lake, Unity Catalog, Workflows, and Databricks Asset Bundles • Preferred: Infrastructure-as-code and platform automation experience • Preferred: Experience supporting commercial and government or segregated environments • Preferred: Experience in defense, aerospace, federal, or another regulated industry

🏖️ Benefits

• Bonus • Benefits for full-time regular employees • Equity • Temporary benefits package applicable after 60 days of employment for temporary employees • Benefits eligibility excluded for military fellows and part-time employees • Equal employment opportunity and disability or special-needs accommodation

Apply Now

Similar Jobs

🔥 4 hours ago

Sonatype

501 - 1000

🔒 Cybersecurity

☁️ SaaS

GCP DevOps Engineer designing secure infrastructure, CI/CD, and Kubernetes platforms for Sonatype. Improving developer delivery, reliability, observability, and software supply-chain security.

🔥 4 hours ago

Summit

51 - 200

💼 Consulting

🏨 Hospitality

📣 Marketing

Site Reliability Architect designing observability and reliability platforms for Summit’s regulated-industry application hosting and cloud services. Improving resilience, automation, and incident response across teams.

🔥 5 hours ago

MeridianLink

501 - 1000

💳 Fintech

🏦 Banking

☁️ SaaS

Senior Site Reliability Engineer operating MeridianLink’s serverless AWS platform. Managing production reliability, databases, backups, monitoring, incident response, and infrastructure automation.

🔥 6 hours ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Senior SRE improving NVIDIA GeForce NOW’s reliable GPU cloud gaming infrastructure. Building observability, automation, Kubernetes, and incident-response tooling for service SLOs.

🔥 7 hours ago

SimpliGov

11 - 50

🏛️ Government

☁️ SaaS

⚡ Productivity

Senior DevOps/MLOps Engineer operating SimpliGov’s Azure AI platform for government customers. Building secure Kubernetes infrastructure, compliant inference paths, observability, releases, and cost controls.