Senior Site Reliability Engineer

🔥 12 hours ago

🗽 New York – Remote

infoinfo

💵 $180k - $190k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Chalice AI

Chalice AI

11 - 50 employees

📣 Marketing

🤖 Artificial Intelligence

Marketing • Artificial Intelligence • Advertising

Chalice AI is a company that offers software solutions empowering businesses with AI-driven advertising tools. Their platform enables users to build and control algorithms trained on their own data, prioritizing transparency, compliance, and auditability. Chalice AI emphasizes the customization and real-time decision-making capabilities of their software, allowing brands to drive specific business outcomes. They focus on providing unique, brand-specific AI models and solutions for data scientists and advertisers, enhancing productivity and competitive advantage in the digital advertising space.

📋 Description

• Define and operate the architectural backbone of Chalice's AI platform • Report directly to the VP of Engineering and work cross-functionally with Engineering, Data Science, Machine Learning, and Product • Mentor and elevate engineers in infrastructure best practices, operational rigor, and architectural thinking • Design and operate scalable, event-driven, multi-tenant ML infrastructure • Support distributed ML training using Databricks, Ray, and Flyte on EKS • Deliver containerized products to external customers • Build and support AWS event-driven systems using EventBridge, MSK/Kafka, Kinesis, Lambda/Fargate, SQS/SNS, and Step Functions • Architect centralized state stores using DynamoDB, Redis, and Postgres • Establish standards for idempotency, replay safety, event schema governance, and traceability • Operate multi-cluster Kubernetes environments in production • Implement GitOps, progressive delivery, cluster-level security policies, and multi-tenant isolation • Develop Terraform module standards and reusable infrastructure primitives • Enforce GitHub guardrails and evolve CI/CD pipelines with infrastructure testing, policy validation, progressive deployment, and rollback • Standardize federated identity management, Azure SSO, API authentication, key management, and resource isolation • Define external model-serving architecture, including OCI packaging, secret injection, network isolation, telemetry, upgrades, and runtime contracts • Define SLIs, SLOs, error budgets, distributed tracing, golden signals, on-call structures, escalation paths, incident response, and postmortem processes • Own Databricks infrastructure, including SSO, authentication, provisioning, Terraform deployments, cluster policies, and Unity Catalog governance • Help build and scale the SRE function and establish organizational reliability standards

🎯 Requirements

• 8–10 years of experience designing and operating production infrastructure at scale • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience • Deep AWS architecture expertise across networking, SSO, compute, storage, and event-driven services • Experience managing Kubernetes control planes in production environments • Experience building or migrating event-driven systems at scale • Experience in ML-heavy or data-heavy environments • Experience replacing or eliminating legacy infrastructure components and simplifying system design • Experience owning production failures and leading postmortems resulting in systemic improvements • Strong architectural judgment and ability to challenge assumptions constructively • Certifications such as AWS, Databricks, or Kubernetes are a plus but not required

🏖️ Benefits

• Equity • Medical, dental, and vision insurance • 401(k) options • Unlimited PTO • 11 company holidays • Office-wide closure between Christmas Eve and New Year's • Office start-up stipend • In-office meal allowance • Professional development and advancement opportunities • Innovative, collaborative, inclusive workplace

Apply Now

Similar Jobs

🔥 15 hours ago

Clair

51 - 200

💳 Fintech

🤝 B2B

👥 HR Tech

DevOps Engineer automating secure, reliable cloud infrastructure for Clair’s embedded-finance and real-time payroll platform. Managing Terraform, monitoring, compliance, and AI/ML environments.

🔥 16 hours ago

Independence Pet Group

1001 - 5000

🛡️ Insurance

👥 B2C

🧘 Wellness

Senior DevOps Engineer building multi-region Azure application platforms for a pet health services company. Enabling safe deployments, reliable AKS workloads, and scalable cloud-native systems.

🔥 19 hours ago

Niche

201 - 500

📚 Education

🏪 Marketplace

🤝 Non-profit

Tech Lead Manager leading hands-on Site Reliability engineering and people management for Niche’s school-search platform. Improving infrastructure reliability, delivery systems, and developer self-service.

🔥 20 hours ago

NetImpact Strategies Inc.

201 - 500

💼 Consulting

🏥 Healthcare

🎖️ Defense

DevSecOps Engineer securing cloud infrastructure and applications for the Federal Government. Supporting Kubernetes, CI/CD, vulnerability remediation, and DoD RMF/ATO documentation.

🔥 21 hours ago

NetImpact Strategies Inc.

201 - 500

💼 Consulting

🏥 Healthcare

🎖️ Defense

DevSecOps Engineer securing cloud infrastructure, applications, and Kubernetes environments for Federal Government agencies. Supporting DoD RMF/ATO documentation and authorization packages.