
11 - 50 employees
📣 Marketing
🤖 Artificial Intelligence
Marketing • Artificial Intelligence • Advertising
Chalice AI is a company that offers software solutions empowering businesses with AI-driven advertising tools. Their platform enables users to build and control algorithms trained on their own data, prioritizing transparency, compliance, and auditability. Chalice AI emphasizes the customization and real-time decision-making capabilities of their software, allowing brands to drive specific business outcomes. They focus on providing unique, brand-specific AI models and solutions for data scientists and advertisers, enhancing productivity and competitive advantage in the digital advertising space.
🔥 12 hours ago
🗽 New York – Remote
💵 $180k - $190k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
👻 Ghost score 0%
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
📣 Marketing
🤖 Artificial Intelligence
Marketing • Artificial Intelligence • Advertising
Chalice AI is a company that offers software solutions empowering businesses with AI-driven advertising tools. Their platform enables users to build and control algorithms trained on their own data, prioritizing transparency, compliance, and auditability. Chalice AI emphasizes the customization and real-time decision-making capabilities of their software, allowing brands to drive specific business outcomes. They focus on providing unique, brand-specific AI models and solutions for data scientists and advertisers, enhancing productivity and competitive advantage in the digital advertising space.
• Define and operate the architectural backbone of Chalice's AI platform • Report directly to the VP of Engineering and work cross-functionally with Engineering, Data Science, Machine Learning, and Product • Mentor and elevate engineers in infrastructure best practices, operational rigor, and architectural thinking • Design and operate scalable, event-driven, multi-tenant ML infrastructure • Support distributed ML training using Databricks, Ray, and Flyte on EKS • Deliver containerized products to external customers • Build and support AWS event-driven systems using EventBridge, MSK/Kafka, Kinesis, Lambda/Fargate, SQS/SNS, and Step Functions • Architect centralized state stores using DynamoDB, Redis, and Postgres • Establish standards for idempotency, replay safety, event schema governance, and traceability • Operate multi-cluster Kubernetes environments in production • Implement GitOps, progressive delivery, cluster-level security policies, and multi-tenant isolation • Develop Terraform module standards and reusable infrastructure primitives • Enforce GitHub guardrails and evolve CI/CD pipelines with infrastructure testing, policy validation, progressive deployment, and rollback • Standardize federated identity management, Azure SSO, API authentication, key management, and resource isolation • Define external model-serving architecture, including OCI packaging, secret injection, network isolation, telemetry, upgrades, and runtime contracts • Define SLIs, SLOs, error budgets, distributed tracing, golden signals, on-call structures, escalation paths, incident response, and postmortem processes • Own Databricks infrastructure, including SSO, authentication, provisioning, Terraform deployments, cluster policies, and Unity Catalog governance • Help build and scale the SRE function and establish organizational reliability standards
• 8–10 years of experience designing and operating production infrastructure at scale • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience • Deep AWS architecture expertise across networking, SSO, compute, storage, and event-driven services • Experience managing Kubernetes control planes in production environments • Experience building or migrating event-driven systems at scale • Experience in ML-heavy or data-heavy environments • Experience replacing or eliminating legacy infrastructure components and simplifying system design • Experience owning production failures and leading postmortems resulting in systemic improvements • Strong architectural judgment and ability to challenge assumptions constructively • Certifications such as AWS, Databricks, or Kubernetes are a plus but not required
• Equity • Medical, dental, and vision insurance • 401(k) options • Unlimited PTO • 11 company holidays • Office-wide closure between Christmas Eve and New Year's • Office start-up stipend • In-office meal allowance • Professional development and advancement opportunities • Innovative, collaborative, inclusive workplace
Apply Now🔥 15 hours ago
DevOps Engineer automating secure, reliable cloud infrastructure for Clair’s embedded-finance and real-time payroll platform. Managing Terraform, monitoring, compliance, and AI/ML environments.
🇺🇸 United States – Remote
💵 $138k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 16 hours ago
Senior DevOps Engineer building multi-region Azure application platforms for a pet health services company. Enabling safe deployments, reliable AKS workloads, and scalable cloud-native systems.
🔥 19 hours ago
Tech Lead Manager leading hands-on Site Reliability engineering and people management for Niche’s school-search platform. Improving infrastructure reliability, delivery systems, and developer self-service.
🔥 20 hours ago
DevSecOps Engineer securing cloud infrastructure and applications for the Federal Government. Supporting Kubernetes, CI/CD, vulnerability remediation, and DoD RMF/ATO documentation.
🇺🇸 United States – Remote
💵 $120k - $165k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 21 hours ago
DevSecOps Engineer securing cloud infrastructure, applications, and Kubernetes environments for Federal Government agencies. Supporting DoD RMF/ATO documentation and authorization packages.
🇺🇸 United States – Remote
💵 $120k - $165k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)