
11 - 50 employees
Founded 2024
đď¸ Construction
đ Manufacturing
đŚ Logistics
Construction ⢠Manufacturing ⢠Logistics
Nexxa. ai is an industrial AI company building agentic automation for industrial operations. Its platform learns from engineers and employees to take over complex, repetitive workflows, autonomously setting goals and execution criteria to increase productivity and resilience. Delivered as an integrated solution that works with existing technology stacks, Nexxa. ai positions itself as a SaaS-like enterprise tool that generates immediate ROI for operations teams.
đĽ 0 minutes ago
đ California â Remote
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 18%
Amazon Redshift
AWS
Azure
BigQuery
Cloud
Distributed Systems
Google Cloud Platform
Grafana
Jenkins
Kubernetes
Prometheus
Python
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
Founded 2024
đď¸ Construction
đ Manufacturing
đŚ Logistics
Construction ⢠Manufacturing ⢠Logistics
Nexxa. ai is an industrial AI company building agentic automation for industrial operations. Its platform learns from engineers and employees to take over complex, repetitive workflows, autonomously setting goals and execution criteria to increase productivity and resilience. Delivered as an integrated solution that works with existing technology stacks, Nexxa. ai positions itself as a SaaS-like enterprise tool that generates immediate ROI for operations teams.
⢠Own and evolve core infrastructure across compute, networking, storage, and deployment systems ⢠Design and operate CI/CD pipelines for AI, data, and product engineering teams ⢠Build and maintain infrastructure-as-code for reproducible, auditable cloud and on-prem/edge environments ⢠Architect and manage Kubernetes platforms for training, inference, and application workloads, including GPU scheduling and autoscaling ⢠Support data warehouses, lakehouse architectures, feature stores, embedding indices, retrieval pipelines, and model training, evaluation, and serving infrastructure ⢠Define and drive observability practices across distributed systems, including metrics, logging, tracing, and alerting ⢠Establish reliability practices including SLOs/SLIs, incident response, postmortems, and on-call rotations ⢠Design security and compliance for cloud infrastructure, secrets management, access control, and industrial or legacy-environment integrations ⢠Make tradeoffs across cost, latency, reliability, and developer velocity ⢠Collaborate with engineering leadership on infrastructure roadmap and platform strategy ⢠Mentor engineers and raise operational excellence across the organization ⢠Own ambiguous, high-stakes infrastructure problems end-to-end and design systems for future scale
⢠6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles ⢠Deep hands-on experience with AWS, GCP, or Azure at production scale ⢠Kubernetes production experience, including GPU workload scheduling ⢠Infrastructure-as-code tooling such as Terraform, Pulumi, or equivalent ⢠CI/CD systems such as GitHub Actions, GitLab CI, CircleCI, Jenkins, or ArgoCD ⢠Experience designing and operating observability stacks such as Prometheus, Grafana, Datadog, or OpenTelemetry ⢠Experience supporting ML/AI infrastructure is a strong plus ⢠Excellent scripting/programming skills in Python, Go, or Bash ⢠Proven ability to independently scope and lead infrastructure projects from design through production rollout ⢠Strong incident management instincts and ability to lead outage response and root-cause analysis ⢠Preferred: experience bridging cloud and edge/on-prem environments, especially industrial or manufacturing contexts ⢠Preferred: familiarity with Snowflake, BigQuery, Redshift, or Databricks ⢠Preferred: service mesh, zero-trust networking, or industrial compliance frameworks such as SOC 2 or IEC 62443 ⢠Preferred: building internal developer platforms or self-service infrastructure tooling ⢠Preferred: scaling infrastructure teams or setting technical direction at Staff level
⢠Comprehensive salary and equity package ⢠Significant opportunities for career development and advancement ⢠Innovative environment focused on transforming heavy industries through AI and automation ⢠Collaborative culture valuing innovation, discipline, and continuous improvement
Apply NowđĽ 17 hours ago
Staff SRE scaling Carrierâs cloud-based building automation SaaS platform. Improving reliability, observability, automation, and resilience across customer-critical services.
đşđ¸ United States â Remote
đľ $96k - $192k / year
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor
đ Yesterday
SRE Delivery Manager leading SRE and Delivery teams for Ninetyâs EOS business-management software. Improving AWS infrastructure, deployments, observability, security, and incident response.
đşđ¸ United States â Remote
đľ $200k - $220k / year
đ° $35M Series B - Ninety on 2023-11
â° Full Time
đ Senior
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đ Yesterday
Principal Deployment Engineer serving as senior engineering contact for sensitive customers. Leading IT deployments, testing, training, and customer support for AIS cyber and information security operations.
đşđ¸ United States â Remote
đľ $126k - $180k / year
đ° Venture Round on 2015-06
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đ 2 days ago
Principal SRE scaling Kubernetes infrastructure and Golang services for Blue Riverâs autonomous robotics platforms. Driving reliability, observability, security, and platform adoption across engineering teams.
đşđ¸ United States â Remote
đľ $174k - $305k / year
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đŚ H1B Visa Sponsor
đ 2 days ago
Staff Site Reliability Engineer operating AWS and Kubernetes infrastructure for Butterfly Networkâs medical ultrasound platform. Improving observability, reliability, and incident response across clinical workflows.
đşđ¸ United States â Remote
đľ $190k - $210k / year
đ° $75.6M Post-IPO Equity - Butterfly Network on 2025-01
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)