
11 - 50 employees
Founded 2023
☁️ SaaS
🤖 Artificial Intelligence
🤝 B2B
💰 Seed on 2024-06
SaaS • Artificial Intelligence • B2B
OpsMill is a company that builds Infrahub, an AI-ready infrastructure data management platform and unified source of truth for network, data center, and cloud automation. Infrahub models infrastructure with flexible, user-defined schemas (graph-based), version-control and GitOps workflows, and native integrations with tools like Ansible and Terraform to enable configuration management, IPAM/resource management, self-service catalogs, and AIOps-driven reasoning and automation. The platform targets network and infrastructure teams, offering enterprise features, open-source/community options, and capabilities for validation, auditing, and automated deployments.
🔥 1 hour ago
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
Founded 2023
☁️ SaaS
🤖 Artificial Intelligence
🤝 B2B
💰 Seed on 2024-06
SaaS • Artificial Intelligence • B2B
OpsMill is a company that builds Infrahub, an AI-ready infrastructure data management platform and unified source of truth for network, data center, and cloud automation. Infrahub models infrastructure with flexible, user-defined schemas (graph-based), version-control and GitOps workflows, and native integrations with tools like Ansible and Terraform to enable configuration management, IPAM/resource management, self-service catalogs, and AIOps-driven reasoning and automation. The platform targets network and infrastructure teams, offering enterprise features, open-source/community options, and capabilities for validation, auditing, and automated deployments.
• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents
• 4-7 years of experience in production engineering, SRE, platform engineering, or similar roles where you've owned reliability and customer escalations • Strong software engineering fundamentals including design, debugging, testing, code review, and a focus on maintainable, production-quality code • Practical Kubernetes expertise sufficient to debug real deployments: troubleshooting resources, networking, storage, RBAC, and platform-specific quirks across different distributions • Deep troubleshooting instincts and observability experience using logs, metrics, and traces to diagnose issues quickly in complex, distributed systems • Experience with at least one of: Python, Go, or Rust for building tooling and contributing to product code (you don't need to be expert in all three) • Excellent problem decomposition and communication skills—you can break down messy, ambiguous issues and clearly explain your findings and recommendations • Self-directed remote work capability with strong async communication skills and the ability to operate independently in a fast-moving environment where priorities shift based on customer needs • Collaborative mindset with experience partnering across product, engineering, and customer-facing teams to drive systematic improvements
• The people: Work alongside world-class engineers who've built and scaled automation platforms in production. Daily technical challenges with smart colleagues who push you to grow. • The product: Shape Infrahub based on real customer needs. Your input directly influences features, integrations, and roadmap priorities. • The mission: We're making enterprise-grade infrastructure automation accessible to any organization. Open-source at the core, production-ready out of the box. This is a multi-year journey, not a quarterly sprint. • The impact: You'll work with teams managing some of the world's most complex infrastructure deployments, solving problems that ripple across entire organizations. • Our Commitment to Diversity and Inclusion: OpsMill is committed to building a diverse and inclusive team. We believe different perspectives make us stronger and more innovative. We encourage applications from candidates of all backgrounds and experiences, and we're committed to providing an inclusive environment where everyone can do their best work.
Apply Now🔥 1 hour ago
Senior Site Reliability Engineer designing and maintaining reliable infrastructure at CertifyOS for healthcare data systems. Influencing architecture and ensuring uptime for millions of provider records.
🇺🇸 United States – Remote
💰 $40M Series B - Certify on 2025-06
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Cloud
Distributed Systems
Google Cloud Platform
Grafana
Java
Linux
Microservices
Node.js
Prometheus
Python
React
Terraform
TypeScript
Go
🔥 1 hour ago
Site Reliability Engineer responsible for improving production operations and data pipelines at Voleon. Collaborating with software engineers to enhance reliability and efficiency across critical systems.
🇺🇸 United States – Remote
💵 $120k - $160k / year
⏰ Full Time
🟢 Junior
🟡 Mid-level
⛑ DevOps & Site Reliability Engineer (SRE)
Linux
Python
SQL
🔥 1 hour ago
Senior Cluster Site Reliability Engineer at Voleon, ensuring robustness and uptime of research compute clusters. Collaborating on systemic improvements and cluster health monitoring.
🇺🇸 United States – Remote
💵 $205k - $235k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Ansible
AWS
Cloud
Google Cloud Platform
Grafana
Prometheus
Python
Ruby
Terraform
🔥 1 hour ago
Site Reliability Engineer managing reliability and observability for GPU One platform. Leading incident response and ensuring SLOs for customer SLAs.
Grafana
Kubernetes
Prometheus
Python
Go
🔥 2 hours ago
Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.
AWS
Kubernetes
Python
Terraform