
51 - 200 employees
🤖 Artificial Intelligence
🤝 B2B
Artificial Intelligence • B2B
Comet is a meta machine learning platform designed to help AI practitioners and teams build reliable machine learning models for real-world applications by streamlining and connecting the machine learning model lifecycle. By leveraging Comet, users can employ machine learning experiment tracking to track, compare, explain and reproduce their models. Backed by thousands of users and multiple Fortune 100 companies, Comet provides insights and data to build better, more accurate AI models while improving productivity, collaboration and visibility across teams.
🔥 8 minutes ago
🇺🇸 United States – Remote
đź’µ $150k - $200k / year
⏰ Full Time
🟡 Mid-level
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
đź‘» Ghost score 1%
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
🤖 Artificial Intelligence
🤝 B2B
Artificial Intelligence • B2B
Comet is a meta machine learning platform designed to help AI practitioners and teams build reliable machine learning models for real-world applications by streamlining and connecting the machine learning model lifecycle. By leveraging Comet, users can employ machine learning experiment tracking to track, compare, explain and reproduce their models. Backed by thousands of users and multiple Fortune 100 companies, Comet provides insights and data to build better, more accurate AI models while improving productivity, collaboration and visibility across teams.
• Work directly with customers to understand their infrastructure and deployment requirements • Deploy and maintain Comet across Kubernetes and cloud environments • Troubleshoot complex infrastructure and deployment issues across Kubernetes, networking, cloud services, databases, and application configuration • Own customer issues from investigation through resolution, collaborating with Engineering, Support, Customer Success, and other teams when needed • Develop and improve Helm charts, Terraform, deployment tooling, automation, and documentation • Participate in technical customer calls and explain complex topics clearly • Use AI tools for troubleshooting, development, automation, and investigation while validating results and maintaining engineering judgment
• Located in the East or Central USA only • Proven customer-facing technical experience is essential • Proven experience using AI tools in real engineering workflows, with the ability to critically evaluate and validate their output • Strong hands-on experience with Kubernetes, Helm, and containerized applications • Production experience with cloud infrastructure, preferably AWS • Experience with Infrastructure as Code, preferably Terraform • Strong Linux, networking, and infrastructure troubleshooting skills • Experience troubleshooting production systems using logs, metrics, and monitoring/observability tools • Excellent written and verbal communication skills • Enterprise or self-hosted software experience (nice to have) • Multi-cloud or on-premises deployment experience (nice to have) • CI/CD and GitOps experience (nice to have) • MLOps or AI infrastructure experience (nice to have) • Software development experience (nice to have) • Startup experience (nice to have)
• Competitive benefits package • Flexible working hours and remote work options • Opportunities for professional growth and development • A collaborative and innovative work environment • The chance to work with cutting-edge technologies and projects
Apply Now🔥 1 hour ago
Senior DevOps Engineer designing AWS/Azure infrastructure, IaC and CI/CD systems. Supporting Element 84’s cloud-based geospatial and Generative AI software projects.
🇺🇸 United States – Remote
đź’µ $145k - $180k / year
⏰ Full Time
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 2 hours ago
DevOps Engineer building secure, scalable AWS/Azure infrastructure and CI/CD pipelines for WorkWave’s web, mobile, and API applications. Automating deployments, monitoring, testing, and operational processes.
🇺🇸 United States – Remote
đź’µ $110k - $140k / year
⏰ Full Time
🟢 Junior
🟡 Mid-level
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🔥 3 hours ago
Site Reliability Engineer building reliable cloud infrastructure for Weave's business communications platform. Automating, scaling, and monitoring GCP and Kubernetes environments.
🔥 4 hours ago
Senior DevOps Engineer leading cloud-native platforms, automation, and CI/CD for Ascensus’s retirement savings technology. Driving secure, reliable delivery across enterprise engineering teams.
🇺🇸 United States – Remote
đź’µ $150k - $200k / year
đź’° Secondary Market on 2019-02
⏰ Full Time
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🔥 6 hours ago
Electrical engineer evaluating Fundrise properties and deploying modular GPU-compute infrastructure. Leading electrical diligence, upgrades, commissioning, and contractor coordination across U.S. sites.
🇺🇸 United States – Remote
đź’µ $200k - $230k / year
đź’° Corporate Round on 2021-08
⏰ Full Time
🟡 Mid-level
đźź Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor