Senior SRE

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CloudFactory

CloudFactory

1001 - 5000 employees

Founded 2010

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

CloudFactory is a company that provides an AI Data Platform to accelerate the journey from AI concepts to fully operational solutions. Their platform enables the creation and preparation of high-quality, structured datasets aimed at improving model accuracy and speeding up deployment. CloudFactory's services include GenAI for enhancing foundational LLM model performance, model oversight for auditing and managing AI models, and professional services guiding projects through value proof and MVP stages to full production. The company places a strong emphasis on inference optimization and the synergy of human and machine intelligence to ensure trustworthy and reliable AI model deployment, focusing on delivering real-world business results. They also prioritize data security and compliance with industry standards like ISO 9001:2015, ISO 27001, SOC 2, HIPAA, and GDPR.

📋 Description

• Design and implement new core infrastructure components with a high degree of autonomy • Optimize and improve deployment pipelines, environment provisioning, and high-throughput batch jobs • Use Infrastructure as Code tools such as Terraform to manage and scale complex infrastructure • Develop CI/CD pipelines for automated build, test, deployment, and monitoring processes • Create and manage multi-step CI/CD pipelines, including environment setup and artifact handling • Support the reliability, availability, and performance of production systems • Apply site-reliability practices across infrastructure • Set up monitoring, alerting, and observability tooling • Collaborate with software engineers, product, and business stakeholders on infrastructure and deployment systems • Communicate complex technical issues clearly to technical and non-technical stakeholders

🎯 Requirements

• 5+ years of experience building and operating infrastructure in production environments • Fluent in Python with strong experience writing production-ready code • Experience with Docker and Kubernetes • Knowledge of cloud platforms such as GCP or AWS • Experience using Infrastructure as Code tools such as Terraform • Experience using CI/CD platforms to automate build, test, and deployment pipelines • Comfortable applying site-reliability principles including availability, observability, and automation • Degree in Computer Science, Engineering, or another quantitative or computational field, or equivalent practical experience • Familiarity with Prometheus or Grafana preferred • Experience with configuration management tools such as Ansible, Chef, or Puppet preferred • Exposure to multi-cloud or hybrid-cloud environments preferred

🏖️ Benefits

• Great Mission and Culture • Meaningful Work • Market competitive salary • Quarterly variable compensation • Hybrid Working Model • Comprehensive medical cover • Group life insurance • Personal development and growth opportunities

Apply Now

Similar Jobs

🔥 3 hours ago

InnoData

2 - 10

🤝 B2B

💼 Consulting

🌍 Social Impact

Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.

🕒 2 days ago

InnoData

2 - 10

🤝 B2B

💼 Consulting

🌍 Social Impact

Lead Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and delivering enhancements across microservices environments.

🕒 3 days ago

WinAir

51 - 200

📦 Logistics

💼 Consulting

🏭 Manufacturing

DevOps Specialist automating CI/CD and infrastructure for WinAir’s aviation maintenance software. Improving Jenkins, Ansible, Linux environments, deployments, and monitoring across development and production systems.

🕒 3 days ago

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Senior SRE maintaining Kubernetes-based UI and AI service reliability for an enterprise cloud software company. Managing incidents, observability, deployments, and runtime troubleshooting in production.

🕒 4 days ago

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Senior SRE supporting Kubernetes production reliability for an enterprise cloud software company. Troubleshooting distributed services, observability, incidents, CI/CD, Node.js, and JVM/Java runtimes.