Search Remote Jobs

Senior DevOps Engineer

🔥 0 minutes ago

🏖️ New Jersey – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akkadian Labs

Akkadian Labs

51 - 200 employees

☁️ SaaS

🏢 Enterprise

📡 Telecommunications

SaaS • Enterprise • Telecommunications

Akkadian Labs is a leading provider of automated provisioning solutions for Unified Communications (UC) environments. They offer a suite of products including Akkadian Provisioning Manager, Console, Contact Manager, and Site Builder, designed to streamline UC adoption, deployment, and management. Their solutions facilitate error-free provisioning across platforms like Webex, Microsoft, and Zoom, helping manage millions of enterprise users efficiently. Akkadian Labs works closely with partners and employs automation to reduce repetitive manual work, enhance security, and achieve quick ROI. They are recognized for their partnership with UC providers and support through an intuitive console that caters to high call volume environments.

📋 Description

• Design, implement, and maintain scalable and secure infrastructure and DevOps processes • Deploy and maintain scalable infrastructure in AWS and hybrid cloud environments • Manage infrastructure-as-code using Terraform, CloudFormation, or similar tools • Maintain Linux-based environments • Design and implement Docker containerization and Kubernetes orchestration • Design, deploy, and manage AI agent workloads, including compute provisioning and resource scaling for inference-heavy tasks • Build and maintain AI model deployment pipelines, including versioning, testing, and production rollback • Monitor AI API consumption and infrastructure costs; implement alerting and usage controls • Implement infrastructure-level security guardrails for AI systems, including access controls and data isolation • Manage monitoring and observability using Prometheus, Grafana, and the ELK stack • Troubleshoot system issues and contribute to incident response and root cause analysis • Improve system reliability, performance, and uptime • Build, maintain, and optimize CI/CD pipelines using Jenkins, Bitbucket CI/CD, or similar tools • Automate builds, testing, deployments, and system updates • Integrate pipelines with Akkadian tools • Implement secure DevOps practices, security controls, compliance initiatives, and vulnerability remediation • Collaborate with DevOps, engineering, QA, and product teams on deployments and releases • Maintain infrastructure, process, and operational documentation • Participate in collaborative team processes and continuous improvement initiatives

🎯 Requirements

• 10+ years of experience in DevOps or Site Reliability Engineering (SRE) • Expertise with AWS, including EC2, ECS, S3, IAM, Lambda, and CloudWatch • Expertise with infrastructure-as-code tools, including Terraform and CloudFormation • Strong knowledge of Linux environments • Experience with Docker and Kubernetes • Scripting ability in Python, Bash, or similar languages • Experience building or maintaining CI/CD pipelines and related tools • Experience with Prometheus, Grafana, and ELK for monitoring and observability • Experience implementing secure DevOps practices and compliance frameworks such as SOC2 and ISO • Experience supporting AI or machine learning workloads and compute environments • Exposure to AI model deployment pipelines and model versioning practices • Familiarity with hybrid cloud or on-premises environments • Exposure to DevOps security best practices, including AI-specific data isolation and access controls • Experience supporting production systems and participating in on-call rotations

🏖️ Benefits

• Fully remote environment • Medical insurance • Dental insurance • Vision insurance • Company-paid life insurance • Company-paid disability policies • 401(k) with a generous matching program • Paid time off

Apply Now

Similar Jobs

🔥 1 hour ago

Casper Studios

2 - 10

🤖 Artificial Intelligence

🏢 Enterprise

📱 Media

AI Security & DevOps Engineer securing AI systems from prototype to production for Casper Studios, an AI services firm. Owning security standards, platform controls, and enterprise client approvals.

🔥 4 hours ago

XBOX

10,000+ employees

🎮 Gaming

🔧 Hardware

Senior SRE improving Blizzard’s large-scale data, analytics, and ML platforms. Building automation, Kubernetes infrastructure, and reliability tooling for gaming services.

🔥 4 hours ago

Trumid

51 - 200

💳 Fintech

💸 Finance

☁️ SaaS

Senior DBRE safeguarding Postgres durability, recovery, and performance for Trumid’s fixed-income trading platform. Owning failover, observability, backups, and data resilience.

🔥 8 hours ago

Replicant

51 - 200

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Senior SRE building CI/CD, observability, and Kubernetes infrastructure for Replicant’s AI-powered customer-service platform. Improving reliability and developer tooling for large-scale conversational traffic.

🔥 9 hours ago

Ancestry

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

Site Reliability Engineer maintaining Ancestry’s family-history platform availability through incident response, AWS operations, and automation. Developing AI agentic workflows and tooling for a 24/7 Command Center.