Senior DevOps – Platform Reliability Engineer

🕒 May 8

🗽 New York – Remote

infoinfo

⏰ Full Time

🟠 Senior

🏗️ Platform Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 49%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Zingtree

Zingtree

11 - 50 employees

🏥 Healthcare

🛡️ Insurance

📦 Logistics

💰 $15M Series A on 2022-01

Healthcare • Insurance • Logistics

Zingtree is a company that empowers customer support through AI process automation, helping enterprises streamline and simplify complex support processes. It offers tools to create, manage, and automate support workflows, making it easier for customer service agents and customers to resolve issues efficiently. Zingtree integrates with various CRM systems and supports multiple industries including contact centers, healthcare, insurance, and home services. Through its dynamic workflows, Author Assist AI, and compliance automation, it helps businesses improve their customer experience with faster resolution times and enhanced compliance controls.

📋 Description

• Own and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments. • Automate infrastructure provisioning using Infrastructure as Code (IaC) tools such as Terraform and CloudFormation. • Operate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers. • Manage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls. • Operate the data and event tier, including Aurora MySQL, ElastiCache/Redis, S3, and MSK (Kafka), with responsibility for backups, point-in-time recovery (PITR), and multi-AZ disaster recovery aligned to defined RTO/RPO objectives. • Build and maintain Lambda workloads where event-driven or serverless architectures are the right fit. • Build observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking. • Strengthen our security and compliance posture for SOC 2 and HIPAA, including least-privilege IAM, SCPs, secrets management, SAST/DAST, dependency and container scanning, image signing, AWS Config, Security Hub, GuardDuty, Inspector, and evidence automation. • Drive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls. • Build and evolve our AI-native DevOps capabilities.

🎯 Requirements

• 5+ years of experience in DevOps, SRE, or Platform Engineering operating production systems on AWS. • Strong experience with CI/CD pipelines and tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI. • Hands-on experience operating production EKS environments, including autoscaling, ingress, secrets management, and cluster upgrades. • Strong AWS networking experience, including multi-account VPC design, subnets, routing, security groups, NACLs, Route 53, ACM, and load balancers. • Deep experience with Terraform and GitHub Actions, ideally using OIDC-based cloud authentication. • Experience with Aurora/RDS MySQL, Redis (ElastiCache), and S3, including backups, PITR, migrations, and lifecycle management. • Strong observability experience using Prometheus, Grafana, and OpenTelemetry. • Experience operating Argo CD at scale. • Experience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible. • Experience managing Cloudflare services including WAF, Bot Management, Rate Limiting, and Zero Trust / Access, along with CloudFront. • Experience operating Kafka/MSK at scale, including topics, consumer groups, and schema registries. • Experience with Lambda and event-driven architectures. • Comfortable working with Python, Bash, and Linux systems. • Strong understanding of security best practices across IAM, KMS, secrets management, networking, and software supply chain security. • Familiarity with vulnerability scanning and compliance tooling.

🏖️ Benefits

• Competitive compensation packages • Comprehensive health benefits: • 100% of employee premiums covered • 75%–80% of dependent premiums covered for most health, dental, and vision plans • 401(k) plans to support retirement planning (no employer matching currently) • Paid parental leave • Unlimited PTO • Flexible remote work from anywhere • Up to $200/month co-working reimbursement • Home office stipend: • Up to $500 for home office setup • $100/month for internet, phone, and related expenses

Apply Now

Similar Jobs

🕒 May 8

Alteryx

1001 - 5000

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Senior Developer Platform Engineer at Alteryx enhancing developer productivity with AI-enabled tools. Focusing on local dev, CI/CD, and release readiness with operational excellence.

🕒 May 6

Turquoise Health

51 - 200

🏥 Healthcare

💼 Consulting

☁️ SaaS

Platform Operations Engineer leading infrastructure operations at Turquoise Health. Responsible for reliability, observability, and platform scalability within a fully remote team.

🇺🇸 United States – Remote

💵 $153k - $170k / year

💰 $30M Series B - Turquoise Health on 2024-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏗️ Platform Engineer

🕒 May 6

TigerData (creators of TimescaleDB)

51 - 200

📦 Logistics

🏭 Manufacturing

💼 Consulting

Senior Platform Engineer maintaining Kubernetes-based infrastructure for Tiger Data cloud services. Collaborating with engineering teams to enhance platform stability and develop new features.

🕒 May 6

Liatrio

51 - 200

🏢 Enterprise

☁️ SaaS

💼 Consulting

Lead Platform Engineer designing and leading technical solutions for AI-native platforms at Liatrio. Overseeing CI/CD pipelines and cloud-native environments for enterprise clients in transformation.

🕒 May 5

Paramount

10,000+ employees

💼 Consulting

📣 Marketing

📱 Media

Senior MLOps Engineer designing agentic systems for Paramount, focusing on production runtimes and multi-agent orchestration. Leading performance management and system governance.