Senior/Staff Platform Engineer

Job not on LinkedIn

🔥 14 hours ago

🌐 Canada, Brazil, +1 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

🏗️ Platform Engineer

👻 Ghost score 13%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Virtasant

Virtasant

51 - 200 employees

💼 Consulting

🏢 Enterprise

🤝 B2B

Consulting • Enterprise • B2B

Virtasant is a cloud services and optimization company that helps businesses migrate to, manage, and build on public cloud platforms. They combine a proprietary automation platform with managed services and FinOps expertise to reduce cloud costs (claiming average savings of over 50%), support Cloud FinOps programs, and deliver outcome-based engagements rather than hourly or seat-based billing. Virtasant also provides enterprise AI guidance and partners with major cloud providers (AWS, Google Cloud, Azure) to deliver migration, optimization, and 24/7 managed operations.

📋 Description

• Design, build, operate, and improve production Kubernetes platforms • Own cluster architecture, networking, workload isolation, resource management, security, upgrades, scaling, and reliability • Troubleshoot Kubernetes, containers, Linux, networking, and underlying infrastructure issues • Operate and improve large-scale, highly available infrastructure across cloud, hybrid, virtualised, and/or bare-metal environments • Write and maintain production tooling and automation using Go, Python, or Java • Build internal services, APIs, integrations, and operational tooling • Automate repetitive operational processes and reduce manual platform intervention • Own reliability and operational health of critical production infrastructure • Lead or significantly contribute to incident response and root-cause remediation • Define and improve SLOs, SLIs, alerting, and operational processes • Use logs, metrics, traces, profiling tools, and system-level diagnostics • Drive improvements in availability, performance, capacity, resilience, and operational readiness • Contribute to disaster recovery planning, testing, and continuous improvement • Build and maintain infrastructure as code using Terraform and related automation technologies • Build and improve CI/CD and deployment workflows • Participate in production cloud or infrastructure migrations, including dependency analysis, networking, cutover, rollback, and validation • Build and maintain monitoring, metrics, dashboards, alerting, logging, and distributed tracing • Work directly with customer and internal engineering teams to investigate problems and drive technical solutions • Communicate architecture, technical decisions, risks, trade-offs, and progress to technical stakeholders • Own complex infrastructure initiatives from problem definition through production operation • Contribute to architecture discussions, RFCs, design reviews, and technical direction • Mentor engineers and improve engineering and operational practices • Operate independently in ambiguous situations and take ownership when direction is unavailable

🎯 Requirements

• 10+ years of professional experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related fields; 10+ years is preferred for Staff-level candidates • Significant hands-on experience operating complex production infrastructure and distributed systems • Demonstrated experience building and operating production Kubernetes platforms, not only deploying applications onto existing clusters • Production programming experience with Go, Python, or Java • Strong experience with production reliability, incident response, troubleshooting, and operational ownership • Experience independently owning complex technical initiatives from an ambiguous starting point through production • Experience working directly with technical stakeholders or customers and communicating complex technical topics effectively • Degree in Computer Science, Engineering, or a related field, or equivalent practical experience • Deep understanding of production Kubernetes infrastructure, including cluster architecture, networking/CNI, NetworkPolicy, scheduling, resource management, nodes, security/RBAC, and cluster behaviour • Strong Linux fundamentals and hands-on production systems troubleshooting • Strong understanding of networking concepts including DNS, routing, load balancing, connectivity, and cloud/Kubernetes networking • Production experience with at least one major cloud platform: AWS, GCP, or Alicloud • Infrastructure as code at scale using Terraform or equivalent tooling • Configuration management and automation experience with technologies such as Ansible, Puppet, or similar • Strong production debugging and root-cause analysis skills across infrastructure and distributed systems • Observability experience using metrics, logs, traces, dashboards, and alerting platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent • Experience with CI/CD infrastructure and modern software delivery practices • Docker and container tooling as part of the production lifecycle • Understanding of high availability, capacity planning, disaster recovery, and production resilience • Exceptional written and verbal technical communication • Strong analytical, debugging, and problem-solving ability • High degree of ownership and ability to operate independently • Comfortable making technical decisions and driving work forward in ambiguous environments • Able to communicate effectively with customers, engineers, and technical leadership • Strong technical judgment and ability to explain trade-offs • Able to lead technically and influence others without formal people-management authority • Comfortable working within a distributed, highly technical team

🏖️ Benefits

• Fully remote work arrangement • Pacific Hours (8:00 AM – 5:00 PM PST) • On-call every 4–5 weeks

Apply Now

Similar Jobs

🔥 16 hours ago

Jane.app

201 - 500

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Developer Operations leader managing GitHub governance, security, and automation for Jane’s healthcare SaaS platform. Extending standardized tooling practices across Jira, Datadog, JFrog, and future SDLC systems.

🕒 September 12

EverCommerce

1001 - 5000

☁️ SaaS

🤝 B2B

📣 Marketing

Senior Software Engineer building cloud infrastructure and deployment systems for Invoice Simple, EverCommerce’s small-business invoicing platform. Improving reliability, observability, and production operations.

🇨🇦 Canada – Remote

💵 $120k - $150k / year

💰 Private equity on 2019-08

⏰ Full Time

🟠 Senior

🏗️ Platform Engineer

🕒 September 11

EverCommerce

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Software Engineer improving Invoice Simple’s CI/CD, cloud infrastructure, and reliability. Supporting EverCommerce’s software platform for home and field service businesses.

🇨🇦 Canada – Remote

💵 $110k - $130k / year

💰 Private Equity Round on 2019-07

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏗️ Platform Engineer

🕒 August 14

Faire

1001 - 5000

🤝 B2B

🛍️ eCommerce

🛒 Retail

Senior Staff ML Platform Engineer architecting Faire’s wholesale technology platform. Defining scalable Databricks, MLOps, and machine learning infrastructure for independent retailers.

🕒 August 7

Midnite

201 - 500

🎲 Gambling

👥 B2C

Senior Platform Engineer strengthening backend reliability, service boundaries, and regional scalability for Midnite’s sports betting and gaming platform. Supporting the Canada launch and future market expansion.