Platform Architect, AI/ML Infrastructure, GCP-focused

Job not on LinkedIn

đŸ”„ 4 minutes ago

🌐 Brazil, Argentina, +3 more countries – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

🔙 Backend Engineer

đŸ‘» Ghost score 14%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Wizdaa

Wizdaa

11 - 50 employees

đŸ’Œ Consulting

📩 Logistics

📣 Marketing

Consulting ‱ Logistics ‱ Marketing

Wizdaa is a company that provides access to top-tier remote developers, specializing in helping startups build their dream development teams in the U. S. time zone. They offer a meticulous six-stage human and AI screening process to ensure access to the top 1% of engineering talent. Wizdaa's services include managing hiring processes, onboarding, payroll, benefits, and taxes, allowing startups to focus on core business matters. Known for competitive rates averaging $30/hour, Wizdaa emphasizes cultural fit, technical excellence, and English fluency among their developers. The company prides itself on delivering tailored, cost-effective solutions that maximize startups' runway and success. They also offer insights into leveraging AI tools and aligning remote teams with U. S. time zones to boost productivity.

📋 Description

‱ Build and operate model and inference serving infrastructure for real-time and batch inference across multiple tenants ‱ Manage latency, throughput, autoscaling, and reliability ‱ Own the ML deployment lifecycle, including model registry, versioning, promotion workflows, rollout strategies, and safe rollback ‱ Operate agentic and LLM workloads in production, including inference providers, gateways, quotas, throttling, guardrails, prompt/version management, and graceful degradation ‱ Build reproducible, automated training, evaluation, and deployment pipelines as code ‱ Extend infrastructure-as-code practices to ML systems using Terraform and multi-project design ‱ Operate GitOps for ML workloads and own ArgoCD configuration and promotion workflows ‱ Run ML and AI workloads on multi-tenant Kubernetes/GKE, managing GPU scheduling, workload placement, tenant isolation, and capacity ‱ Own ML reliability and observability, including inference SLOs, model/data drift detection, regression monitoring, alert quality, on-call ergonomics, and runbooks ‱ Drive ML cost efficiency through accelerator right-sizing, committed-use and Spot VM capacity management, and cost attribution ‱ Use agentic coding tools to scaffold environments, generate/review IaC and pipeline code, and accelerate automation ‱ Identify platform problems proactively and shape platform evolution

🎯 Requirements

‱ 5+ years in platform engineering, SRE, MLOps, or infrastructure, including meaningful time operating production systems at scale ‱ Hands-on experience deploying and operating ML or AI workloads in production ‱ Strong SRE/DevOps foundation, including ownership of reliability, SLOs, post-mortems, and measurable improvements ‱ Deep Terraform expertise, including complex state, reusable modules, multi-project configurations, and CI-driven plan/apply workflows ‱ Strong GitOps background with ArgoCD or Flux in production ‱ Deep Kubernetes knowledge and production cluster operations, including control-plane-level troubleshooting ‱ Production GKE experience is strongly preferred ‱ Strong GCP background, including VPC networking, Compute Engine, IAM, Cloud Storage, and multi-project/organization design ‱ Hands-on production BigQuery experience, including partitioning, clustering, query cost/performance tuning, and dataset-level IAM ‱ Familiarity with Dataflow, Pub/Sub, or Dataproc ‱ Hands-on experience building and operating CI/CD pipelines ‱ Understanding of differences between ML pipelines and standard application CI/CD ‱ Senior-level automation-first thinking ‱ Active use of agentic coding tools ‱ Strong communication skills ‱ Experience with GPU/accelerator scheduling and node lifecycle management is nice to have ‱ Experience operating LLM inference at scale is nice to have ‱ Experience with ML pipeline and orchestration tooling is nice to have ‱ Experience with model registries, feature stores, and experiment tracking is nice to have ‱ Familiarity with model and data drift monitoring and ML-specific observability is nice to have ‱ FinOps background is nice to have ‱ Familiarity with data infrastructure is nice to have ‱ Experience with multi-tenant infrastructure is nice to have ‱ Prior startup scaling experience is nice to have ‱ Bachelor's Degree or equivalent experience ‱ Degree in IT or Computer Science or equivalent experience ‱ English conversational skills

đŸ–ïž Benefits

‱ Payment in USD ‱ Remote work arrangement ‱ Working hours aligned with EST time zone

Apply Now

Similar Jobs

đŸ”„ 11 minutes ago

PrideLogic

11 - 50

Platform Architect scaling Wizdaa’s GCP AI/ML infrastructure from model pipelines to reliable production services. Operating Kubernetes, Terraform, GitOps, and LLM workloads across LATAM.

BigQuery

Cloud

Flux

Google Cloud Platform

Kubernetes

Terraform

đŸ”„ 24 minutes ago

SysMap Solutions

1001 - 5000

đŸ’Œ Consulting

📣 Marketing

đŸ€– Artificial Intelligence

Senior Backend Python developer building APIs and scalable solutions for Brazilian technology group. Delivering integrations, architecture, performance, automation, and AI-assisted development.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

AWS

Django

Docker

Flask

JavaScript

Kafka

Linux

MongoDB

MySQL

NGINX

Node.js

NoSQL

Postgres

Python

RabbitMQ

React

Redis

Terraform

đŸ”„ 1 hour ago

Stone & Company

1 - 10

đŸ’Œ Consulting

đŸ€ B2B

Senior Go Software Engineer developing critical systems for Stone, a Brazilian payments technology and financial services company. Building scalable, secure production features and collaborating through code reviews.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

AWS

Azure

Cloud

Kafka

MySQL

OpenStack

Postgres

RabbitMQ

Go

đŸ”„ 4 hours ago

WEX

5001 - 10000

đŸ„ Healthcare

📩 Logistics

✈ Travel

Mid .NET Developer building cloud-based employee benefits software for WEX’s employee benefits technology team. Developing and testing full-stack solutions with C#, ASP.NET, Angular, and Azure.

Angular

ASP.NET

Azure

Cloud

Docker

JavaScript

SQL

.NET

đŸ”„ 6 hours ago

Stefanini Brasil

10,000+ employees

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Desenvolvedor Java sĂȘnior criando soluçÔes cloud, APIs e microsserviços para a Stefanini. Atuando com Kubernetes, CI/CD, Kafka e evolução arquitetural.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

AWS

Azure

Cloud

Docker

Google Cloud Platform

Java

JavaScript

Jenkins

Kafka

Kubernetes

Node.js

OpenShift

Spring

Spring Boot

SpringBoot

SQL