Platform Architect, AI/ML Infrastructure, GCP-focused

Job not on LinkedIn

đŸ”„ 0 minutes ago

🌐 Brazil, Argentina, +3 more countries – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

🔙 Backend Engineer

đŸ‘» Ghost score 16%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of PrideLogic

PrideLogic

11 - 50 employees

PrideLogic is a company that does not currently have detailed information available as their website is under construction. More details may be provided in the future once updates are made available on their site.

📋 Description

‱ Build and operate model and inference serving infrastructure for real-time and batch inference across multiple tenants ‱ Manage latency, throughput, autoscaling, and reliability ‱ Own the ML deployment lifecycle, including model registry, versioning, promotion workflows, canary/shadow/A-B rollouts, and rollback ‱ Operate agentic and LLM workloads in production, including providers, gateways, quotas, throttling, guardrails, prompt/version management, and graceful degradation ‱ Build reproducible training, evaluation, and deployment pipelines as code with lineage and reproducibility ‱ Extend Terraform infrastructure-as-code patterns to ML systems and multi-project designs ‱ Operate GitOps for ML workloads using ArgoCD configuration and promotion workflows ‱ Run ML and AI workloads on multi-tenant GKE, managing GPU scheduling, workload placement, tenant isolation, and cost-aware capacity ‱ Own ML reliability and observability, including inference SLOs, model/data drift detection, regression monitoring, alert quality, on-call ergonomics, and runbooks ‱ Drive ML cost efficiency through accelerator right-sizing, committed-use and Spot VM capacity, and tenant/workload cost attribution ‱ Use agentic coding tools to scaffold environments, generate/review IaC and pipeline code, and accelerate automation ‱ Identify platform problems proactively and shape platform evolution

🎯 Requirements

‱ 5+ years in platform engineering, SRE, MLOps, or infrastructure, including meaningful time operating production systems at scale ‱ Hands-on experience deploying and operating ML or AI workloads in production ‱ Strong SRE/DevOps foundation, including ownership of production reliability, SLOs, post-mortems, and measurable improvements ‱ Deep Terraform expertise, including complex state, reusable modules, multi-project configurations, and CI-driven plan/apply workflows ‱ Strong GitOps background with ArgoCD or Flux in production ‱ Deep Kubernetes knowledge, including production cluster operations, failure modes, and control-plane understanding ‱ Production GKE experience is strongly preferred ‱ Strong GCP background: VPC networking, Compute Engine, IAM, Cloud Storage, and multi-project/organization design ‱ Hands-on production experience with BigQuery, including partitioning, clustering, query cost/performance tuning, and dataset-level IAM ‱ Familiarity with Dataflow, Pub/Sub, or Dataproc ‱ Hands-on experience building and operating CI/CD pipelines, including understanding of ML pipeline differences ‱ Senior-level automation-first approach ‱ Active use of agentic coding tools ‱ Strong written and verbal communication ‱ Bachelor's degree or equivalent experience in IT or Computer Science indicated in application questions ‱ Ability to work EST time zone ‱ Preferred experience with GPU/accelerator scheduling, LLM inference at scale, ML orchestration, model registries, drift monitoring, FinOps, data infrastructure, multi-tenant infrastructure, and startup-to-enterprise scaling

đŸ–ïž Benefits

‱ Payment in USD ‱ Remote work arrangement in LATAM ‱ Working hours aligned with EST time zone

Apply Now

Similar Jobs

đŸ”„ 13 minutes ago

SysMap Solutions

1001 - 5000

đŸ’Œ Consulting

📣 Marketing

đŸ€– Artificial Intelligence

Senior Backend Python developer building APIs and scalable solutions for Brazilian technology group. Delivering integrations, architecture, performance, automation, and AI-assisted development.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

AWS

Django

Docker

Flask

JavaScript

Kafka

Linux

MongoDB

MySQL

NGINX

Node.js

NoSQL

Postgres

Python

RabbitMQ

React

Redis

Terraform

đŸ”„ 1 hour ago

Stone & Company

1 - 10

đŸ’Œ Consulting

đŸ€ B2B

Senior Go Software Engineer developing critical systems for Stone, a Brazilian payments technology and financial services company. Building scalable, secure production features and collaborating through code reviews.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

AWS

Azure

Cloud

Kafka

MySQL

OpenStack

Postgres

RabbitMQ

Go

đŸ”„ 3 hours ago

WEX

5001 - 10000

đŸ„ Healthcare

📩 Logistics

✈ Travel

Mid .NET Developer building cloud-based employee benefits software for WEX’s employee benefits technology team. Developing and testing full-stack solutions with C#, ASP.NET, Angular, and Azure.

Angular

ASP.NET

Azure

Cloud

Docker

JavaScript

SQL

.NET

đŸ”„ 6 hours ago

Stefanini Brasil

10,000+ employees

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Desenvolvedor Java sĂȘnior criando soluçÔes cloud, APIs e microsserviços para a Stefanini. Atuando com Kubernetes, CI/CD, Kafka e evolução arquitetural.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

AWS

Azure

Cloud

Docker

Google Cloud Platform

Java

JavaScript

Jenkins

Kafka

Kubernetes

Node.js

OpenShift

Spring

Spring Boot

SpringBoot

SQL

đŸ”„ 6 hours ago

Alterdata Software

1001 - 5000

đŸ’Œ Consulting

⚡ Productivity

Desenvolvedor C#/.NET evoluindo sistemas Prosoft da Alterdata, empresa de soluçÔes empresariais em software. Construção de APIs, correção de bugs e melhorias de arquitetura.

đŸ—ŁïžđŸ‡§đŸ‡·đŸ‡”đŸ‡č Portuguese Required

Angular

AWS

Azure

Docker

Kubernetes

MongoDB

NoSQL

Postgres

React

Redis

SQL

TypeScript

.NET