Senior Software Engineer, Data Processing

Job not on LinkedIn

🕒 June 3

🇧🇷 Brazil – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 23%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Grupo Protege

Grupo Protege

10,000+ employees

Founded 1971

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Artificial Intelligence • B2B • SaaS

Grupo Protege is an AI training data platform that connects AI developers with high-quality, ethically sourced training data. It serves both AI developers by providing a vast and rich collection of data for model training and data holders by enabling them to monetize their data while maintaining governance and control. The platform aims to streamline the data procurement process significantly, making it easier for developers to access the data they need efficiently.

📋 Description

• Design, build, and operate the ingestion systems that process large volumes of multimodal data into usable, well-structured datasets • Own the ingestion path end to end, from how data lands to how it is validated, processed, tracked, and made available downstream • Build modality-specific processing steps for real-world source data, such as medical imaging processing, audio and video metadata extraction, quality validation, and notes processing • Build parsers, validators, and normalization logic that can systematically handle messy, non-standard, and high-variance source formats • Turn repeated one-off data handling work into reusable processing patterns, internal tooling, and platform capabilities • Build for high volume and high throughput, optimizing systems for reliability, cost, and speed • Work across distributed and parallel compute systems to process workloads that do not fit well on a single machine • Choose the right execution model for the workload, including batch processing, distributed execution, and modern compute patterns for unstructured data and inference-heavy processing • Diagnose and resolve bottlenecks across ingestion and processing systems, and keep performance from degrading as volume and modality complexity grow • Build validation and quality checks that catch bad, incomplete, or malformed data before it propagates downstream • Handle sensitive and regulated data, including PHI, with the security and care the domain demands, including de-identification where required • Track provenance, metadata, and usage constraints through the ingestion path so downstream use remains compliant and auditable • Raise the quality bar for observability, debuggability, and operational reliability across the ingestion layer • Partner with product and Data Lab to support new modalities, new partner requirements, and non-standard source data • Work directly with partner engineering teams when needed to translate source-system realities into robust ingestion and processing design • Surface recurring patterns that are worth standardizing into reusable transforms, validators, and internal tooling • Help shape how Protege handles new data types as the platform expands into more complex data environments

🎯 Requirements

• 5+ years building and operating production backend or data systems, with real experience in data processing at scale • Hands-on experience designing and running large-scale data pipelines • Strong programming skills in Python • Experience with distributed data processing • Strong proficiency with AWS • Comfort with messy, varied, high-volume data and high ambiguity, with a knack for finding patterns in complex environments • Attention to detail without losing speed, and a bias to action • Excited to work on a product built around moving and processing large volumes of data • Curious, tenacious, and proactive

🏖️ Benefits

• Health insurance • Professional development opportunities • Flexible working hours

Apply Now

Similar Jobs

🕒 June 1

Chainlink Labs

201 - 500

💸 Finance

💳 Fintech

🌐 Web 3

Software Engineer in Data Growth at Chainlink building automation and self-service experiences for web3 customers to improve operational efficiency and scaling processes.

TypeScript

Web3

🕒 May 26

CI&T

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Data Architect uniting human expertise with AI to create scalable tech solutions for clients globally. Managing data lifecycle, architecture, and leading cross-functional teams.

🇧🇷 Brazil – Remote

💰 $5.5M Venture Round on 2014-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Azure

Cloud

ETL

Google Cloud Platform

🕒 May 22

Arco Educação

1001 - 5000

📚 Education

🤝 B2B

Senior Data Platform Engineer designing and implementing solutions for Arco Educação's education technology platform. Collaborating with teams to optimize data model and performance in a complex AI-driven environment.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

BigQuery

Cloud

Docker

Google Cloud Platform

Java

Kubernetes

Python

Scala

Spark

SQL

Tableau

Go

🕒 May 15

Review ALL

11 - 50

💼 Consulting

🎯 Recruiter

🤝 B2B

Senior Data Engineer responsible for optimizing data pipelines and maintaining data quality. Collaborating on data strategies and implementing analytical models for business needs.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

ETL

Grafana

Kafka

Kubernetes

Prometheus

Python

SQL

Terraform

Unity

🕒 May 14

Jusbrasil

201 - 500

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior Data Engineer at Jusbrasil working on data pipelines and supporting data-driven decisions. Collaborating with cross-functional teams to ensure data quality and performance.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

BigQuery

Cloud

Google Cloud Platform

Java

Kafka

Python

Scala

SQL

Terraform

Vault