Senior Data Engineer

🕒 Julho 14

🇨🇴 Colômbia – Remoto

⏰ Tempo Integral

🟠 Sênior

🚰 Engenheiro de Dados

👻 Score fantasma 11%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🗣️🇪🇸 Espanhol obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of 8020REI

8020REI

11 - 50 funcionários

Fundada em 2017

💼 Consultoria

📦 Logística

🏠 Imobiliário

Consulting • Logistics • Real Estate

8020REI é uma empresa de investimentos imobiliários que utiliza tecnologia de ponta e soluções de dados para otimizar o processo de investimento. Ela se especializa em fornecer estratégias de dados e marketing impulsionadas por AI para ajudar investidores imobiliários a identificar propriedades com desconto e fortalecer seus esforços de outbound marketing. Com foco em inovação e eficiência, a 8020REI busca entregar ROI significativo para seus clientes ao usar predictive analytics e machine learning para simplificar a jornada de investimento.

Descrição

• Own Big Data pipelines end to end. • Design, build, and optimize PySpark ETL/ELT pipelines on Amazon EMR and AWS Glue that process nationwide, county-partitioned property data (First American, BuildZoom permits, market comps) on daily and monthly cadences. • Run our lakehouse. • Operate and evolve our Bronze → Silver → Gold data lake on S3 with Apache Hudi and the AWS Glue Data Catalog, queried through Athena, including schema contracts, partitioning strategy, compaction, and performance tuning. • Orchestrate and automate. • Build reliable orchestration with AWS Step Functions, EventBridge, and Lambda; make reruns, backfills, and failure recovery boring and documented. • Enforce data quality. • Implement and extend our Data QA Audit Standard: layer contracts, write-audit-publish gating, quarantine flows, drift monitoring, and actionable Slack alerting, so bad data never reaches a client list. • Operate production databases and APIs. • Manage Aurora PostgreSQL and DynamoDB workloads, and run data-serving APIs (API Gateway, SQS-backed async workers), such as our Address Resolution Service, to production SLAs with dashboards and runbooks. • Own cloud cost. • Monitor, report, and reduce the AWS data-platform bill (EMR cluster sizing, Glue/Lambda usage, S3 lifecycle, Athena scan costs) as a first-class engineering responsibility. • Ship infrastructure as code. • Define infrastructure with Terraform and CloudFormation, delivered through GitHub Actions CI/CD with tests, linting, and coverage gates, we run a disciplined PR, branch-policy, and code-review culture. • Partner with Data Science. • Build the feature pipelines, training datasets, and serving paths behind our ML scoring and valuation models (scikit-learn/XGBoost-family stack), and co-own the handoff contracts between DS and DE. • Document like a pro. • Maintain runbooks, architecture docs, and data dictionaries (Confluence) so any teammate can operate what you build.

🎯 Requisitos

• 4+ years of hands-on data engineering with large-scale distributed data systems and a track record of production ownership (not just development). • Advanced PySpark : performance tuning, partitioning strategy, and cost-aware cluster sizing on real workloads (EMR or equivalent). • Strong Python : clean, tested, production-grade code (we use pytest, ruff, mypy, and coverage gates in CI). • Deep AWS experience : EMR, Glue, Lambda, S3, Athena, Step Functions, EventBridge, IAM, and VPC networking; comfort operating (not just deploying to) these services. • Advanced SQL : complex analytical queries, query optimization, and data modeling on both a warehouse/lake engine (Athena/Presto) and PostgreSQL. • Lakehouse experience : hands-on production work with at least one open table format (Apache Hudi strongly preferred; Iceberg or Delta Lake also valued) and medallion-style architecture. • Data quality mindset : experience implementing validation, quality gates, monitoring, and incident response for production data. • Infrastructure as Code : Terraform and/or CloudFormation in a CI/CD workflow. • English and Spanish : professional working proficiency in both (B2+). • Bachelor’s degree in Computer Science, Systems Engineering, Data Engineering, or equivalent practical experience.

🏖️ Benefícios

• Competitive Base Compensation • Profit Share Bonus • Flex PTO (up to 26 days per year) • Home Office Upgrade Bonus • HMO Bonus • Full-time Remote Work • Opportunity for growth and team-building potential • Ongoing support and budget to develop new skills

Candidatar-se

Vagas Similares

🕒 Julho 10

Caseware

201 - 500

💸 Finanças

🏢 Corporativo

☁️ SaaS

Senior Software Developer focusing on AI Data Engineering at Caseware. Responsible for building scalable data ingestion and processing systems.

🇨🇴 Colômbia – Remoto

⏰ Tempo Integral

🟠 Sênior

🚰 Engenheiro de Dados

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 3

CI&T

5001 - 10000

💼 Consultoria

🏥 Saúde

📣 Marketing

Senior Data Architect enabling clients to modernize their data infrastructure with AWS solutions. Leading architecture design and team collaboration for cloud-native data platforms.

🇨🇴 Colômbia – Remoto

💰 $5.500.000 Venture Round em 2014-04

⏰ Tempo Integral

🟠 Sênior

🚰 Engenheiro de Dados

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 29

Truelogic Software

501 - 1000

☁️ SaaS

🤝 B2B

🏢 Corporativo

Senior Data Engineer developing and maintaining a data platform for veterinary technology. Collaborating remotely with a skilled engineering team to ensure data insights and reliability.

🇨🇴 Colômbia – Remoto

⏰ Tempo Integral

🟠 Sênior

🚰 Engenheiro de Dados

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 17

BlueCloud

501 - 1000

💼 Consultoria

📣 Marketing

📦 Logística

Senior Data Engineer designing and implementing Snowflake data solutions for enterprise organizations. Leading the architecture of scalable data pipelines and enforcing data governance standards.

🇨🇴 Colômbia – Remoto

⏰ Tempo Integral

🟠 Sênior

🚰 Engenheiro de Dados

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 12

Blend360

501 - 1000

🏥 Saúde

🏨 Hospitalidade

✈️ Turismo

Senior Data Engineer at Blend responsible for designing and optimizing data pipelines. Collaborating with teams to ensure data accuracy and quality for enterprise initiatives.

🇨🇴 Colômbia – Remoto

💰 $100.000.000 Private Equity Round em 2022-08

⏰ Tempo Integral

🟠 Sênior

🚰 Engenheiro de Dados

🗣️🇺🇸🇬🇧 Inglês obrigatório