Senior Data Engineer

🔥 14 hours ago

🇨🇴 Colombia – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 10%

infoinfo

🗣️🇪🇸 Spanish Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 8020REI

8020REI

11 - 50 employees

Founded 2017

💼 Consulting

📦 Logistics

🏠 Real Estate

Consulting • Logistics • Real Estate

8020REI is a company that specializes in providing AI-driven predictive data and advanced marketing strategies for real estate investors. Their services include direct mail, SMS, and cold calling to generate more motivated seller leads, allowing clients to maximize their return on investment. By leveraging machine learning, predictive analytics, and a proven marketing cadence plan, 8020REI helps clients save money and focus on the most promising properties likely to sell at a discount. Their robust data services and marketing insights support transparency, efficiency, and effective targeting in real estate investment.

📋 Description

• Design, build, and optimize PySpark ETL/ELT pipelines on Amazon EMR and AWS Glue • Process nationwide, county-partitioned property data on daily and monthly cadences • Operate and evolve the Bronze → Silver → Gold data lake on S3 with Apache Hudi and AWS Glue Data Catalog • Manage schema contracts, partitioning, compaction, and performance tuning • Build orchestration with AWS Step Functions, EventBridge, and Lambda • Implement reruns, backfills, failure recovery, and documentation • Enforce data quality through contracts, write-audit-publish gating, quarantine flows, drift monitoring, and Slack alerting • Manage Aurora PostgreSQL and DynamoDB workloads • Run data-serving APIs and async workers to production SLAs with dashboards and runbooks • Monitor, report, and reduce AWS data-platform costs • Define infrastructure with Terraform and CloudFormation through GitHub Actions CI/CD • Build feature pipelines, training datasets, and serving paths for ML scoring and valuation models • Maintain runbooks, architecture documentation, and Confluence data dictionaries • Make architecture decisions and partner directly with Data Science

🎯 Requirements

• 4+ years of hands-on data engineering with large-scale distributed data systems and a track record of production ownership • Advanced PySpark, including performance tuning, partitioning strategy, and cost-aware cluster sizing on EMR or equivalent • Strong Python with clean, tested, production-grade code • Deep AWS experience with EMR, Glue, Lambda, S3, Athena, Step Functions, EventBridge, IAM, and VPC networking • Advanced SQL, including complex analytical queries, query optimization, and data modeling on Athena/Presto and PostgreSQL • Production experience with at least one open table format; Apache Hudi strongly preferred, with Iceberg or Delta Lake also valued • Experience implementing data validation, quality gates, monitoring, and incident response for production data • Terraform and/or CloudFormation in a CI/CD workflow • Professional working proficiency in English and Spanish (B2+) • Bachelor’s degree in Computer Science, Systems Engineering, Data Engineering, or equivalent practical experience • Real estate, property, or geospatial data experience is a plus • Experience building or operating public/internal data APIs is a plus • Observability tooling experience is a plus • ML-adjacent engineering experience is a plus • Modern Python tooling, Agile/SCRUM experience, and documentation practices are a plus

🏖️ Benefits

• Performance Share Bonus • Flex PTO (up to 26 days per year) • Home Office Upgrade Bonus • HMO Bonus • Full-time Remote Work • Impact Moments • Opportunity for growth and team-building potential • Ongoing support and budget to develop new skills

Apply Now

Similar Jobs

🔥 15 hours ago

Capgemini

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Data Engineer managing data application releases, deployments, and DevOps pipelines for Capgemini Engineering. Supporting migration, architecture, and environment management for client solutions.

🇨🇴 Colombia – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

Airflow

Azure

PySpark

Python

Spark

Tableau

Unity

🕒 Yesterday

ARKHO

51 - 200

🤖 Artificial Intelligence

🤝 B2B

🏢 Enterprise

Data Engineer remoto apoyando procesos ETL/ELT, validación de datos y troubleshooting en ARKHO. Contribución operativa a la migración y apagado de DataMarts.

🇨🇴 Colombia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

🗣️🇪🇸 Spanish Required

ETL

SQL

🕒 5 days ago

Valtech

5001 - 10000

💼 Consulting

📣 Marketing

☁️ SaaS

Senior Data Engineer building Snowflake and Agentic AI data pipelines at Valtech, an experience innovation company. Developing scalable systems with Databricks, Azure, Python, and SQL.

🇨🇴 Colombia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

AWS

Azure

Cloud

Kafka

Postgres

Python

Scikit-Learn

Spark

SQL

Terraform

🕒 September 1

Gorilla Logic

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

FinOps Data Engineer building BigQuery pipelines and Power BI dashboards for Gorilla Logic’s AI usage and cost analytics. Supporting FinOps, forecasting, allocation, and optimization.

🇨🇴 Colombia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

Airflow

Amazon Redshift

Apache

AWS

Azure

BigQuery

Cloud

Google Cloud Platform

Python

SQL

Terraform

🕒 August 31

Gorilla Logic

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

Senior DataOps Engineer building governed Snowflake data platforms for Gorilla Logic. Transforming CRM, ERP, SaaS, and product data into trusted analytics and AI-ready business products.

🇨🇴 Colombia – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

Cloud

ERP

ETL

Python

SQL