Senior Data Engineer

🕒 July 14

🇨🇴 Colombia – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 11%

infoinfo

🗣️🇪🇸 Spanish Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 8020REI

8020REI

11 - 50 employees

Founded 2017

💼 Consulting

📦 Logistics

🏠 Real Estate

Consulting • Logistics • Real Estate

8020REI is a company that specializes in providing AI-driven predictive data and advanced marketing strategies for real estate investors. Their services include direct mail, SMS, and cold calling to generate more motivated seller leads, allowing clients to maximize their return on investment. By leveraging machine learning, predictive analytics, and a proven marketing cadence plan, 8020REI helps clients save money and focus on the most promising properties likely to sell at a discount. Their robust data services and marketing insights support transparency, efficiency, and effective targeting in real estate investment.

📋 Description

• Own Big Data pipelines end to end. • Design, build, and optimize PySpark ETL/ELT pipelines on Amazon EMR and AWS Glue that process nationwide, county-partitioned property data (First American, BuildZoom permits, market comps) on daily and monthly cadences. • Run our lakehouse. • Operate and evolve our Bronze → Silver → Gold data lake on S3 with Apache Hudi and the AWS Glue Data Catalog, queried through Athena, including schema contracts, partitioning strategy, compaction, and performance tuning. • Orchestrate and automate. • Build reliable orchestration with AWS Step Functions, EventBridge, and Lambda; make reruns, backfills, and failure recovery boring and documented. • Enforce data quality. • Implement and extend our Data QA Audit Standard: layer contracts, write-audit-publish gating, quarantine flows, drift monitoring, and actionable Slack alerting, so bad data never reaches a client list. • Operate production databases and APIs. • Manage Aurora PostgreSQL and DynamoDB workloads, and run data-serving APIs (API Gateway, SQS-backed async workers), such as our Address Resolution Service, to production SLAs with dashboards and runbooks. • Own cloud cost. • Monitor, report, and reduce the AWS data-platform bill (EMR cluster sizing, Glue/Lambda usage, S3 lifecycle, Athena scan costs) as a first-class engineering responsibility. • Ship infrastructure as code. • Define infrastructure with Terraform and CloudFormation, delivered through GitHub Actions CI/CD with tests, linting, and coverage gates, we run a disciplined PR, branch-policy, and code-review culture. • Partner with Data Science. • Build the feature pipelines, training datasets, and serving paths behind our ML scoring and valuation models (scikit-learn/XGBoost-family stack), and co-own the handoff contracts between DS and DE. • Document like a pro. • Maintain runbooks, architecture docs, and data dictionaries (Confluence) so any teammate can operate what you build.

🎯 Requirements

• 4+ years of hands-on data engineering with large-scale distributed data systems and a track record of production ownership (not just development). • Advanced PySpark : performance tuning, partitioning strategy, and cost-aware cluster sizing on real workloads (EMR or equivalent). • Strong Python : clean, tested, production-grade code (we use pytest, ruff, mypy, and coverage gates in CI). • Deep AWS experience : EMR, Glue, Lambda, S3, Athena, Step Functions, EventBridge, IAM, and VPC networking; comfort operating (not just deploying to) these services. • Advanced SQL : complex analytical queries, query optimization, and data modeling on both a warehouse/lake engine (Athena/Presto) and PostgreSQL. • Lakehouse experience : hands-on production work with at least one open table format (Apache Hudi strongly preferred; Iceberg or Delta Lake also valued) and medallion-style architecture. • Data quality mindset : experience implementing validation, quality gates, monitoring, and incident response for production data. • Infrastructure as Code : Terraform and/or CloudFormation in a CI/CD workflow. • English and Spanish : professional working proficiency in both (B2+). • Bachelor’s degree in Computer Science, Systems Engineering, Data Engineering, or equivalent practical experience.

🏖️ Benefits

• Competitive Base Compensation • Profit Share Bonus • Flex PTO (up to 26 days per year) • Home Office Upgrade Bonus • HMO Bonus • Full-time Remote Work • Opportunity for growth and team-building potential • Ongoing support and budget to develop new skills

Apply Now

Similar Jobs

🕒 July 10

Caseware

201 - 500

💸 Finance

🏢 Enterprise

☁️ SaaS

Senior Software Developer focusing on AI Data Engineering at Caseware. Responsible for building scalable data ingestion and processing systems.

🇨🇴 Colombia – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

AWS

Cloud

Distributed Systems

Terraform

🕒 July 3

CI&T

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior Data Architect enabling clients to modernize their data infrastructure with AWS solutions. Leading architecture design and team collaboration for cloud-native data platforms.

🇨🇴 Colombia – Remote

💰 $5.5M Venture Round on 2014-04

⏰ Full Time

🟠 Senior

🚰 Data Engineer

AWS

Cloud

🕒 June 29

Truelogic Software

501 - 1000

☁️ SaaS

🤝 B2B

🏢 Enterprise

Senior Data Engineer developing and maintaining a data platform for veterinary technology. Collaborating remotely with a skilled engineering team to ensure data insights and reliability.

🇨🇴 Colombia – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

Airflow

Amazon Redshift

Apache

AWS

Elixir

Erlang

ETL

Kafka

Kubernetes

Postgres

Python

SQL

Terraform

🕒 June 17

BlueCloud

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

Senior Data Engineer designing and implementing Snowflake data solutions for enterprise organizations. Leading the architecture of scalable data pipelines and enforcing data governance standards.

🇨🇴 Colombia – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

Airflow

AWS

Azure

Cloud

ETL

Google Cloud Platform

Python

SQL

Vault

🕒 June 12

Blend360

501 - 1000

🏥 Healthcare

🏨 Hospitality

✈️ Travel

Senior Data Engineer at Blend responsible for designing and optimizing data pipelines. Collaborating with teams to ensure data accuracy and quality for enterprise initiatives.

🇨🇴 Colombia – Remote

💰 $100M Private Equity Round on 2022-08

⏰ Full Time

🟠 Senior

🚰 Data Engineer

ETL

Python

SQL