Data Engineer – Mid Level

🔥 0 minutes ago

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

👻 Ghost score 16%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Irth Solutions

Irth Solutions

51 - 200 employees

Founded 1995

☁️ SaaS

⚡ Energy

📡 Telecommunications

💰 Private Equity Round on 2015-07

SaaS • Energy • Telecommunications

Irth Solutions is a market-leading provider of a SaaS platform focused on enhancing resilience and reducing risk in the management of critical network infrastructure. Their solutions are trusted by energy, utility, and telecom companies across the U. S. and Canada, offering capabilities in damage prevention, training, asset inspections, land management, and 811 ticket management. By leveraging business intelligence, analytics, and geospatial data, Irth Solutions provides comprehensive situational awareness to proactively manage and mitigate risks in network infrastructure. The acquisition of OneBridge Solutions enhances their asset performance management offerings.

📋 Description

• Build, maintain, and enhance batch and streaming data-ingestion pipelines across AWS, Azure, and GCP • Develop pipelines using Databricks Workflows, Apache Spark/PySpark, SQL, Delta Live Tables, and Databricks Lakeflow components • Implement Bronze–Silver–Gold medallion architecture, CDC, SCD Type 1 and Type 2, schema evolution, validation, reconciliation, and data-quality rules • Configure and maintain Delta Lake storage structures, tables, schemas, partitions, and optimization routines • Apply OPTIMIZE, Z-ORDER, VACUUM, partitioning, and file-management strategies • Support Unity Catalog metadata, cataloging, lineage, governance, and integration with Microsoft Purview • Integrate Amazon S3, Azure Storage, and Google Cloud Storage with Databricks • Implement data-quality checks, profiling, validation, monitoring, RBAC policies, security controls, and data-classification tags • Build, schedule, monitor, and maintain production workflows using Databricks Workflows, Delta Live Tables, Azure Data Factory, and other approved tools • Contribute to CI/CD pipelines, automated testing, deployment, environment management, and DEV–QA–PROD promotion • Monitor production pipelines, troubleshoot failures, investigate root causes, support recovery, and tune performance • Collaborate with the Senior Data Architect, Data Scientists, ML Engineers, Analysts, Product teams, and engineering stakeholders • Participate in architecture reviews, technical design discussions, coding reviews, and engineering standards meetings • Document pipelines, data flows, data dictionaries, transformation logic, data-quality rules, test cases, job schedules, and operational procedures • Translate architecture designs and technical standards into reliable, production-ready pipelines and platform capabilities

🎯 Requirements

• 3–5 years of experience in Data Engineering, ETL development, or cloud data platform engineering • Hands-on experience with Databricks, Apache Spark, PySpark, or other distributed data-processing technologies • Strong proficiency in SQL, including structured data transformation, joins, aggregations, and performance-aware query development • Experience with at least one major cloud platform; Microsoft Azure preferred, with AWS and/or GCP also valuable • Understanding of data modeling, data quality, schema evolution, data validation, pipeline monitoring, and troubleshooting • Basic understanding of RBAC, encryption, credential and secret management, and secure access to cloud and data-platform resources • Working knowledge of Delta Lake, medallion architecture, and modern lakehouse best practices preferred • Experience with metadata, cataloging, and governance platforms such as Unity Catalog, Microsoft Purview, or AWS Glue Data Catalog preferred • Experience with workflow orchestration and scheduling technologies such as Azure Data Factory, Databricks Workflows, Apache Airflow, or similar frameworks preferred • Experience with Git-based development, CI/CD, and DevOps practices preferred • Knowledge or experience in geospatial/GIS data, BI semantic layers particularly Power BI, or data preparation for AI/ML workloads preferred • Relevant cloud or Databricks certifications preferred • Understanding of asset integrity management concepts is nice to have • Experience with oil & gas, utility, infrastructure, or pipeline asset data is nice to have • Familiarity with regulatory, compliance, and audit-reporting requirements is nice to have

🏖️ Benefits

• Competitive compensation package based on experience and qualifications • Medical, Dental, and Vision Insurance • 401(k) Plan with Company Match • Generous Paid Time Off (PTO) • Company-Paid Holidays • Flexible Work Options / work-from-home opportunities, depending on role and business needs • On-Call Compensation for eligible on-call shifts

Apply Now

Similar Jobs

🔥 18 hours ago

Harris Global Business Services (GBS)

501 - 1000

🏥 Healthcare

📦 Logistics

📣 Marketing

Software Engineer developing data systems and ETL software remotely from India. Producing code, unit tests, documentation, technical specifications, and proof of concepts.

🔥 18 hours ago

Harris Computer

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

Software Engineer developing ETL and data systems. Translating requirements into code, tests, documentation, and production-ready software components.

🕒 Yesterday

Fusion Consulting

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Data Developer building ETL pipelines, SQL models, and Snowflake solutions for Fusion Consulting’s global life sciences consulting clients. Supporting SAP, analytics, and machine learning initiatives.

Azure

Cloud

ERP

ETL

SQL

SSIS

🕒 2 days ago

StarTree

11 - 50

💼 Consulting

📦 Logistics

📣 Marketing

Senior Software Engineer solving distributed-systems challenges in Apache Pinot for StarTree’s real-time analytics cloud. Improving performance, reliability, concurrency, and data infrastructure at massive scale.

Distributed Systems

Java

🕒 6 days ago

Palo Alto Labs

51 - 200

💼 Consulting

🔒 Cybersecurity

🤖 Artificial Intelligence

Data Engineer optimizing Databricks platforms and Azure data pipelines for IT operations. Building CI/CD, orchestration, infrastructure automation, governance, and reliability practices.

Airflow

Azure

Cloud

Cyber Security

Python

SQL

Terraform