Senior Data Engineer, Consultant

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of DiligenceVault

DiligenceVault

51 - 200 employees

💸 Finance

💳 Fintech

🏢 Enterprise

Finance • Fintech • Enterprise

DiligenceVault is a secure digital diligence platform designed to streamline and digitize the due diligence processes for investment and management research teams. The platform is trusted by over 70,000 users globally to simplify and enhance activities such as manager selection, monitoring, RFP automation, ESG data collection, and regulatory compliance. By leveraging technology, including AI, DiligenceVault helps users to efficiently manage due diligence tasks, centralize content, and improve collaboration through its comprehensive digital ecosystem. The company is backed by Goldman Sachs and provides global support with teams located in major financial hubs, including New York, London, Singapore, and India.

📋 Description

• Educate leadership and senior architects on traditional and AI-native data engineering concepts • Assess the current data infrastructure end to end, map data flows, identify gaps and technical debt, and produce a current-state/target-state assessment and prioritized roadmap • Design the target data platform architecture across ingestion, transformation, storage, serving, and observability • Architect the PostgreSQL migration and multi-workload environment, including pgvector, analytical workloads, transactional queries, pooling, replicas, partitioning, cutover, validation, query translation, and benchmarking • Design a canonical data layer for entity resolution, schema alignment, conflict resolution, temporal alignment, versioning, and auditability • Design data-residency architecture for a multi-tenant, data-sharing platform across regions and jurisdictions • Design the data governance framework covering access control, classification, lineage, retention, consent, quality accountability, and security controls • Define and prioritize data-platform use cases including cross-source intelligence, behavioral insights, enrichment, semantic search, compliance detection, and analytics/reporting • Produce architecture decision records, data-flow diagrams, tool evaluation guides, migration runbooks, and training decks • Provide ongoing architecture reviews, design consultations, and progress check-ins during execution • Deliver knowledge transfer, architectural guidance, strategic documents, and reference materials rather than primarily writing production code

🎯 Requirements

• 8–10+ years building data platforms across heterogeneous sources at meaningful scale • Deep expertise in relational databases, specifically PostgreSQL • Hands-on experience with PostgreSQL extensions such as pgvector, Citus, PostGIS, or similar • Experience with replication topologies, partitioning, and performance tuning • Experience migrating from SQL Server to PostgreSQL strongly preferred • Experience designing canonical data models across disparate sources, including entity resolution, master data management, and conflict resolution at scale • Experience architecting multi-region or data-residency-compliant systems, ideally in a multi-tenant SaaS context • Strong understanding of data governance, including access control, data classification, lineage, retention policies, and regulatory compliance • Knowledge of dimensional modeling, ETL/ELT, CDC, orchestration, and query optimization • Active, informed engagement with AI-native approaches including ML-driven quality, semantic matching, embedding pipelines, and LLM-assisted development • Ability to design multi-layer platforms and make defensible technology choices considering scale, cost, team size, and maintainability • Exceptional communication and teaching ability with senior architects and leadership • Experience defining data use cases tied to business outcomes • Availability for approximately 4–5 hours per day, 20–25 hours per week, with overlap during US working hours (6:00 PM–11:00 PM IST) • Strong-to-have experience with Python/Celery, Elasticsearch, Azure cloud services, .NET APIs, Kestra or similar orchestration • Familiarity with OCR, layout-aware extraction, and table parsing from financial PDFs and Word documents • Familiarity with dbt, Airflow/Dagster, Airbyte/dlt, Snowflake/Databricks • Familiarity with vector databases, RAG pipelines, feature stores, and embedding workflows • Financial services, investment management, or due diligence background is advantageous but not required • Track record of producing technical documentation and training materials

🏖️ Benefits

• Remote flexibility: work from anywhere in India • Startup energy and flat hierarchy • Opportunity to design for a platform used by financial institutions in 150+ countries • Work with a global team • Direct impact and quick shipping of work

Apply Now

Similar Jobs

🕒 July 27

Tech Minds Agency

1 - 10

💼 Consulting

📣 Marketing

🛍️ eCommerce

Data Engineer with expertise in BlackRock Aladdin and Snowflake development. Building data pipelines and ensuring high-quality data solutions while working in Bangalore.

Cloud

ETL

SQL

🕒 June 24

Smart Working

51 - 200

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior Data Engineer enhancing data infrastructure and performance at Smart Working for iGaming brands. Collaborating with product and data teams while leveraging AI and AWS technologies.

Airflow

AWS

ETL

Kafka

Python

SQL

🕒 June 8

Tech Minds Agency

1 - 10

💼 Consulting

📣 Marketing

🛍️ eCommerce

Data Architect role focusing on designing data solutions with 12+ years of experience. Involves data architecture development, data strategy implementation, and leading data migration projects.

Apache

Azure

MariaDB

MySQL

Postgres

Spark

SQL

Vault

🕒 June 4

ProArch

201 - 500

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Microsoft Fabric Data Engineer developing scalable data solutions and engaging with US-based clients. Key role in data modeling and client engagement to optimize data architectures.

Azure

Cloud

ETL

PySpark

Python

SQL

🕒 June 4

Mactores

51 - 200

💼 Consulting

🏢 Enterprise

SAP Data Engineer responsible for building extraction pipelines from SAP HANA to AWS S3. Collaborating with team members to ensure effective data management and reporting.

AWS

Cloud

PySpark

Spark