Senior Data Engineer – AI, LLM Solutions

Job not on LinkedIn

🔥 19 hours ago

🌐 Pakistan, Mexico, +2 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 15%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of AIM Qualifications and Assessment Group

AIM Qualifications and Assessment Group

51 - 200 employees

Founded 2011

📚 Education

🤝 Non-profit

👥 HR Tech

Education • Non-profit • HR Tech

AIM Qualifications and Assessment Group is a not-for-profit awarding organisation and Access Validating Agency with over 40 years of experience in creating and awarding qualifications, including Access to HE Diplomas. The organisation is dedicated to promoting inclusion, integrity, respect, and empowerment through flexible, professional, and technical qualifications ranging from Entry Level to Level 6. AIM also specializes in end-point assessments for apprenticeships in the creative, cultural, digital technologies, and professional services sectors, supporting various organisations in delivering quality training and assessments.

📋 Description

• Design and develop enterprise-scale data solutions using Microsoft Fabric, Azure Data Factory, Azure Synapse Analytics, Azure Data Lake Storage, and Azure Databricks • Build and optimize scalable ETL/ELT pipelines for structured, semi-structured, and unstructured data • Develop Lakehouse, Data Warehouse, and Medallion Architecture solutions • Design and implement batch and real-time data processing frameworks • Lead data modernization, migration, and integration initiatives from legacy platforms • Develop data foundations for AI, Machine Learning, and Generative AI solutions • Prepare, transform, and optimize datasets for AI applications, model training, and fine-tuning • Design and implement RAG architectures and vector search solutions • Support Azure OpenAI, Microsoft Copilot, and enterprise LLM implementations • Collaborate with Data Scientists and AI Engineers to operationalize AI and machine learning solutions • Build distributed data processing solutions using PySpark and Spark SQL • Develop automated integrations connecting APIs, SaaS applications, databases, and cloud services • Design reusable data engineering frameworks, components, and standards • Optimize workloads for performance, scalability, reliability, and cost efficiency • Implement data quality, validation, monitoring, and error-handling processes • Implement data governance and cataloging using Microsoft Purview • Establish data lineage, classification, cataloging, and access controls • Ensure alignment with security, privacy, compliance, SOC 2, NIST, and CMMC requirements • Design and support semantic data models for Power BI and enterprise analytics • Enable self-service analytics and support KPI frameworks, executive dashboards, reporting, and predictive analytics • Provide technical guidance and mentorship to junior and mid-level data engineers • Participate in solution architecture, technical design reviews, implementation planning, and presales activities • Work with clients and stakeholders to translate business needs into scalable technical solutions • Develop technical documentation, engineering standards, and reusable implementation patterns • Troubleshoot data pipeline, platform, and integration issues • Monitor and optimize data workloads and cloud resources • Support enhancements, upgrades, and ongoing maintenance of data platforms • Identify improvements in data quality, automation, performance, and operational efficiency • Stay current with Microsoft Fabric, Azure, AI, Generative AI, and data engineering technologies

🎯 Requirements

• Bachelor’s degree in Computer Science, Software Engineering, Information Technology, Data Science, or a related discipline • 8+ years of professional experience in Data Engineering, Data Platforms, or Cloud Data Solutions • 5+ years of hands-on experience with Microsoft Azure Data Services • Proven experience delivering enterprise-scale data, analytics, and AI solutions • Strong experience working with modern cloud data platforms and data architectures • Experience working within consulting, professional services, or enterprise environments • Strong communication, stakeholder management, and problem-solving skills • Proficiency with Microsoft Fabric, Azure Data Factory, Azure Synapse Analytics, Azure Data Lake Storage Gen2, Azure Databricks, and Azure Event Hubs • Knowledge of Data Warehouse, Lakehouse, and Medallion Architecture • Proficiency in Python, SQL, PySpark, Spark SQL, Kusto Query Language, REST API integrations, and ETL/ELT development • Experience with SQL Server, PostgreSQL, MySQL, NoSQL databases, data modeling, and optimization • Knowledge of Azure OpenAI, Generative AI, Prompt Engineering, Retrieval-Augmented Generation, Vector Databases/Vector Search, Machine Learning Pipelines, NLP fundamentals, LLM integration patterns, and AI-ready data architecture • Experience with Power BI, Tableau, Semantic Models, Executive Reporting, and KPI Frameworks • Knowledge of Microsoft Purview, Data Governance, Data Classification, Data Lineage, Data Cataloging, Access Control, Enterprise Security Standards, and SOC 2, NIST, and CMMC • Experience with Azure DevOps, Git, CI/CD Pipelines, Infrastructure as Code concepts, Containerization concepts, and cloud resource optimization • Ability to design and implement scalable data solutions and understand enterprise data architecture • Ability to work independently in a remote and cross-functional environment • Ability to mentor engineers and provide technical leadership • Strong focus on security, quality, performance, and reliability • Ability to translate business requirements into technical solutions • Preferred certifications include Microsoft Certified: Fabric Analytics Engineer Associate, Microsoft Certified: Azure Data Engineer Associate, Microsoft Certified: Azure Data Scientist Associate, Microsoft Certified: Azure Solutions Architect Expert, Databricks Certified Data Engineer, and relevant AI, Machine Learning, or Generative AI certifications

🏖️ Benefits

• Remote-first work environment • Opportunity to work on enterprise-scale Microsoft Fabric and Azure data platforms • Exposure to AI, Generative AI, Azure OpenAI, and LLM initiatives • Opportunity to work with modern cloud and data technologies • Collaborative, innovation-driven culture • Exposure to diverse enterprise and consulting projects

Apply Now

Similar Jobs

🕒 September 2

HR POD - Hiring Talent Globally

11 - 50

👥 HR Tech

🎯 Recruiter

🤝 B2B

SSIS/SQL Developer migrating Oracle schemas, Java ETL applications, and large datasets to SQL Server. Building, tuning, and supporting robust SSIS/T-SQL pipelines remotely from Pakistan.

Apache

Azure

Cloud

ETL

Java

Maven

MS SQL Server

Oracle

Spring

SQL

SSIS

TFS

🕒 August 20

Smart Working

51 - 200

💼 Consulting

🏥 Healthcare

📣 Marketing

Lead Data Engineer building real-time pipelines, vector databases, and ML workflows. Scaling AI infrastructure that automates operations for property management businesses.

Airflow

Apache

AWS

Azure

Cloud

ElasticSearch

Google Cloud Platform

Kafka

MongoDB

MySQL

NoSQL

Pandas

Postgres

Python

Spark

🕒 June 5

Darkroom

51 - 200

📣 Marketing

🤝 B2B

☁️ SaaS

Senior Data Engineer managing marketing data pipelines for consumer brands. Building ingestion layers and ensuring data reliability across multiple platforms in a fast-paced environment.

Amazon Redshift

BigQuery

ETL

Python

SQL

🕒 March 31

Creative Chaos

201 - 500

💼 Consulting

📣 Marketing

📦 Logistics

Data Engineer responsible for architecting and developing data pipelines within Azure Data Lake for Creative Chaos. Working with cross-functional teams on data integration and governance practices.

Azure

ETL

Java

Python

SQL

🕒 March 31

Creative Chaos

201 - 500

💼 Consulting

📣 Marketing

📦 Logistics

Senior Data Engineer at a dynamic team, responsible for developing data infrastructure and optimizing workflows. Collaborating across teams to fulfill data infrastructure needs.

Apache

AWS

Azure

Cloud

ETL

Java

Kafka

MongoDB

MySQL

Postgres

Python

Scala

Spark

SQL