Lead Data Engineer – GenAI, LLM Applications

Job not on LinkedIn

🕒 April 29

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 55%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Clario

Clario

5001 - 10000 employees

Founded 1973

🏥 Healthcare

💼 Consulting

📦 Logistics

💰 Private Equity Round on 2019-10

Healthcare • Consulting • Logistics

Clario is a company specializing in accelerating clinical trials from initiation to implementation through advanced technologies and services. Since 2018, Clario has revolutionized endpoint analyses in clinical trials by integrating over 30 artificial intelligence-enabled solutions across more than 600 active trials, enhancing data quality and patient privacy while expediting data collection processes. Clario provides a comprehensive clinical trial management platform, offering solutions such as eCOA, cardiac safety, medical imaging, precision motion, and respiratory services in various therapeutic areas including oncology, cardiology, and neurology. Known for its global reach, Clario supports clinical trials in over 100 countries with a strong focus on decentralized and hybrid trial models. The company's commitment to patient safety and innovation is reflected in their over 26,000 trials and involvement in numerous new drug approvals.

📋 Description

• Design, develop, and maintain scalable software architectures and data pipelines • Write clean, reusable, and well-tested Python code using Flask and related libraries • Build and integrate LLM-powered solutions, including RAG pipelines, intelligent agents, and automated workflows using AWS Bedrock or similar services • Develop and optimize complex SQL across Oracle, MS SQL Server, PostgreSQL, and Snowflake • Design and implement ETL pipelines using Snowflake and related technologies • Implement scheduling and orchestration with Apache Airflow or similar frameworks • Establish data quality, versioning, and governance practices • Develop and maintain architectures and models for structured and unstructured data • Troubleshoot production issues and improve software quality, performance, and reliability • Deploy, manage, and support AWS solutions • Create source-to-target mappings and support data and code migration initiatives • Gather requirements, translate business needs into technical solutions, and produce documentation • Collaborate with product managers, analysts, and cross-functional teams to deliver data-driven insights and reporting using Plotly and Power BI

🎯 Requirements

• Bachelor’s or higher degree in Computer Science, Information Technology, or a related technical field • 5+ years of professional experience in software engineering, data engineering, or data-focused development roles • Strong proficiency in Python, including Django or Flask, pandas, NumPy, Plotly, and ag-Grid • Strong SQL expertise with Oracle, MS SQL Server, PostgreSQL, and/or Snowflake • Proven experience writing complex SQL, including analytical and window functions, subqueries, all join types, DML/DDL/TCL statements, CASE expressions, and performance tuning • Working knowledge of cloud platforms, with a preference for AWS, including S3, EC2, Secrets Manager, Bedrock, and Lambda • Experience using GitHub Copilot and LangChain for LLM-powered applications and workflows • Experience with Git-based version control systems and CI/CD pipelines • Familiarity with data modeling for structured and unstructured data • Willingness to work across all phases of the SDLC • Preferred exposure to clinical trial lifecycle or clinical data management, Plotly, Power BI, HTML5, CSS3, JavaScript, Jira, Confluence, Microsoft Teams, data analysis, data cleansing, SQL, and Excel

🏖️ Benefits

• Competitive compensation aligned with local market practices • Comprehensive health and wellness benefits • Paid time off and company holidays • Opportunities for professional development, learning, and career growth • The flexibility of working from Bangalore or remotely within India

Apply Now

Similar Jobs

🕒 April 28

Forbes

201 - 500

💼 Consulting

📣 Marketing

📦 Logistics

Data Engineering Lead needed to scale and manage a team at Forbes Advisor. Oversee delivery of data pipelines and ensure engineering excellence in a remote-first environment.

Airflow

BigQuery

ETL

Google Cloud Platform

Kafka

Python

Spark

SQL

🕒 April 27

Mactores

51 - 200

💼 Consulting

🏢 Enterprise

AWS Data Engineer (Senior) designing and maintaining data pipelines using AWS technologies. Collaborating with teams to optimize data models and solutions in a remote working environment.

Airflow

Amazon Redshift

AWS

PySpark

SQL

🕒 April 24

Codvo.ai

51 - 200

🤖 Artificial Intelligence

🔒 Cybersecurity

☁️ SaaS

Full Stack Data Engineer at Codvo responsible for designing and maintaining Databricks data pipelines. Collaborating with data scientists and implementing best-practice DevOps/MLOps processes for operationalizing machine learning models.

AWS

ETL

Flask

PySpark

Python

Spark

Unity

🕒 April 23

NextGen Healthcare

1001 - 5000

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Senior Staff Data Engineer guiding the design of enterprise data platforms for analytics and AI/ML systems. Mentor engineers while driving standards and best practices.

Cloud

🕒 April 23

Forbes

201 - 500

💼 Consulting

📣 Marketing

📦 Logistics

Data Engineer (L2) at Forbes Advisor building data pipelines for marketing analytics and reporting. Requires experience in working with Python, SQL, and Meta Ads ecosystem.

🇮🇳 India – Remote

💰 $200M Corporate Round on 2022-02

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

Airflow

BigQuery

Cloud

ETL

Microservices

Python

SQL