Senior Data Engineer

🔥 11 minutes ago

🇺🇸 United States – Remote

💵 $130k - $160k / year

⏰ Full Time

🟠 Senior

🚰 Data Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Zifo

Zifo

1001 - 5000 employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Zifo is a global specialist scientific and process informatics service provider working in research, development, manufacturing, and clinical domains. They offer a wide range of services including scientific IT services, cloud migration, bioinformatics consulting, data science, and digital transformation. Zifo serves industries such as pharmaceuticals, biotechnology, contract research organizations, and more. They are committed to advancing scientific discovery through innovative solutions and expert consulting, catering to the needs of science-driven organizations.

📋 Description

• Design, build, and operate data infrastructure for a clinical trial design platform • Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles from PubMed, CTIS, and ClinicalTrials.gov • Handle PDF parsing, text extraction, and structured data normalization • Design and implement relational and knowledge graph data models in Amazon Aurora and GraphDB • Develop embedding and vectorization pipelines for RAG-based retrieval in LangGraph agentic workflows • Build and maintain ETL/ELT workflows that transform unstructured clinical content into queryable, linked data • Implement clinical data quality validation, including protocol section classification, entity extraction completeness, and cross-reference integrity • Build data-serving APIs with Python/FastAPI for Angular frontend and LangGraph agent-layer integration • Set up data lineage tracking and audit trails to support regulatory traceability of AI-generated trial design recommendations

🎯 Requirements

• Strong Python development experience, including PDF/document parsing libraries such as PyMuPDF, pdfplumber, unstructured.io, or similar • Advanced PostgreSQL-compatible SQL, including Amazon Aurora; experience with schema design, migrations, query optimization, and indexing strategies for large clinical datasets • Hands-on experience with Neptune, Neo4j, or similar graph databases • Proficiency in SPARQL or Cypher • Experience with ontology and knowledge graph modeling for biomedical entities • Experience with AWS services including Aurora PostgreSQL, S3, Lambda, Step Functions, SQS/SNS, and IAM • Experience with PDF text extraction, document section classification, and named entity recognition (NER) for clinical/biomedical text • Familiarity with embedding models and vector stores such as OpenSearch, pgvector, or Pinecone • Experience building data-serving APIs using FastAPI, including asynchronous programming patterns and backend integration • Experience preparing data for LangChain/LangGraph applications and designing RAG pipelines, including chunking, retrieval, reranking, and prompt-data integration • Experience with Airflow, Prefect, AWS Step Functions, Temporal, or similar workflow orchestration tools • Ability to design multi-stage DAGs with dependency management, retry logic, monitoring, and error handling • Experience with Terraform or AWS CDK, Docker, and Git, including automated pipeline testing and deployment on AWS • Understanding of clinical trial structure, including protocol sections, objectives, endpoints, eligibility criteria, study design, and statistical considerations • Familiarity with clinical data standards or terminologies such as MeSH, MedDRA, SNOMED, ATC codes, or CDISC is a strong plus • Awareness of regulatory data integrity requirements, including 21 CFR Part 11, EU Annex 11, and ALCOA+ principles

🏖️ Benefits

• Accrued vacation • Medical insurance • Dental insurance • Vision insurance • 401(k) with company matching • Life insurance • Flexible spending accounts • Equal opportunity employment and diversity commitment

Apply Now

Similar Jobs

🔥 1 hour ago

Intertwine Associates

1 - 10

🏥 Healthcare

📦 Logistics

📣 Marketing

Senior AI/Data Engineer leading Azure data platforms, machine learning, and LLM applications. Delivering secure, governed solutions for federal agencies and other clients.

🔥 1 hour ago

Intertwine Associates

1 - 10

🏥 Healthcare

📦 Logistics

📣 Marketing

AI/Data Engineer building Python ETL workflows, Azure data services, and ML-ready infrastructure. Supporting federal and regulated clients with scalable, governed data platforms.

🔥 3 hours ago

WISEcode

11 - 50

🍽️ Food & Beverage

🏥 Healthcare

🤖 Artificial Intelligence

Software Developer & Data Engineer building WISEcode’s food intelligence lakehouse. Designing ingestion, entity resolution, and data pipelines powering food decisions.

🔥 6 hours ago

Stellic

51 - 200

📚 Education

☁️ SaaS

Platform Data Architect integrating university student information systems for Stellic’s student-success platform. Owning partner deployments, data feeds, migrations, and scalable integration standards.

🇺🇸 United States – Remote

💵 $118k - $173k / year

💰 $11M Series A - Stellic on 2022-03

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

🔥 7 hours ago

Robots & Pencils

51 - 200

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Senior Data Engineer building scalable pipelines, warehouses, and AI/ML data platforms. Supporting Robots and Pencils’ production-ready enterprise AI systems.

🇺🇸 United States – Remote

💵 $119.5k - $164.9k / year

💰 Venture Round on 2022-04

⏰ Full Time

🟠 Senior

🚰 Data Engineer