Senior Scientific Data Engineer, R&D Data Platform

🔥 0 minutes ago

🌽 Illinois – Remote

infoinfo

💵 $78k - $156k / year

⏰ Full Time

🟠 Senior

🚰 Data Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Abbott

Abbott

10,000+ employees

Founded 1888

🏥 Healthcare

⚕️ Healthcare Insurance

🧬 Biotechnology

Healthcare • Healthcare Insurance • Biotechnology

Abbott is a global healthcare company committed to advancing medical technologies and improving lives around the world. It offers a broad range of leading products in diagnostics, medical devices, nutrition, and branded generic medicines. Abbott's innovations such as the FreeStyle Libre glucose monitoring systems and BinaxNOW rapid antigen tests are transforming diabetes management and COVID-19 response. Through its partnerships and initiatives, Abbott aims to foster health equity, improve access to healthcare, and address critical health challenges like malnutrition and infectious diseases. Abbott is also dedicated to sustainability and social responsibility, striving to make life-changing technologies accessible and affordable.

📋 Description

• Lead the design and delivery of reusable tools and services for ingesting, validating, transforming, documenting, discovering, and sharing scientific data • Own platform capability areas end to end, including design, implementation, adoption, operational support, and long-term maintainability • Develop Python and SQL solutions including software packages, data pipelines, APIs, notebooks, workflow utilities, and lightweight internal applications • Create self-service workflows enabling researchers to prepare and share data consistently • Partner with scientific teams to understand studies, analytical workflows, data sources, and technical challenges, translating needs into a prioritized technical roadmap • Establish standards and reusable patterns for organizing and harmonizing data from disparate sources and drive adoption • Design automated data-quality and validation frameworks • Improve documentation, traceability, and discoverability of scientific datasets • Evaluate AWS services and features and partner with R&D DevOps on architecture, deployment patterns, and operational ownership • Develop solutions using Amazon S3, Athena, Glue, EMR, Lambda, and SageMaker • Prototype research-program solutions and generalize successful approaches into reusable platform capabilities • Provide technical leadership on cross-project and cross-team designs, including design reviews and trade-off documentation • Mentor engineers through code review, pairing, design feedback, and documentation • Support hands-on preparation and analysis of scientific data • Apply quantitative and scientific judgment to data and technical solutions • Use Spark or PySpark for appropriate distributed processing • Apply software-engineering practices including version control, testing, code review, documentation, dependency management, continuous integration, and reproducible development • Communicate technical concepts, decisions, trade-offs, limitations, and project status to technical, scientific, and leadership audiences • Operate independently in an evolving environment by scoping ambiguous problems, sequencing work, and making defensible decisions • Help scientists reduce manual data locating, cleaning, interpretation, and restructuring • Enable research teams to use documented, consistent data-preparation and sharing tools

🎯 Requirements

• Bachelor’s degree in computer science, data science, engineering, statistics, mathematics, bioinformatics, computational science, or another relevant quantitative discipline • Five or more years of relevant professional or applied research experience, or three or more years with an advanced degree in a relevant field • Advanced programming skills in Python • Strong SQL skills and experience working with structured and semi-structured data • Demonstrated track record of building reusable, maintainable software that others depend on • Experience designing and delivering several of the following: data pipelines, Python packages, APIs, analytical workflows, notebooks, or internal software tools • Substantial hands-on experience using AWS for data processing, analytics, scientific computing, or software development • Sufficient depth in AWS services and architecture to evaluate technical options, justify design recommendations, and define infrastructure requirements • Experience conducting or supporting quantitative research, such as statistical analysis, machine learning, computational modeling, or another data-intensive research activity • Experience cleaning, integrating, standardizing, or validating data from multiple sources at meaningful scale • Fluency with Git, automated testing, technical documentation, code review, and continuous integration • Demonstrated ability to investigate ambiguous problems, define an approach, and deliver a working solution with little guidance • Experience mentoring or providing technical guidance to other engineers, scientists, or analysts • Strong communication and collaboration skills across scientific and technical disciplines • Preferred: Advanced degree in a quantitative, computational, or life-science discipline • Preferred: Experience with biomedical, genomic, clinical, proteomic, imaging, laboratory, or other complex scientific data • Preferred: Experience supporting research in life sciences, healthcare, diagnostics, or a similarly data-intensive and regulated scientific environment • Preferred: Production experience with Spark or PySpark and distributed data processing • Preferred: Depth in AWS services such as Athena, Glue, EMR, SageMaker, Lambda, Step Functions, Lake Formation, or related data and analytics technologies • Preferred: Experience developing REST APIs or lightweight web applications used by non-engineering audiences • Preferred: Experience designing automated validation frameworks, data contracts, reusable data-processing libraries, or researcher-facing workflow tools • Preferred: Experience with metadata-management, data-catalog, or data-discovery platforms • Preferred: Experience with containerization, continuous integration and deployment, or infrastructure-as-code • Preferred: Experience supporting machine-learning workflows or preparing data for model development and evaluation • Preferred: Experience working with large files or multimodal datasets • Preferred: Familiarity with governance considerations for research data • Preferred: Experience working within a data mesh, data product, or federated data-ownership model

🏖️ Benefits

• Remote work arrangement • Standard work shift • Travel opportunity: 10% of the time

Apply Now

Similar Jobs

🔥 1 hour ago

Aptive Resources

501 - 1000

🎖️ Defense

📦 Logistics

📣 Marketing

Data Engineer building cloud ETL pipelines and healthcare analytics for the Department of Veterans Affairs. Supporting patient matching, data harmonization, and Power BI reporting at Aptive.

🔥 2 hours ago

CVS Health

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

🛒 Retail

Senior Data Engineer building Oracle-based reporting data layers and pipelines for CVS Health’s healthcare operations. Supporting near real-time business intelligence, production reliability, and data quality.

🔥 3 hours ago

Eli Lilly and Company

10,000+ employees

💊 Pharmaceuticals

🏥 Healthcare

🧬 Biotechnology

Data engineering lead building HIPAA-compliant patient analytics for LillyDirect, Lilly’s direct-to-patient pharmacy platform. Designing consent, ingestion, identity resolution, and semantic data products.

🇺🇸 United States – Remote

💵 $132k - $193.6k / year

💰 $6.5M Post-IPO Debt - Eli Lilly on 2024-02

⏰ Full Time

🟠 Senior

🚰 Data Engineer

🔥 3 hours ago

SHI International Corp.

5001 - 10000

💼 Consulting

📦 Logistics

🏭 Manufacturing

Data Engineer building scalable data pipelines and analytics assets for SHI International, a global IT solutions provider. Establishing data engineering practices and long-term platform architecture.

🔥 5 hours ago

Fuze Health

1001 - 5000

🏥 Healthcare

☁️ SaaS

💊 Pharmaceuticals

Data Engineer building pipelines, observability platforms, and integrations for Fuze Health’s digital healthcare ecosystem. Supporting accurate operational data and scalable infrastructure.