Healthcare Data Scientist

🕒 August 4

🏛️ District of Columbia, Virginia, +1 more states – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

📊 Data Scientist

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cherokee Federal

Cherokee Federal

5001 - 10000 employees

Founded 1969

🏥 Healthcare

📦 Logistics

🏭 Manufacturing

Healthcare • Logistics • Manufacturing

Cherokee Federal is a U. S. federal systems integrator and government contractor that empowers mission success for more than 60 U. S. federal agencies. With a global workforce of over 5,000, it delivers advanced technology (cloud, cybersecurity, data & analytics), health services, intelligence analysis and operational support, logistics and sustainment, mission-critical manufacturing, program and engineering technical services, and dynamic contracting solutions to support federal priorities and national security. Cherokee Federal is part of Cherokee Nation Businesses and focuses on mission-focused, U. S. -made solutions.

📋 Description

• Develop, maintain, and optimize data pipelines using SQL, Python (pandas), and PySpark • Execute and manage notebook-based workflows within Azure Synapse, including debugging and documentation • Process and transform structured and semi-structured data in CSV, JSON/NDJSON, and Parquet formats • Work within ETL/ELT pipelines across raw, curated, and production data layers • Ingest, profile, map, and transform healthcare data from EHR systems and interface feeds while preserving source lineage and clinical context • Perform structured data validation, including row counts, null checks, duplicate detection, schema validation, and allowed value enforcement • Identify and resolve schema drift, inconsistencies, and transformation errors across pipeline stages • Apply repeatable testing and validation practices, reproduce issues, verify fixes, and ensure data reliability • Validate source-to-target mappings and reconcile records across source and destination systems during data conversion and migration • Conduct exploratory data analysis to identify patterns, anomalies, and data quality concerns • Develop derived datasets for reporting, analytics, and downstream data use cases • Collaborate with stakeholders to translate data requirements into usable datasets and metrics • Interpret schemas, column definitions, data types, keys, and table relationships • Manage dataset grain and assess how join strategies affect row counts and outputs • Trace data issues from source ingestion through transformation logic to final outputs • Use logs and debugging approaches to diagnose and resolve pipeline issues • Document transformations, assumptions, mappings, and validation results • Collaborate with engineers, analysts, and stakeholders to ensure data usability, integrity, and requirements alignment • Communicate data issues, findings, and workflow updates with technical team members

🎯 Requirements

• Hands-on experience writing SQL queries for data transformation and analysis • Experience using Python, such as pandas, for data processing • Hands-on experience using PySpark for distributed data processing • Experience working within cloud-based data platforms, preferably Azure Synapse or similar • Understanding of ETL/ELT concepts and data pipeline architecture • Experience with structured and semi-structured data formats, including CSV, JSON, and Parquet • Familiarity with Git and collaborative development workflows • Strong problem-solving and debugging skills across data pipelines • Ability to validate and ensure data quality through structured checks and testing practices • Strong written and verbal communication skills

🏖️ Benefits

• Generous paid time-off • Employee incentive program • Continuous learning culture • Internal Investment Projects (IIP) • Virtual brown-bags/level-ups • Other professional development activities • Recruiting bonuses • 3% 401k Safe Harbor contributions • Medical insurance • Dental insurance • Vision insurance • Long-term disability insurance • Short-term disability insurance • AD&D insurance • Life insurance

Apply Now

Similar Jobs

🕒 August 4

GondolaBio

11 - 50

🏥 Healthcare

💼 Consulting

🏭 Manufacturing

Computational biology scientist integrating multiomics and AI analyses for GondolaBio’s genetic-disease therapeutics. Supporting target identification, biomarker discovery, and translational research.

🇺🇸 United States – Remote

💵 $155k - $205k / year

💰 $300M Venture Round - GondolaBio on 2024-08

⏰ Full Time

🟠 Senior

📊 Data Scientist

🕒 August 4

Aptive Resources

501 - 1000

🎖️ Defense

📦 Logistics

📣 Marketing

Senior Data Scientist building AI learning tools and cloud data models for the Veterans Health Administration. Integrating VA systems and advancing responsible AI education across federal healthcare.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

📊 Data Scientist

🕒 August 4

OneStudyTeam

201 - 500

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior Data Scientist developing predictive models and patient-matching algorithms for OneStudyTeam’s clinical-trial platform. Improving enrollment, trial management, model fairness, and compliant healthcare analytics.

🇺🇸 United States – Remote

💵 $140k - $190k / year

⏰ Full Time

🟠 Senior

📊 Data Scientist

🕒 August 4

HubSpot

1001 - 5000

🤝 B2B

☁️ SaaS

📣 Marketing

Senior Product Manager owning CRM data associations and record creation for HubSpot’s AI-powered customer platform. Driving cross-platform roadmap, relational data quality, and trustworthy AI experiences across HubSpot products.

🇺🇸 United States – Remote

💵 $140k - $175k / year

⏰ Full Time

🟠 Senior

📊 Data Scientist

🦅 H1B Visa Sponsor

infoinfo

🕒 August 4

SpyCloud

51 - 200

🔒 Cybersecurity

🔐 Security

🏢 Enterprise

Senior Data Scientist building applied ML models for SpyCloud, which protects identities from cybercrime. Owning data pipelines, model deployment, monitoring, and cybersecurity feature development.

🇺🇸 United States – Remote

💵 $154k - $200k / year

⏰ Full Time

🟠 Senior

📊 Data Scientist

🦅 H1B Visa Sponsor

infoinfo