Search Remote Jobs

Healthcare Data Scientist

🔥 1 minute ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cherokee Federal

Cherokee Federal

5001 - 10000 employees

Founded 1969

🏥 Healthcare

📦 Logistics

🏭 Manufacturing

Healthcare • Logistics • Manufacturing

Cherokee Federal is a U. S. federal systems integrator and government contractor that empowers mission success for more than 60 U. S. federal agencies. With a global workforce of over 5,000, it delivers advanced technology (cloud, cybersecurity, data & analytics), health services, intelligence analysis and operational support, logistics and sustainment, mission-critical manufacturing, program and engineering technical services, and dynamic contracting solutions to support federal priorities and national security. Cherokee Federal is part of Cherokee Nation Businesses and focuses on mission-focused, U. S. -made solutions.

📋 Description

• Develop, maintain, and optimize data pipelines using SQL, Python (pandas), and PySpark • Execute and manage notebook-based workflows within Azure Synapse, including debugging and documentation • Process and transform structured and semi-structured data in CSV, JSON/NDJSON, and Parquet formats • Work within ETL/ELT pipelines across raw, curated, and production data layers • Ingest, profile, map, and transform healthcare data from EHR systems and interface feeds while preserving source lineage and clinical context • Perform structured data validation, including row counts, null checks, duplicate detection, schema validation, and allowed value enforcement • Identify and resolve schema drift, inconsistencies, and transformation errors across pipeline stages • Apply repeatable testing and validation practices, reproduce issues, verify fixes, and ensure data reliability • Validate source-to-target mappings and reconcile records across source and destination systems during data conversion and migration • Conduct exploratory data analysis to identify patterns, anomalies, and data quality concerns • Develop derived datasets for reporting, analytics, and downstream data use cases • Collaborate with stakeholders to translate data requirements into usable datasets and metrics • Interpret schemas, column definitions, data types, keys, and table relationships • Manage dataset grain and assess how join strategies affect row counts and outputs • Trace data issues from source ingestion through transformation logic to final outputs • Use logs and debugging approaches to diagnose and resolve pipeline issues • Document transformations, assumptions, mappings, and validation results • Collaborate with engineers, analysts, and stakeholders to ensure data usability, integrity, and requirements alignment • Communicate data issues, findings, and workflow updates with technical team members

🎯 Requirements

• Hands-on experience writing SQL queries for data transformation and analysis • Experience using Python, such as pandas, for data processing • Hands-on experience using PySpark for distributed data processing • Experience working within cloud-based data platforms, preferably Azure Synapse or similar • Understanding of ETL/ELT concepts and data pipeline architecture • Experience with structured and semi-structured data formats, including CSV, JSON, and Parquet • Familiarity with Git and collaborative development workflows • Strong problem-solving and debugging skills across data pipelines • Ability to validate and ensure data quality through structured checks and testing practices • Strong written and verbal communication skills

🏖️ Benefits

• Generous paid time-off • Employee incentive program • Continuous learning culture • Internal Investment Projects (IIP) • Virtual brown-bags/level-ups • Other professional development activities • Recruiting bonuses • 3% 401k Safe Harbor contributions • Medical insurance • Dental insurance • Vision insurance • Long-term disability insurance • Short-term disability insurance • AD&D insurance • Life insurance

Apply Now

Similar Jobs

🔥 41 minutes ago

Humana

10,000+ employees

🏥 Healthcare

🛡️ Insurance

⚕️ Healthcare Insurance

AI and Automation Lead improving encounter data reconciliation and quality at Humana, a U.S. healthcare company. Owning model performance, automation roadmaps, and cross-functional healthcare data improvements.

🔥 51 minutes ago

TD

10,000+ employees

🏦 Banking

💸 Finance

🛡️ Insurance

Data Scientist III developing AML/CTF models and AI solutions for TD Bank. Applying Python, SQL, machine learning, and LLM technologies to financial crime risk analytics.

🔥 1 hour ago

Jerry

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Data Scientist driving analytics across Jerry.ai’s car and home asset-management platform. Defining metrics, running experiments, and influencing growth, product, AI, and partnership decisions.

🔥 1 hour ago

Jerry

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Data Scientist driving analytics across growth, product, AI, and partnerships for Jerry.ai’s physical-asset management app. Defining metrics, running experiments, and turning insights into business decisions.

🔥 1 hour ago

Jerry

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Associate Data Scientist driving analytics for Jerry.ai’s car and home asset-management app. Defining metrics, experiments, and recommendations across growth, product, AI, or partnerships.