
5001 - 10000 employees
Founded 1969
đĽ Healthcare
đŚ Logistics
đ Manufacturing
Healthcare ⢠Logistics ⢠Manufacturing
Cherokee Federal is a U. S. federal systems integrator and government contractor that empowers mission success for more than 60 U. S. federal agencies. With a global workforce of over 5,000, it delivers advanced technology (cloud, cybersecurity, data & analytics), health services, intelligence analysis and operational support, logistics and sustainment, mission-critical manufacturing, program and engineering technical services, and dynamic contracting solutions to support federal priorities and national security. Cherokee Federal is part of Cherokee Nation Businesses and focuses on mission-focused, U. S. -made solutions.
đ August 4
đď¸ District of Columbia, Virginia, +1 more states â Remote
â° Full Time
đĄ Mid-level
đ Senior
đ Data Scientist
đť Ghost score 11%
Improve your chances of getting an interview by checking your resume score before you apply.

5001 - 10000 employees
Founded 1969
đĽ Healthcare
đŚ Logistics
đ Manufacturing
Healthcare ⢠Logistics ⢠Manufacturing
Cherokee Federal is a U. S. federal systems integrator and government contractor that empowers mission success for more than 60 U. S. federal agencies. With a global workforce of over 5,000, it delivers advanced technology (cloud, cybersecurity, data & analytics), health services, intelligence analysis and operational support, logistics and sustainment, mission-critical manufacturing, program and engineering technical services, and dynamic contracting solutions to support federal priorities and national security. Cherokee Federal is part of Cherokee Nation Businesses and focuses on mission-focused, U. S. -made solutions.
⢠Develop, maintain, and optimize data pipelines using SQL, Python (pandas), and PySpark ⢠Execute and manage notebook-based workflows within Azure Synapse, including debugging and documentation ⢠Process and transform structured and semi-structured data in CSV, JSON/NDJSON, and Parquet formats ⢠Work within ETL/ELT pipelines across raw, curated, and production data layers ⢠Ingest, profile, map, and transform healthcare data from EHR systems and interface feeds while preserving source lineage and clinical context ⢠Perform structured data validation, including row counts, null checks, duplicate detection, schema validation, and allowed value enforcement ⢠Identify and resolve schema drift, inconsistencies, and transformation errors across pipeline stages ⢠Apply repeatable testing and validation practices, reproduce issues, verify fixes, and ensure data reliability ⢠Validate source-to-target mappings and reconcile records across source and destination systems during data conversion and migration ⢠Conduct exploratory data analysis to identify patterns, anomalies, and data quality concerns ⢠Develop derived datasets for reporting, analytics, and downstream data use cases ⢠Collaborate with stakeholders to translate data requirements into usable datasets and metrics ⢠Interpret schemas, column definitions, data types, keys, and table relationships ⢠Manage dataset grain and assess how join strategies affect row counts and outputs ⢠Trace data issues from source ingestion through transformation logic to final outputs ⢠Use logs and debugging approaches to diagnose and resolve pipeline issues ⢠Document transformations, assumptions, mappings, and validation results ⢠Collaborate with engineers, analysts, and stakeholders to ensure data usability, integrity, and requirements alignment ⢠Communicate data issues, findings, and workflow updates with technical team members
⢠Hands-on experience writing SQL queries for data transformation and analysis ⢠Experience using Python, such as pandas, for data processing ⢠Hands-on experience using PySpark for distributed data processing ⢠Experience working within cloud-based data platforms, preferably Azure Synapse or similar ⢠Understanding of ETL/ELT concepts and data pipeline architecture ⢠Experience with structured and semi-structured data formats, including CSV, JSON, and Parquet ⢠Familiarity with Git and collaborative development workflows ⢠Strong problem-solving and debugging skills across data pipelines ⢠Ability to validate and ensure data quality through structured checks and testing practices ⢠Strong written and verbal communication skills
⢠Generous paid time-off ⢠Employee incentive program ⢠Continuous learning culture ⢠Internal Investment Projects (IIP) ⢠Virtual brown-bags/level-ups ⢠Other professional development activities ⢠Recruiting bonuses ⢠3% 401k Safe Harbor contributions ⢠Medical insurance ⢠Dental insurance ⢠Vision insurance ⢠Long-term disability insurance ⢠Short-term disability insurance ⢠AD&D insurance ⢠Life insurance
Apply Nowđ August 4
Computational biology scientist integrating multiomics and AI analyses for GondolaBioâs genetic-disease therapeutics. Supporting target identification, biomarker discovery, and translational research.
đşđ¸ United States â Remote
đľ $155k - $205k / year
đ° $300M Venture Round - GondolaBio on 2024-08
â° Full Time
đ Senior
đ Data Scientist
đ August 4
Senior Data Scientist building AI learning tools and cloud data models for the Veterans Health Administration. Integrating VA systems and advancing responsible AI education across federal healthcare.
đ August 4
Senior Data Scientist developing predictive models and patient-matching algorithms for OneStudyTeamâs clinical-trial platform. Improving enrollment, trial management, model fairness, and compliant healthcare analytics.
đşđ¸ United States â Remote
đľ $140k - $190k / year
â° Full Time
đ Senior
đ Data Scientist
đ August 4
Senior Product Manager owning CRM data associations and record creation for HubSpotâs AI-powered customer platform. Driving cross-platform roadmap, relational data quality, and trustworthy AI experiences across HubSpot products.
đşđ¸ United States â Remote
đľ $140k - $175k / year
â° Full Time
đ Senior
đ Data Scientist
đŚ H1B Visa Sponsor
đ August 4
Senior Data Scientist building applied ML models for SpyCloud, which protects identities from cybercrime. Owning data pipelines, model deployment, monitoring, and cybersecurity feature development.
đşđ¸ United States â Remote
đľ $154k - $200k / year
â° Full Time
đ Senior
đ Data Scientist
đŚ H1B Visa Sponsor