Senior Data Engineer – Design, Architecture

🕒 vor 1 Monat

🇺🇸 Vereinigte Staaten – Remote

💵 $120.000 - $140.000 / Jahr

⏰ Vollzeit

🟠 Senior

🚰 Dateningenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

AWS

Cloud

Cyber Security

EC2

ETL

Pandas

Postgres

PySpark

Python

Spark

Unity

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of K2United

K2United

51 - 200 Mitarbeiter

Als Knowledge2Share vor 22 Jahren gegründet wurde, trieben leidenschaftliche Individuen das Startup mit ihrem Innovationsdrang voran. Seitdem hat sich das Unternehmen in zwei Firmen entwickelt: K2Share, eine Cybersecurity-Beratung, und CareerSafe, ein führendes Unternehmen im Bereich Online-OSHA-Training und karriereorientierte Bildung. Während unsere Marken K2Share und CareerSafe jeweils repräsentieren, was wir tun (und das wirklich gut, wenn wir das hinzufügen dürfen!), stellt keine von beiden vollständig dar, wer wir sind. Im Kern sind wir genau das, was unser Gründer, Dr. Larry Teverbaugh, schaffen wollte: einen großartigen Arbeitsplatz, an dem Mitarbeiter Spaß haben können, während sie etwas bewirken. Unter unserer neuen Organisation, K2United, können unsere beiden Unternehmen weiter wachsen und sich entwickeln, vereint durch das, was uns verbindet: unser Ziel. Gemeinsam schaffen wir Lösungen, damit diejenigen, denen wir dienen, gedeihen.

Beschreibung

• The Senior Data Engineer will own the data engineering function on K2Share's Federal Team, partnering with technical and product leadership to deliver data products that support mission-critical decision-making for federal agency clients. • Design and build relational data layers that handle OSCAL and other structured compliance data - including ingestion, validation, transformation, and export workflows that preserve fidelity to source schemas across the full data lifecycle • Design and maintain data models that support governance, risk, compliance, scoring, and reporting workflows for federal cybersecurity programs, with OSCAL as the connective layer across them - including long-term retention and archival policies that align with federal recordkeeping and audit requirements • Design and build big-data processing pipelines on Databricks (PySpark, Delta Lake, Unity Catalog) that normalize cybersecurity data from across federal agency environments and produce analytical layers for trend analysis, executive reporting, and cross-program insights • Optimize data systems for performance and cost - identifying I/O and compute bottlenecks, scaling compute responsibly, and balancing throughput against the cost discipline federal engagements require • Architect, build and maintain AWS data infrastructure that meets federal security and operational requirements - working across services such as S3, Bedrock, Lambda, Fargate, and EC2 in support of compliance and analytical workloads • Design and implement audit-ready data primitives - change capture, access controls, validation, and lineage — that support agency reporting and continuous monitoring needs • Lead AI-first development and responsible AI deployment on the data team - using AI development tools as a standard part of the engineering loop, prototyping AI-assisted compliance workflows, and designing the production AI systems behind them (RAG architectures, vector store management, conversational agents, prompt and output guardrails, and evaluation pipelines), in alignment with federal AI governance guidance (OMB, NIST AI RMF) • Engage with federal agency stakeholders and internal teams during requirement discovery, delivery, and ongoing support — translating compliance needs into data products and customer feedback into improvements

🎯 Anforderungen

• 5+ years of production data engineering experience, with a track record of designing and owning data systems end-to-end • Strong relational database expertise with PostgreSQL or equivalent — schema design at scale, indexing and partitioning strategy, access control, and patterns for handling semi-structured data including JSON. • Strong system design and architecture instincts - able to translate business and compliance requirements into data system designs, document tradeoffs, and lead design reviews with technical and non-technical stakeholders. • 3+ years of big-data processing on Databricks, Spark, or equivalent — PySpark, Delta Lake, Unity Catalog, and medallion (bronze/silver/gold) architecture patterns • Strong AWS experience including S3, Bedrock, Lambda, Fargate, EC2, relational database services, change-data-capture services, serverless compute, IAM, and KMS — ideally in GovCloud or other regulated-cloud environments • Strong Python development skills, including data manipulation libraries such as pandas, for ETL, transformation, and analytical workflows • Proficient with Git and modern version-control practices — branching strategies, code review discipline, and collaborative workflows in a team setting • Experience working with structured external schemas - OSCAL or similar standards-based data — including the discipline of preserving fidelity through transformation • Demonstrated focus on data system optimization - identifying I/O bottlenecks, remediating performance issues, balancing cost and performance, and scaling compute responsibly • Schema evolution discipline - migration strategy, backward compatibility, change-data-capture-friendly design, and the operational rigor of running production schemas under change control • Experience with async data processing patterns — task queues, message-based pipelines, idempotent task design • Active use of AI development tools as a routine part of the engineering workflow, with informed views on where they accelerate and where they need supervision • Familiarity with responsible AI deployment patterns — RAG architectures, vector databases, embedding management, prompt and output guardrails, and evaluation methods • Working knowledge of federal cybersecurity frameworks: FISMA, NIST RMF, NIST SP 800-53, NIST CSF • Demonstrated ability to interpret regulatory and policy guidance and translate it into product or data-product requirements • Comfort working across ambiguous, fast-moving federal programs with minimal supervision and strong collaborative instincts

🏖️ Vorteile

• 401(k) plan with employer matching contributions • Low-cost, comprehensive medical benefits for employees and their families • Flexibility for those needing time off for jury duty, voting, military leave, etc. • Paid time off • Wellness stipend program (includes fitness reimbursement program) • Tuition stipend • Casual dress work environment • Technical training and certifications as required • Any of our CareerSafe Online training courses for free to employees and their immediate family

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

Marqeta

501 - 1000

💳 Fintech

🤝 B2B

Senior Staff Software Engineer at Marqeta, building and operating the data platform infrastructure. Collaborating with teams to ensure reliable data pipelines and architect technical solutions.

🇺🇸 Vereinigte Staaten – Remote

💵 $200.500 - $275.000 / Jahr

💰 Post-IPO Equity im 2021-06

⏰ Vollzeit

🟠 Senior

🚰 Dateningenieur

🦅 H1B-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Setton Industries Inc.

1 - 10

🚀 Luft- und Raumfahrt

Senior Data Engineer building Scorpion's analytical data platform to support AI products. Collaborating with teams to design scalable systems and ensure data quality standards.

🇺🇸 Vereinigte Staaten – Remote

💵 $155.000 - $185.000 / Jahr

⏰ Vollzeit

🟠 Senior

🚰 Dateningenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Zocdoc

501 - 1000

⚕️ Krankenversicherung

🏪 Marktplatz

🧘 Wellness

Senior Staff Engineer focusing on analytics and data infrastructure at Zocdoc. Leading initiatives and enhancing data systems to serve stakeholders' needs.

🇺🇸 Vereinigte Staaten – Remote

💵 $235.000 - $300.000 / Jahr

💰 €150.000.000 Private Equity Round im 2021-02

⏰ Vollzeit

🟠 Senior

🚰 Dateningenieur

🦅 H1B-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich

Airflow

Amazon Redshift

BigQuery

Python

SQL

🕒 vor 1 Monat

LeafLink

201 - 500

🛍️ eCommerce

🏪 Marktplatz

🤝 B2B

Senior Data Engineer leading the development of scalable data solutions at LeafLink, the premier B2B cannabis platform. Collaborating with cross-functional teams to enhance analytics and operational insights.

🇺🇸 Vereinigte Staaten – Remote

💵 $125.000 - $155.000 / Jahr

💰 €100.000.000 Series D im 2023-02

⏰ Vollzeit

🟠 Senior

🚰 Dateningenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

Airflow

Amazon Redshift

AWS

Python

SQL

🕒 vor 1 Monat

MCG Health

201 - 500

⚕️ Krankenversicherung

☁️ SaaS

🏢 Unternehmen

Senior Product Manager leading data platform direction and strategy for MCG Health. Overseeing data usability across payer, provider, and partner environments.

🇺🇸 Vereinigte Staaten – Remote

💵 $150.000 - $210.000 / Jahr

⏰ Vollzeit

🟠 Senior

🚰 Dateningenieur

🦅 H1B-Visum-Sponsor

info

🗣️🇺🇸🇬🇧 Englisch erforderlich