Data Engineer – Onboarding

🔥 9 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Sardine

Sardine

51 - 200 employees

Founded 2020

🔒 Cybersecurity

📋 Compliance

💳 Fintech

Cybersecurity • Compliance • Fintech

Sardine is a cutting-edge platform focused on fraud prevention and compliance. The company offers a behavior-based system for fraud detection, identity verification, and transaction monitoring that helps leading banks, online retailers, and fintechs protect themselves from scams and financial crime. Sardine's technology integrates advanced behavioral biometrics and device intelligence to combat identity fraud, payment fraud, and account takeovers. The platform also streamlines compliance checks such as KYC (Know Your Customer) and AML (Anti-Money Laundering) transaction monitoring. Sardine's comprehensive solution empowers its clients to automate risk decisioning, identify high-risk users early, and manage fraud across the customer journey effectively.

📋 Description

• Own the data ingestion layer that brings device telemetry, transaction events, KYC/identity signals, and third-party enrichment into the platform — designing streaming pipelines (Pub/Sub, Apache Beam on Dataflow, Flink) and batch pipelines (Python, Airflow on Cloud Composer, Spark on Dataproc) that are correct, observable, and cheap to extend. • Build and evolve our feature platform, where the same Chronon feature definitions are computed by Flink for streaming and Spark for batch, with aggregation windows from one hour to 300 days, served to the rules engine and to models under a sub-second budget. • Establish feature correctness as an engineering discipline: streaming-versus-batch reconciliation, recomputation tests against the warehouse, train/serve parity checks, and drift monitoring that catches a broken feature before an analyst does. • Productionize fraud and identity ML models — training pipelines on Vertex AI and Kubeflow, gradient-boosted and tree-based models (XGBoost, LightGBM, CatBoost, scikit-learn), hyperparameter search, SHAP-based explanations, and score normalization — and build the automated retraining, champion/challenger promotion, and rollback machinery we don't yet have. • Engineer KYC, AML, and identity risk signals: document verification and doc-KYC outcomes, sanctions/PEP/adverse-media screening results, email and phone risk, synthetic identity indicators, bank and account verification, and periodic customer due diligence — turning noisy, multi-vendor, multi-jurisdiction data into features a model can actually learn from. • Integrate and harden new data sources, including 30+ third-party enrichment providers called in parallel on the request path, plus our cross-client consortium network — owning failover behavior, timeout budgets, graceful degradation, caching, and cost. • Own the warehouse and modeling layer in BigQuery — partitioning strategy, the staging-to-mart layer cake, training datasets, and the in-flight migration off dbt onto scheduled SQL and Python pipelines. • Design the entity resolution and graph data that link customers, devices, emails, phones, cards, bank accounts, and crypto addresses across clients, including large-scale connected-components work. • Make the platform safe by construction: field-level encryption for sensitive identifiers, regional data residency enforced in the pipeline definitions, PII handling and deletion paths, and feature-level gating so a bad signal can be turned off without a deploy. • Set technical direction and raise the team's ceiling — write the design docs, run the reviews, mentor engineers and data scientists, and decide what we build versus buy.

🎯 Requirements

• 8+ years building production data and ML systems, with real ownership of both the pipeline side and the model side. You have shipped models that made consequential automated decisions, not just dashboards. • Deep Python and strong SQL. You are fluent in a distributed processing framework (Spark, Beam, or Flink) and comfortable reasoning about streaming semantics — windowing, watermarks, late data, exactly-once versus at-least-once, and where correctness actually breaks. • Hands-on experience with a modern cloud data stack: GCP strongly preferred (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable, Composer, Vertex AI) or the AWS equivalents, plus Docker, Kubernetes, Terraform, and CI/CD. • Practical ML engineering depth: feature stores and feature pipelines, training/serving skew, gradient-boosted tree models, class imbalance and rare-event modeling, threshold and cost-sensitive tuning, model monitoring and drift detection, and explainability. • Experience with high-volume, low-latency serving where a feature fetch has a few hundred milliseconds and there is no retry budget. • Domain experience in fraud, risk, payments, lending, or identity/KYC — or the demonstrated ability to get fluent in a regulated domain fast. You understand why label latency, feedback loops, and adversarial drift make fraud modeling different from ordinary supervised learning. • Comfort with data governance in a regulated environment: PII, encryption, access control, regional data residency, auditability. • Strong written communication. You can explain a modeling tradeoff to a fraud analyst and a pipeline design to a backend engineer, and you write things down. • A bias toward action and comfort in ambiguity. Much of this role is deciding what should exist, then building it.

🏖️ Benefits

• Generous compensation in cash and equity • Early exercise for all options, including pre-vested • Work from anywhere: Remote-first Culture • Flexible paid time off and Year-end break • Health insurance, dental, and vision coverage for employees and dependents - *US and Canada specific* • 4% matching in 401k / RRSP - *US and Canada specific* • MacBook Pro delivered to your door • One-time stipend to set up a home office — desk, chair, screen, etc. • Monthly meal stipend • Monthly social meet-up stipend • Annual health and wellness stipend • Annual Learning stipend

Apply Now

Similar Jobs

🔥 19 minutes ago

Danaher Corporation

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

Solutions Architect handling sales and technical support for medical imaging products at Leica Biosystems. Achieving sales goals while ensuring customer satisfaction through product demonstrations and support.

🔥 28 minutes ago

CVS Health

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

🛒 Retail

Senior Manager overseeing engineers to enhance member experience at CVS Health. Leading technical solutions and mentoring high-performing engineering teams in healthcare contexts.

🔥 29 minutes ago

GCG

1001 - 5000

📦 Logistics

🏭 Manufacturing

🏛️ Government

Inside Sales Account Manager focused on developing B2B relationships in Data Center solutions. Engaging with customers to support strategic growth and sales initiatives.

🔥 38 minutes ago

Databricks

1001 - 5000

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Solutions Architect working with large retail and hospitality clients. Leading technical strategies for Databricks products and mentoring team members.

🔥 47 minutes ago

SYNCREON

10,000+ employees

🚘 Automotive

📦 Logistics

🚗 Transport

Solutions Engineer curating media datasets from ingestion to delivery for a recruitment staffing solutions company. Collaborating with cross-functional teams to enhance data handling.