Sardine Logo

Sardine

Data Engineer - Onboarding

Posted Yesterday
Remote
Hiring Remotely in USA
150K-205K Annually
Senior level
Remote
Hiring Remotely in USA
150K-205K Annually
Senior level
Lead end-to-end data and ML pipelines for onboarding and fraud decisions: ingest streaming and batch sources, build feature platform and entity resolution, productionize ML models, enforce PII/data governance, and set technical direction while mentoring engineers and ensuring low-latency, reliable serving.
The summary above was generated by AI

Who we are:

Sardine is the leading agentic risk platform for fighting financial crime. Our integrated solution unifies data across risk teams to help organizations stop fraud in real time, prevent AI-driven attacks, and automate fraud and AML operations. Sardine’s platform is strengthened by one of the fastest-growing fraud consortiums in the market, spanning more than 6 billion profiled devices, 800 million consumers, and 3 million businesses worldwide. Leading companies including FIS, GoDaddy, Intuit, Edward Jones, ZoomInfo, and Checkout.com rely on Sardine to secure and grow trust in their products.

Our culture:

  • We have hubs in the Bay Area, NYC, Austin, Toronto, and São Paulo. However, we maintain a remote-first work culture. #WorkFromAnywhere

  • We hire talented, self-motivated individuals with extreme ownership and high growth orientation.

  • We value performance and not hours worked. We believe you shouldn't have to miss your family dinner, your kid's school play, friends get-together, or doctor's appointments for the sake of adhering to an arbitrary work schedule.

Location:

  • Remote - United States or Canada

  • From Home / Beach / Mountain / Cafe / Anywhere!

  • We are a remote-first company with a globally distributed team. You can find your productive zone and work from there.

About the role

We are looking for a Senior Data/ML Engineer to own the data and machine learning foundation that Sardine's compliance decisions run on. Every onboarding decision we make — a payment approved, an account blocked, a KYC case escalated — is the output of a pipeline someone built. This role owns those pipelines end to end: how data arrives, how it becomes a feature, how that feature becomes a model, and how that model stays correct in production.

This is a high-impact, highly technical IC role sitting at the intersection of data engineering and ML engineering. We need someone at the senior level to set technical direction for the next order of magnitude: new feature generation, build specific models around KYC onboarding, in house entity matcher for the sanctions and more

You will write production code, make architectural calls that outlive your tenure, and raise the bar for how a small team ships fraud ML. You will work directly with data scientists, backend engineers, and the fraud analysts who use what you build.

What you'll be doing

  • Own the data ingestion layer that brings device telemetry, transaction events, KYC/identity signals, and third-party enrichment into the platform — designing streaming pipelines (Pub/Sub, Apache Beam on Dataflow, Flink) and batch pipelines (Python, Airflow on Cloud Composer, Spark on Dataproc) that are correct, observable, and cheap to extend.

  • Build and evolve our feature platform, where the same Chronon feature definitions are computed by Flink for streaming and Spark for batch, with aggregation windows from one hour to 300 days, served to the rules engine and to models under a sub-second budget.

  • Establish feature correctness as an engineering discipline: streaming-versus-batch reconciliation, recomputation tests against the warehouse, train/serve parity checks, and drift monitoring that catches a broken feature before an analyst does.

  • Productionize fraud and identity ML models — training pipelines on Vertex AI and Kubeflow, gradient-boosted and tree-based models (XGBoost, LightGBM, CatBoost, scikit-learn), hyperparameter search, SHAP-based explanations, and score normalization — and build the automated retraining, champion/challenger promotion, and rollback machinery we don't yet have.

  • Engineer KYC, AML, and identity risk signals: document verification and doc-KYC outcomes, sanctions/PEP/adverse-media screening results, email and phone risk, synthetic identity indicators, bank and account verification, and periodic customer due diligence — turning noisy, multi-vendor, multi-jurisdiction data into features a model can actually learn from.

  • Integrate and harden new data sources, including 30+ third-party enrichment providers called in parallel on the request path, plus our cross-client consortium network — owning failover behavior, timeout budgets, graceful degradation, caching, and cost.

  • Own the warehouse and modeling layer in BigQuery — partitioning strategy, the staging-to-mart layer cake, training datasets, and the in-flight migration off dbt onto scheduled SQL and Python pipelines.

  • Design the entity resolution and graph data that link customers, devices, emails, phones, cards, bank accounts, and crypto addresses across clients, including large-scale connected-components work.

  • Make the platform safe by construction: field-level encryption for sensitive identifiers, regional data residency enforced in the pipeline definitions, PII handling and deletion paths, and feature-level gating so a bad signal can be turned off without a deploy.

  • Set technical direction and raise the team's ceiling — write the design docs, run the reviews, mentor engineers and data scientists, and decide what we build versus buy.

What you'll need

  • 8+ years building production data and ML systems, with real ownership of both the pipeline side and the model side. You have shipped models that made consequential automated decisions, not just dashboards.

  • Deep Python and strong SQL. You are fluent in a distributed processing framework (Spark, Beam, or Flink) and comfortable reasoning about streaming semantics — windowing, watermarks, late data, exactly-once versus at-least-once, and where correctness actually breaks.

  • Hands-on experience with a modern cloud data stack: GCP strongly preferred (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable, Composer, Vertex AI) or the AWS equivalents, plus Docker, Kubernetes, Terraform, and CI/CD.

  • Practical ML engineering depth: feature stores and feature pipelines, training/serving skew, gradient-boosted tree models, class imbalance and rare-event modeling, threshold and cost-sensitive tuning, model monitoring and drift detection, and explainability.

  • Experience with high-volume, low-latency serving where a feature fetch has a few hundred milliseconds and there is no retry budget.

  • Domain experience in fraud, risk, payments, lending, or identity/KYC — or the demonstrated ability to get fluent in a regulated domain fast. You understand why label latency, feedback loops, and adversarial drift make fraud modeling different from ordinary supervised learning.

  • Comfort with data governance in a regulated environment: PII, encryption, access control, regional data residency, auditability.

  • Strong written communication. You can explain a modeling tradeoff to a fraud analyst and a pipeline design to a backend engineer, and you write things down.

  • A bias toward action and comfort in ambiguity. Much of this role is deciding what should exist, then building it.

Bonus points for

  • Experience supporting customer-facing ML — bring-your-own-model integrations, model explainability for adverse action or regulatory review, or shadow/challenger scoring frameworks.

  • Experience in high-growth B2B SaaS, or as an early data/ML hire who built the function rather than inherited it.

Benefits we offer:

  • Generous compensation in cash and equity

  • Early exercise for all options, including pre-vested

  • Work from anywhere: Remote-first Culture

  • Flexible paid time off and Year-end break

  • Health insurance, dental, and vision coverage for employees and dependents - US and Canada specific

  • 4% matching in 401k / RRSP - US and Canada specific

  • MacBook Pro delivered to your door

  • One-time stipend to set up a home office — desk, chair, screen, etc.

  • Monthly meal stipend

  • Monthly social meet-up stipend

  • Annual health and wellness stipend

  • Annual Learning stipend

Join a fast-growing company with world-class professionals from around the world. If you are seeking a meaningful career, you found the right place, and we would love to hear from you.

To learn more about how we process your personal information and your rights in regards to your personal information as an applicant and Sardine employee, please visit our Applicant and Worker Privacy Notice.

Similar Jobs

2 Hours Ago
Remote or Hybrid
90K-135K Annually
Entry level
90K-135K Annually
Entry level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Detect, contain, and remediate security incidents across Windows, macOS, and Linux. Perform malware and forensic analysis, develop detection and remediation processes, produce customer-facing reports and recommendations, and contribute to public thought leadership. Use scripting/programming and AI tools to enhance investigations and response.
Top Skills: .NetAi TechnologiesCC#Forensic Analysis ToolsLinuxmacOSMalware AnalysisNetwork Analysis ToolsPerlPythonRuby On RailsVbWindows
3 Hours Ago
Easy Apply
Remote
Easy Apply
191K-191K Annually
Senior level
191K-191K Annually
Senior level
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Lead design and delivery of reliability projects for Coinbase's platform: secure service configuration and secrets management, improve canary-based deployments, partner with core services to increase scalability and reduce incidents, drive reliability best practices, and participate in on-call rotations.
Top Skills: AWSAzureDatadogGCPGoKibanaRubyTerraform
4 Hours Ago
In-Office or Remote
CA, USA
164K-297K Annually
Senior level
164K-297K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Lead end-to-end delivery of high-priority revenue programs, coordinating Product, Sales, Marketing, Finance, Legal, Risk, and Operations. Drive launch readiness for products and partnerships, establish governance and operating cadences, create repeatable launch frameworks and executive communications, and continuously improve cross-functional execution.

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account