K2 Integrity Logo

K2 Integrity

Lead Data Engineer

Posted 6 Days Ago
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
Lead the design and operation of a scalable data platform, owning ingestion, orchestration, observability, lineage, data modeling, contracts, and ETL/ELT pipelines. Build Databricks and PostgreSQL solutions, implement data-quality controls, anomaly detection, governance, privacy, and compliance practices, and establish reusable engineering standards. Mentor engineers and ensure reliable, traceable data across markets and product lines.
The summary above was generated by AI
We are looking for a Lead Data Engineer who can own more than pipelines. Beyond building and maintaining ETL and integrations, you'll own and implement data engineering best practices that make a platform trustworthy at scale: orchestration, observability, lineage, data modeling, and data contracts.
This is a design-forward, lead level role. You'll set the patterns the rest of the team builds on, and you'll be the person who can trace any record end to end, where it came from, every transformation it passed through, and every downstream consumer that depends on it so that when something upstream changes, you know exactly what it touches and can be confident the change propagates cleanly.
What you'll own:
  • Ingestion. Own how data enters the platform across all channels: automated feeds, batch loads, and source-system extracts from CRM and case-management systems. Design ingestion that validates structure and mandatory elements at the door, handles varying formats and volumes across markets and products. You'll build ingestion to be reusable and configurable, so onboarding a new source or market is a known pattern rather than a one-off, and you'll make sure every inbound record carries the lineage and contract metadata the downstream layers depend on.
  • Orchestration. Design the data flow layer that routes datasets through the correct processing sequence and propagates change to everything that depends on it. You'll build dependency-aware refresh, ordered execution, retry and error handling, and clean rollback, with controls on changes to core routing and strong backup/disaster recovery practices. You'll favor versioned, auditable orchestration-as-code in the existing stack over bolt-on tooling.
  • Observability. Stand up the monitoring that proves the platform is healthy and ensure data integrity. This means data-quality checks at key transformation points, deviation monitoring for volume and expectations, automated flagging of anomalies, and exception reporting on breaches. You'll put monitoring on the seams where systems hand off to each other.
  • Lineage. Build native lineage from ingestion through to reporting, so every element is traceable across transformations. You'll make lineage a first-class capability that supports debugging, impact analysis, audit, and compliance.
  • Data modeling. Own the models that underpin the platform: the medallion layers (raw → validated → interim → canonical, with history/SCD handling), canonical entity and master-ID design, and the linkage between datasets. You'll design master-data and naming conventions and build models flexible enough to absorb differing data requirements by market and product rather than locking to a rigid schema. Sound entity/master-ID modeling is central to this role.
  • Data contracts. Establish declared contracts between producers and consumers so interfaces are explicit and changes are safe.
  • Pipelines and integrations. Design, build, and harden ETL/ELT across ingestion channels into the lake house and out to the operational serving layer. Own reliability, performance, and reusability.

Job responsibilities
  • Architect orchestration, observability, lineage, and data contracts as reusable platform capabilities, and document the patterns for the team.
  • Design and evolve logical and physical data models across the platform's layers, balancing consistency with flexibility to onboard new markets and products.
  • Build and harden ingestion and transformation pipelines; manage loads into and out of the operational serving layer.
  • Implement automated data-quality and anomaly detection; define what “healthy” looks like for each key dataset.
  • Establish data best practices: naming conventions, design standards, entity linkage that other services build on.
  • Contribute to governance: PII masking, retention rules, and compliance controls that vary by market.
  • Mentor engineers and raise the bar on rigor, testing, and operational readiness.

Requirements
  • 6+ years in data engineering, with senior-level ownership of production data platforms.
  • Deep hands-on experience with Databricks (Spark, notebooks, Delta Lake, orchestration/workflows, cluster management) and PostgreSQL (complex SQL, stored procedures, materialized views, performance tuning).
  • Proven experience building orchestration for complex, dependency-heavy flows, including dependency-aware refresh, error handling, and rollback.
  • Demonstrated data observability / data quality work: automated checks, anomaly and deviation detection, alerting — with a clear understanding that job success is not the same as correct data.
  • Experience implementing or operating data lineage and a grasp of why traceability matters for debugging, impact analysis, and audit.
  • Strong data modeling across relational and lake-house paradigms: layered/medallion architectures, slowly changing dimensions, and entity/master-data modeling.
  • Strong Python and SQL; solid software-engineering fundamentals (version control, testing, code review, infrastructure-as-code).
  • Experience in regulated or compliance-heavy domains (financial crime, AML, KYC, sanctions screening, risk) preferred.
  • Experience with data contracts and producer/consumer interface design preferred.
  • Familiarity with entity resolution and the downstream propagation challenges it creates preferred.
  • Exposure to data governance and privacy regimes (ISO, GDPR) preferred.
  • Experience designing platforms intended to scale across multiple markets and product lines preferred.

What we value:
  • Rigor over throughput: you build controls and contracts, not just pipelines.
  • Change by design: you plan for the platform to evolve, and make evolution safe.
  • Ownership of correctness — you treat the accuracy of what the platform serves as your responsibility.
  • Clear thinking about interfaces and structure: you close gaps between systems before they become incidents.

Similar Jobs

Yesterday
In-Office or Remote
Nebraska, USA
110K-168K Annually
Senior level
110K-168K Annually
Senior level
Insurance
Lead the technical design and operation of consumer-facing data products on a Databricks lakehouse. Own medallion architecture, dimensional modeling, ODCS data contracts, data quality, observability, CI/CD, incident response, and production standards. Write and review Python and SQL, guide code reviews, mentor engineers, collaborate with product, platform, governance, and ML teams, and ensure compliant, reliable delivery in a regulated insurance environment.
Top Skills: Apache AirflowSparkAzureAzure DevopsCi/CdDagsterDagster CloudDatabricksDatabricks WorkflowsDbtDelta LakeDltGitGreat ExpectationsMlflowMonte CarloPythonSQLUnity Catalog
Yesterday
Remote
California, USA
130K-155K Annually
Senior level
130K-155K Annually
Senior level
Aerospace • Information Technology • Professional Services • Analytics
Leads end-to-end data engineering and analytics platform development, including Power BI dashboards, semantic models, cloud data pipelines, governance, quality monitoring, CI/CD, and legacy workload modernization. Partners with stakeholders to define scalable solutions, validates KPIs, troubleshoots production issues, establishes engineering standards, and mentors team members.
Top Skills: Amazon AthenaAmazon S3Aws GlueAws LambdaAzure Data FactoryAzure DevopsAzure Key VaultCi/CdData Vault 2.0DatabricksDaxDbtDelta LakeDomoGitMicrosoft FabricMicrosoft PurviewPower BIPythonSnowflakeSparkSQLSsisTableau
Yesterday
Remote
United States
Senior level
Senior level
Cloud • Information Technology • Infrastructure as a Service (IaaS)
Leads and develops the Data Engineering team while remaining hands-on with scalable pipelines, ETL/ELT processes, data warehouses, integrations, and infrastructure. Establishes architecture, reliability, quality, and performance standards; guides technical strategy; improves automation and monitoring; and partners with Finance and business stakeholders to translate data needs into practical solutions.
Top Skills: Apache AirflowApache KafkaSparkAWSCi/CdDbtGoogle Cloud PlatformInfrastructure As CodeJavaAzurePythonScalaSQL

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account