Aarden AI Logo

Aarden AI

Sr. Data Platform Engineer

Posted One Month Ago
Hybrid
Seattle, WA, USA
150K-180K Annually
Senior level
Hybrid
Seattle, WA, USA
150K-180K Annually
Senior level
Lead development and operations of geospatial data pipelines and lakehouse systems. Migrate orchestration to Prefect, extend sources and transforms, optimize Postgres and Iceberg stores, build observability and AI-agent readiness, and integrate cross-team tooling with product and ML teams.
The summary above was generated by AI
About Us
Aarden is a land intelligence platform that helps landowners, investors, and developers figure out what a piece of land can actually be used for, and how to market it. We turn messy parcel, infrastructure, market, community, and ecological data into clear, bankable answers for land-dependent assets. Our goal is to become the default decision layer for land: helping physical projects start in places where they can be built and supported for decades.

We’ve built out a suite of data products to support that goal — pipelines, databases, and AI/ML models that power our maps and In-app agents. We have strong product-market fit, and we're now focused on augmenting our data systems. That’s where you come in.

The role
We're looking for a Product-Focused data engineer to maintain and evolve our geospatial data pipeline. Our core data asset is a unique blend of property data, geospatial data, and AI-native derived data. Alongside advocating for data excellence, you’ll be empowered to opportunistically contribute to our user-facing product.

What you’ll do
Pipeline modernization
  • Continue our migration of pipeline orchestration to Prefect
  • Own day-to-day operations of our data infrastructure
  • Extend the pipeline to include new data sources and transformations
  • Maintain, expand and optimize our postgres database and Iceberg datalake
  • Create the 'connective tissue' for data at Aarden
Cross-team integration
  • Partner with product on new feature-driven datasets.
  • Collaborate with the ML/analytics team to close the loop: anomaly detection → ticket → fix → validation → promotion to production
  • Develop cross-team tooling/infra to keep GitHub, Notion, Linear, and Slack connected so pipeline issues, docs, and fixes stay linked
Observability & AI-agent readiness
  • Implement run-over-run data observability (row counts, key column distributions) to catch anomalies and bugs
  • Expose accuracy/quality metrics as first-class artifacts so changes can be evaluated automatically, by a human or an agent
  • Write and maintain AI-context documentation (schema docs, pipeline architecture, known patterns/quirks, "what not to do")

You might be a good fit if you…
Must-have
  • Have strong Python skills & are comfortable with PySpark or similar distributed data processing
  • Have a strong sense of how the data you’re working with impacts the end-user
  • Are curious and excited about AI and the impact it can have on our ways of working as developers
  • Have experience with geospatial data (GeoParquet, PostGIS, Apache Sedona, or similar)
  • Have worked with table formats like Apache Iceberg and lakehouse architectures
  • Have worked on workflow orchestration (Prefect, Airflow, Dagster, or similar)
  • Are comfortable working in a git-based, CI-friendly workflow
Strongly preferred
  • Have worked in full-stack environments, where your work can directly impact the application layer
  • Have experience with Apache Sedona or other cloud spatial-compute platforms
  • Have built observability/logging layers for data pipelines (not just app services)
  • Have experience with property, parcel, real estate, or land data specifically
Nice To have
  • Have experience using AI agents to improve data architecture in a real production codebase
  • Experience in real estate/land, energy, forestry, or agriculture tech

Our Stack
Languages: Python and SQL. TypeScript/Node is a plus for our Application layer and AWS ingest paths.
Orchestration & compute
  • Prefect 3 for pipeline orchestration (YAML/config-driven flows, retries, logging)
  • Coiled for elastic EC2 workers on GDAL-heavy and batch Python jobs
  • Wherobots (managed Apache Sedona / PySpark) for large-scale spatial joins, parcel ingest, and lakehouse work
Data lake & formats: Apache Iceberg, Cloud-Optimized GeoTIFF (COG), and PMTiles. Queried with PySpark and DuckDB.
Geospatial: GDAL, rasterio, GeoPandas, and tippecanoe. Large-scale spatial work runs on Sedona/Spark via Wherobots.
Databases & serving: PostgreSQL + PostGIS (and pgvector on the app side) as the production store.

Working at Aarden
Aarden is a high-trust, high-output team. We’re striving to be intentional about our team growth. This allows us to test the outer boundaries of our individual capabilities, while also going deeper on developer tooling and support. You’ll work hard here, and we’ve got your back.
Practically, this means you’ll be asked to take on large projects, have a high bar of expectations to meet, and have a strong support system to help you meet that high bar. That support system includes:
  • At least 2 in-person days per week at our office in Capitol Hill | We’ve found that while heads-down time at home is fantastic for task-related productivity, in-person time is magic for longer-form productivity. Our in-person days are used to plan, troubleshoot, and check-in with each other on progress and questions. Expect team lunches and whiteboarding.
  • Focused ownership in your role | The rest of the team is here to help you and cares deeply about the long-term functionality of our applications. With that said, we’ll be looking to you to own your lane, go deep, and develop a strong stance on what it takes to make our applications best-in-class.
  • Dedicated monthly AI tooling budget | We’re in a golden era of AI-powered developer tooling. We strongly encourage augmenting your output with AI tools, and have a dedicated & flexible budget for every team member to support that setup. We care about what you ship, not how.
Compensation
The base pay range for this role is $150,000 – $180,000 per year.

Similar Jobs

10 Days Ago
In-Office or Remote
Seattle, WA, USA
180K-240K Annually
Senior level
180K-240K Annually
Senior level
Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Architect and build high-performance distributed infrastructure for batch and streaming data across global regions. Optimize state management, checkpointing, JVM performance, and exactly-once processing. Develop AI infrastructure including vector indexing, agentic workflows, and real-time data streaming. Own the full software development lifecycle, contribute to shared tooling and SDKs, and mentor engineers through technical leadership and design reviews.
Top Skills: Apache FlinkSparkGoJavaJvmKafkaKotlinKubernetesOlap EnginesVector Indexing
5 Days Ago
In-Office or Remote
United States
38K-133K Annually
Senior level
38K-133K Annually
Senior level
Agency • Information Technology
Designs, builds, and operates scalable enterprise data platforms supporting ingestion, processing, storage, governance, analytics, and AI/ML. Responsibilities include developing batch and real-time pipelines, lakehouse and warehouse architectures, orchestration, infrastructure automation, CI/CD, monitoring, security, cost optimization, and observability. The role partners with data, ML, DevOps, security, and application teams, troubleshoots complex platform issues, establishes reusable engineering standards, contributes to architecture strategy, and mentors engineers.
Top Skills: Access ControlAlertingCi/CdCloud Data PlatformsData GovernanceData LakesData OrchestrationData PipelinesData SecurityData WarehousesDevOpsDistributed SystemsEncryptionInfrastructure As CodeLakehousesLoggingMonitoringObservabilityPlatform EngineeringStreamingWorkflow Automation
6 Days Ago
In-Office
Seattle, WA, USA
Senior level
Senior level
Artificial Intelligence • Robotics
Build and operate scalable data-platform systems for multimodal robotics and machine-learning datasets. Responsibilities include ingestion, processing, validation, labeling, dataset generation, metadata, lineage, versioning, data quality, observability, and workflow reliability. Partner with ML, robotics, infrastructure, and labeling teams to create reusable platform capabilities supporting large-scale data processing, reprocessing, backfills, training, and evaluation.
Top Skills: AirflowC++DagsterGoJavaKubernetesParquetPythonRayS3Spark

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account