Skip to main content
Starlake Starflow

Test on your laptop.
Ship to your warehouse.

Starflow turns the extract, load, transform, and orchestration boilerplate every data team rewrites into one YAML file per table: you declare what, it generates the how for your warehouse. And because it transpiles your warehouse SQL to DuckDB, the whole pipeline, loads included, runs and tests locally in seconds. Engine choice becomes an environment variable, not a replatforming program.

Apache-2.0 · In production at BPCE Payment Services, Estreem, Axereal, ZE Energy, and Asendia

by hand
-- merge_orders.sql, one warehouse of three
CREATE TEMP TABLE orders_stage AS
SELECT * FROM read_json('incoming/orders_*.json');

-- reject rows with bad types, log them somewhere
-- handle the column that marketing renamed last week

MERGE INTO analytics.orders t
USING orders_stage s ON t.order_id = s.order_id
WHEN MATCHED AND s.order_date > t.order_date
THEN UPDATE SET quantity = s.quantity, ...
WHEN NOT MATCHED THEN INSERT (order_id, ...)

-- plus dag.py: sensors, retries, alerting
-- plus the audit table nobody backfilled
-- now repeat for Snowflake and BigQuery
per table, per warehouse, forever
with Starflow
# metadata/load/starbake/orders.sl.yml
table:
name: orders
pattern: "orders.*.json"
metadata:
format: JSON_FLAT
schedule: "0 * * * *"
writeStrategy:
type: UPSERT_BY_KEY_AND_TIMESTAMP
key: [order_id]
timestamp: order_date
attributes:
- name: order_id
type: long
- name: customer_id
type: long
foreignKey: starbake.customers.id
- name: order_date
type: date
parse, validate, merge, schedule: the whole pipeline
Apache-2.0, for goodA standing public commitment: no BSL, no SSPL, ever.
No telemetryNothing phones home. Verify it in the source.
Your infrastructureRuns entirely where your data lives. EU-sovereignty friendly.
Open standardsArrow Flight SQL and DuckLake, not proprietary protocols.

Test locally. Run anywhere.

Your transforms are written for BigQuery or Snowflake. Starflow transpiles them to DuckDB, so the whole pipeline runs on your laptop. No dev warehouse. No waiting. No bill.

a test is a folder
metadata/tests/transform/sales_kpi/byseller_kpi/test1/
├── sales.orders.csv # input fixture
├── sales.customers.json # input fixture
└── _expected.csv # what the transform must produce
the input you feed in, the output you expect back
run it
starlake test                                  # every load and transform, on local DuckDB
starlake test --transform --domain sales_kpi # just this one
starlake test --site # HTML report with coverage
seconds, in CI or on your machine
🔁

No SQL rewriting

Your warehouse dialect is transpiled to DuckDB, not reimplemented by hand. The SQL you test is the SQL you ship.

📥

Loads are tested too

Parsing, type validation, merge strategy, and rejected rows, not just transforms. The half no SQL-only tool can reach.

🎚️

One project, any engine

SL_ENV=DUCKDB on the laptop, SL_ENV=BQ in production. Same YAML, same SQL, same tests.

See It in Action

Natural-language commands that produce production-ready configurations.

# Bootstrap a new project targeting BigQuery with Airflow
> /bootstrap a new project targeting BigQuery with Airflow orchestration

# Configure ingestion for CSV files
> /load CSV files from GCS into the customers domain with OVERWRITE strategy

# Generate column-level lineage
> /col-lineage for the revenue_summary transform

# Generate Airflow DAGs from your pipeline config
> /dag-generate for all domains using Airflow with daily schedule

# Or use Starflow for a guided lifecycle

# Talk to the data architect persona
> /starflow-data-architect Design a data platform for our e-commerce analytics

# Ask Starflow what to do next based on your project state
> /starflow-help What should I work on next?

The Starlake Stack

One bundle, every layer.

AI Assistants

where you talk to Starlake
Claude CodeGitHub CopilotGemini CLI
↓

Starlake Skills

this bundle
49 CLI skillsStarflow methodology5 expert personas
↓

Orchestration

scheduling and DAGs
AirflowDagster
↓

Data Warehouses & Compute

where your data lives
BigQuerySnowflakeDuckDBPostgreSQLRedshiftDatabricks