Test on your laptop.
Ship to your warehouse.
Starflow turns the extract, load, transform, and orchestration boilerplate every data team rewrites into one YAML file per table: you declare what, it generates the how for your warehouse. And because it transpiles your warehouse SQL to DuckDB, the whole pipeline, loads included, runs and tests locally in seconds. Engine choice becomes an environment variable, not a replatforming program.
Apache-2.0 · In production at BPCE Payment Services, Estreem, Axereal, ZE Energy, and Asendia
-- merge_orders.sql, one warehouse of three
CREATE TEMP TABLE orders_stage AS
SELECT * FROM read_json('incoming/orders_*.json');
-- reject rows with bad types, log them somewhere
-- handle the column that marketing renamed last week
MERGE INTO analytics.orders t
USING orders_stage s ON t.order_id = s.order_id
WHEN MATCHED AND s.order_date > t.order_date
THEN UPDATE SET quantity = s.quantity, ...
WHEN NOT MATCHED THEN INSERT (order_id, ...)
-- plus dag.py: sensors, retries, alerting
-- plus the audit table nobody backfilled
-- now repeat for Snowflake and BigQuery
# metadata/load/starbake/orders.sl.yml
table:
name: orders
pattern: "orders.*.json"
metadata:
format: JSON_FLAT
schedule: "0 * * * *"
writeStrategy:
type: UPSERT_BY_KEY_AND_TIMESTAMP
key: [order_id]
timestamp: order_date
attributes:
- name: order_id
type: long
- name: customer_id
type: long
foreignKey: starbake.customers.id
- name: order_date
type: date
One declaration, four stages
The same YAML drives every step, on DuckDB, BigQuery, Snowflake, Redshift, PostgreSQL, or Spark.
Extract
Pull schemas and data from JDBC databases, REST APIs, and OpenAPI specs into versioned YAML.
Tutorial →Load
Ingest CSV, JSON, XML, and Parquet with type validation, rejection reports, and merge strategies.
Tutorial →Transform
Write plain SQL; Starflow resolves dependencies and generates the right MERGE for each engine.
Tutorial →Orchestrate
Dependency-ordered DAGs generated for Airflow, Dagster, or Snowflake Tasks. No hand-written graphs.
Tutorial →Test locally. Run anywhere.
Your transforms are written for BigQuery or Snowflake. Starflow transpiles them to DuckDB, so the whole pipeline runs on your laptop. No dev warehouse. No waiting. No bill.
metadata/tests/transform/sales_kpi/byseller_kpi/test1/
├── sales.orders.csv # input fixture
├── sales.customers.json # input fixture
└── _expected.csv # what the transform must produce
starlake test # every load and transform, on local DuckDB
starlake test --transform --domain sales_kpi # just this one
starlake test --site # HTML report with coverage
No SQL rewriting
Your warehouse dialect is transpiled to DuckDB, not reimplemented by hand. The SQL you test is the SQL you ship.
Loads are tested too
Parsing, type validation, merge strategy, and rejected rows, not just transforms. The half no SQL-only tool can reach.
One project, any engine
SL_ENV=DUCKDB on the laptop, SL_ENV=BQ in production. Same YAML, same SQL, same tests.
Two Ways to Use the Skills
Greenfield project or migration? Start with Starflow for the full lifecycle. Quick targeted task? Use a CLI skill directly.
Starflow
Guided methodology layer. Five expert personas (Lea, Winston, Amelia, Quinn, Max) walk you through Discovery → Architecture → Pipeline Design → Implementation, with adversarial code review and end-of-epic retrospectives.
Open the Starflow guide →
Direct CLI Skills
One skill per Starlake command: load, transform, extract, dag-generate, and 45 more. Ask in natural language; get production-ready YAML, SQL, or shell.
Browse the catalog →
Guided methodology, five expert personas
Four phases, five expert personas, persistent step-file workflows that resume across sessions.
1. Discovery
Map data domains, sources, and ownership before writing any configuration.
- starflow-domain-discovery
- starflow-source-analysis
2. Architecture
Design the platform, layers, engines, and table schemas that will support your pipelines.
- starflow-create-data-architecture
- starflow-schema-design
3. Pipeline Design
Specify pipelines end-to-end (extract, load, transform, orchestrate) before implementation.
- starflow-create-pipeline-spec
- starflow-transform-design
- starflow-orchestration-design
4. Implementation
Build, review, deploy, and reflect. Adversarial parallel code review and end-of-epic retros.
- starflow-sprint-planning
- starflow-dev-pipeline
- starflow-code-review
- starflow-retrospective
Plus five agent personas (Lea, Winston, Amelia, Quinn, Max) covering data analysis, architecture, engineering, quality, and platform; and the cross-cutting data-quality-review, lineage-review, and adaptive starflow-help skills.
Skill Catalog
49 skills across 11 categories, one per Starlake CLI command, with the configuration patterns to match.
Ingestion & Loading
8 skills- autoload
- load
- cnxload
- esload
- kafkaload
- ingest
- preload
- stage
Transformation
2 skills- transform
- job
Extraction
7 skills- extract
- extract-schema
- extract-data
- extract-bq-schema
- extract-rest-schema
- extract-rest-data
- extract-script
Schema Management
6 skills- bootstrap
- infer-schema
- xls2yml
- xls2ymljob
- yml2ddl
- yml2xls
Data Quality
1 skills- expectations
Lineage
4 skills- lineage
- col-lineage
- table-dependencies
- acl-dependencies
Orchestration
2 skills- dag-generate
- dag-deploy
Operations
7 skills- validate
- metrics
- freshness
- console
- serve
- settings
- migrate
Security
2 skills- secure
- iam-policies
Configuration
2 skills- config
- connection
Utilities
6 skills- bq-info
- compare
- parquet2csv
- site
- summarize
- test
See It in Action
Natural-language commands that produce production-ready configurations.
# Bootstrap a new project targeting BigQuery with Airflow
> /bootstrap a new project targeting BigQuery with Airflow orchestration
# Configure ingestion for CSV files
> /load CSV files from GCS into the customers domain with OVERWRITE strategy
# Generate column-level lineage
> /col-lineage for the revenue_summary transform
# Generate Airflow DAGs from your pipeline config
> /dag-generate for all domains using Airflow with daily schedule
# Or use Starflow for a guided lifecycle
# Talk to the data architect persona
> /starflow-data-architect Design a data platform for our e-commerce analytics
# Ask Starflow what to do next based on your project state
> /starflow-help What should I work on next?The Starlake Stack
One bundle, every layer.