Skip to content
vmskonakanchiPublic

About

Config-driven data pipeline platform

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

Dataflow

A self-hosted data pipeline engine for moving and transforming files with DuckDB.

Dataflow dashboard

Dataflow gives teams a focused control plane for file-based data pipelines: define a source, transform it with SQL or Python, validate the result, and write Parquet or Delta Lake output. It runs as a single self-hosted service with a web UI, durable jobs, and scheduled execution.

Why Dataflow

  • Simple to operate: FastAPI, DuckDB, SQLite, and no frontend build step.
  • Built for files: local paths, S3-compatible storage, HTTPS, CSV, JSON, and Parquet inputs.
  • Reliable by default: durable job queue, retries, checkpoints, worker isolation, and data-quality checks before writes.
  • Flexible execution: SQL transformations, Python plugins, and an ad-hoc DuckDB query tool.
  • Self-hosted: keep data, configuration, and authentication under your control.

Quick Start

git clone https://github.com/vmskonakanchi/dataflow.git
cd dataflow
uv sync
uv run python main.py server --reload

Open http://localhost:8000 and complete the initial administrator setup.

For Docker, production setup, and your first pipeline, start with the documentation.

Documentation

Contributing

Contributions are welcome. Read CONTRIBUTING.md for local development, testing, migrations, and project conventions.

Project Information

About

Config-driven data pipeline platform

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages