A self-hosted data pipeline engine for moving and transforming files with DuckDB.
Dataflow gives teams a focused control plane for file-based data pipelines: define a source, transform it with SQL or Python, validate the result, and write Parquet or Delta Lake output. It runs as a single self-hosted service with a web UI, durable jobs, and scheduled execution.
- Simple to operate: FastAPI, DuckDB, SQLite, and no frontend build step.
- Built for files: local paths, S3-compatible storage, HTTPS, CSV, JSON, and Parquet inputs.
- Reliable by default: durable job queue, retries, checkpoints, worker isolation, and data-quality checks before writes.
- Flexible execution: SQL transformations, Python plugins, and an ad-hoc DuckDB query tool.
- Self-hosted: keep data, configuration, and authentication under your control.
git clone https://github.com/vmskonakanchi/dataflow.git
cd dataflow
uv sync
uv run python main.py server --reloadOpen http://localhost:8000 and complete the initial administrator setup.
For Docker, production setup, and your first pipeline, start with the documentation.
- Getting Started
- Core Concepts
- Template Variables and Timezones
- Scheduling
- Sources
- Transformations
- Sinks
- Examples
- Enterprise Operations
- Troubleshooting
Contributions are welcome. Read CONTRIBUTING.md for local development, testing, migrations, and project conventions.
