Datapipe
Agent-Driven ETL Pipeline Builder
A self-hosted platform where AI agents research an API, write and validate the pipeline, and keep it healthy, so a new data source goes from a name to a running sync into your warehouse without hand-written connectors.
- Role
- Sole engineer — design, build, deploy
- Timeline
- 2026
Why teams use it
- From an app name to a running pipeline. Pick a connector from the catalog, or type the name of any API: a research agent reads its documentation, finds the authentication, endpoints and incremental filters, and proposes the setup.
- Pipelines are tested before they ship. A codegen agent writes the Python pipeline, loads a small sample into a test dataset and validates it against your mapping, retrying with the error in hand until it passes.
- It keeps working when sources change. New fields, disappeared fields, row-count drift and volume anomalies are detected automatically. When a run fails, an ops agent writes a diagnosis with a root cause and a recommended fix.
- Real-time when you need it. Postgres change data capture with resumable, chunked snapshots and a live progress bar, then continuous streaming a few seconds behind the source.
- Your warehouse, your infrastructure. It syncs into BigQuery, Postgres, Snowflake or DuckDB and runs on your own machine or cloud. DuckDB writes to a local file, so you can try it with no cloud account.
- One console for the whole fleet. Run health, lag, replication slots, queues and alerts in a single operations view.
All screenshots below use synthetic data: sources, tables and figures are invented.
Product tour
Your whole fleet at a glance
Rows loaded, runs, capture lag and replication-slot size at the top. Below, every source sorted by attention, with its mode, status, last run, lag and next run, a live progress bar for a snapshot in flight, and the alerts and workers beside it.
A catalog of pre-researched connectors
Ready-to-use connectors for common SaaS tools and databases, each with its authentication type and endpoint count. Anything not listed, the research agent can learn from its API documentation.

Map fields to your warehouse, visually
A five-step wizard walks from app and credentials to endpoints, fields and mapping. Pick the fields, let the agent propose the column mapping, and add transforms such as converting cents to dollars. Tables without a primary key are flagged before they cause trouble.

Generated, tested and validated
The agent writes the pipeline, test-runs a sample into a scratch dataset and validates it. In this run the first attempt failed on a type mismatch, the agent applied a transform and the second attempt passed. The generated code is there to read.

Change data capture, with the numbers that matter
Capture lag, pending changes and replication-slot size for a streaming Postgres source, a run history, and a panel for schema and data changes with row-count checks between source and destination.

Every run on one timeline
A lane per source, runs coloured by outcome, queued runs hatched and a line for now. Select a run to see its duration, rows, attempts, throughput and per-table progress.

Failures are diagnosed, not just reported
When a run fails, the ops agent says what happened, classifies the root cause with a confidence level and recommends a fix. A human clicks Apply fix, or turns on auto-heal per source.

Under the hood
- Research agent. Finds the authentication type, base URL and endpoints with documentation links, and detects incremental filters.
- Codegen agent. Writes the Python pipeline from the chosen endpoints and mapping, validating against a sample load with up to three attempts.
- Drift and ops agents. Detect schema and data changes, diagnose failures into a root-cause category and propose a fix. Healing can be automatic or approved by a person.
- Change data capture. Postgres logical replication with chunked, resumable snapshots, optional parallelism and a BigQuery native apply mode or MERGE. CDC pipelines are rendered from a template rather than generated.
- Stack. FastAPI backend on SQLite or Postgres, React front end, light and dark themes, a command palette and an audit log.
Results
- A new source goes from a name to a validated, scheduled sync without a hand-written connector.
- Failures arrive with a diagnosis and a one-click fix rather than a stack trace.
- Pipelines run on the customer’s own infrastructure into the warehouse they already use.
What I’d do differently
Keep the model on the parts that are genuinely unpredictable (reading API documentation, proposing mappings) and use deterministic templates everywhere else. CDC pipelines are rendered from a template with no model call, and they are the most predictable part of the system. I would have taken that approach for more of the pipeline types sooner.