Data Ingestion / ETL

Libraries for data extraction, transformation, and loading pipelines across multiple sources and destinations.

dlt loads data from APIs, databases, and files into your warehouse. You pip install the Python ETL library into your own code, with no platform to run.

How to choose:

  • Data from APIs, databases, or files, loaded into a warehouse or DuckDB: dlt
  • pandas DataFrames in and out of S3, Athena, Redshift, and other AWS services: awswrangler
  • Stock prices, financials, and options from Yahoo Finance, for personal research: yfinance
  • Chinese stocks, futures, and funds, for academic research: AKShare
  • US company filings and financial statements from SEC EDGAR: EdgarTools
  • Many data providers behind one Python API, also served over REST and MCP: OpenBB

Listed in editorial order, grouped by use case. Click a column to re-sort the whole list.

Press / to search. Tap a tag to filter. Click any row for details.

Search and filter

Results

Row number Tags

General 2 projects

Pandas integration with AWS services like Athena, Glue, Redshift, S3, and DynamoDB.
aws/github.com/aws/aws-sdk-pandas / /81,913,117 downloads/month
A Python library for building data pipelines with automatic schema inference, incremental loading, and support for multiple sources and destinations.
dlt-hub/github.com/dlt-hub/dlt / /4,829,719 downloads/month

Financial Data 4 projects

Easy Pythonic way to download market and financial data from Yahoo Finance.
ranaroussi/github.com/ranaroussi/yfinance / /18,321,524 downloads/month
A financial data interface library, with data provided for academic research only.
akfamily/github.com/akfamily/akshare / /1,060,057 downloads/month
Library for downloading structured data from SEC EDGAR filings and XBRL financial statements.
dgunning/github.com/dgunning/edgartools / /835,246 downloads/month
A financial data platform for analysts, quants and AI agents.
openbq-org/github.com/openbq-org/OpenBB / /59,206 downloads/month

Data Ingestion / ETL guide

dlt loads data from messy sources into well-structured datasets: it infers the schema and data types, normalizes the data, and handles nested structures. Start a project with dlt init rest_api duckdb, which sets up a pipeline script with a REST API source and a DuckDB destination. Build and test on DuckDB, then switch out the destination when you deploy. Beyond the built-in sources, any Python iterable can feed it, since a resource is just a generator. dlt runs anywhere Python runs, so schedule it with the orchestrator you already have.

AWS SDK for pandas extends pandas to AWS: import it as awswrangler, and it connects DataFrames to S3, Athena, Glue, Redshift, and DynamoDB. Its quick start stores a DataFrame on S3 as a Parquet dataset with wr.s3.to_parquet(), passing a database and table name. Then wr.athena.read_sql_query() queries that table with SQL.

yfinance offers a Pythonic way to fetch financial and market data from Yahoo Finance. Its quick start uses Ticker for one symbol's prices, financial statements, and options, and download for prices across several symbols at once.

AKShare aims to simplify fetching financial data, with one function call per dataset, like ak.stock_zh_a_hist() for the daily history of a Chinese stock. Its README credits mostly Chinese exchanges and finance sites as its sources. Its docs are in Chinese, and each data interface comes with an example you can copy and paste.

EdgarTools turns any SEC filing into a typed Python object, so a 10-K's revenue is one line instead of an afternoon of XBRL parsing. EDGAR requires an email with every request, so set your identity with set_identity() before anything else. Everything starts with a Company or a Filing. Company("AAPL").get_financials().income_statement() returns a standardized income statement, and .obj() turns a filing into an object built for its form, with its data as pandas DataFrames.

OpenBB is a "connect once, consume everywhere" layer over financial data sources. Each provider package connects one source and adds its own namespace to the obb client. The same commands run from Python, as a REST API, and as MCP tools. A call returns an OBBject, and to_dataframe() turns it into a pandas DataFrame.

Before you build on financial data, check whose terms you're under. yfinance isn't affiliated with Yahoo and is meant for research and education, and Yahoo's API is for personal use only. AKShare's data is for academic research only. OpenBB hosts no data itself: each provider sets its own coverage and terms of use.

Know a project that belongs here?

Tell us what it does and why it stands out.