| 1 |
charset-normalizer
Encoding and Unicode
|
1,274,812,966 |
795 |
|
Encoding and Unicode
Text Processing
Text & Documents
|
→ |
|
Universal character encoding detector, the default of the requests ecosystem.
|
|
|
|
| 2 |
pygments
Parser
|
976,544,263 |
2,206 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
A generic syntax highlighter.
|
|
|
|
| 3 |
pyyaml
Data Formats
|
924,420,846 |
2,942 |
|
Data Formats
File Format Processing
Text & Documents
|
→ |
|
YAML implementations for Python.
|
|
|
|
| 4 |
markupsafe
Text & Documents
|
644,857,930 |
697 |
|
HTML Manipulation
Text & Documents
|
→ |
|
Implements a XML/HTML/XHTML Markup safe string for Python.
|
|
|
|
| 5 |
markdown-it-py
Markdown
|
456,891,455 |
1,361 |
|
Markdown
File Format Processing
Text & Documents
|
→ |
|
Markdown parser with 100% CommonMark support, extensions, and syntax plugins.
|
|
|
|
| 6 |
pyparsing
Parser
|
322,999,823 |
2,489 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
A general purpose framework for generating parsers.
|
|
|
|
| 7 |
beautifulsoup4
Text & Documents
|
319,896,891 |
External |
— |
HTML Manipulation
Text & Documents
|
→ |
|
Providing Pythonic idioms for iterating, searching, and modifying HTML or XML.
|
|
|
|
| 8 |
lxml
Text & Documents
|
315,155,782 |
3,055 |
|
HTML Manipulation
Text & Documents
|
→ |
|
A very fast, easy-to-use and versatile library for handling HTML and XML.
|
|
|
|
| 9 |
watchfiles
Text & Documents
|
300,696,963 |
2,534 |
|
File Manipulation
Text & Documents
|
→ |
|
Simple, modern and fast file watching and code reload in python.
|
|
|
|
| 10 |
openpyxl
Excel
|
278,934,193 |
External |
— |
Excel
File Format Processing
Text & Documents
|
→ |
|
A library for reading and writing Excel 2010 xlsx/xlsm/xltx/xltm files.
|
|
|
|
| 11 |
chardet
Encoding and Unicode
|
178,822,028 |
2,667 |
|
Encoding and Unicode
Text Processing
Text & Documents
|
→ |
|
Python character encoding detector.
|
|
|
|
| 12 |
rapidfuzz
Fuzzy Matching
|
126,381,069 |
4,125 |
|
Fuzzy Matching
Text Processing
Text & Documents
|
→ |
|
Rapid fuzzy string matching using various string metrics, with a C++ core.
|
|
|
|
| 13 |
pypdf
PDF
|
122,097,092 |
10,202 |
|
PDF
File Format Processing
Text & Documents
|
→ |
|
A library capable of splitting, merging, cropping, and transforming PDF pages.
|
|
|
|
| 14 |
sqlparse
Parser
|
116,208,522 |
4,018 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
A non-validating SQL parser.
|
|
|
|
| 15 |
babel
Internationalization
|
106,700,435 |
1,465 |
|
Internationalization
Text Processing
Text & Documents
|
→ |
|
An internationalization library for Python.
|
|
|
|
| 16 |
markdown
Markdown
|
96,967,670 |
4,246 |
|
Markdown
File Format Processing
Text & Documents
|
→ |
|
A Python implementation of John Gruber’s Markdown.
|
|
|
|
| 17 |
xmltodict
Text & Documents
|
96,410,117 |
5,756 |
|
HTML Manipulation
Text & Documents
|
→ |
|
Working with XML feel like you are working with JSON.
|
|
|
|
| 18 |
python-docx
Word
|
90,235,183 |
5,718 |
|
Word
File Format Processing
Text & Documents
|
→ |
|
Reads, queries and modifies Microsoft Word 2007/2008 docx files.
|
|
|
|
| 19 |
watchdog
Text & Documents
|
87,674,725 |
7,413 |
|
File Manipulation
Text & Documents
|
→ |
|
API and shell utilities to monitor file system events.
|
|
|
|
| 20 |
xlsxwriter
Excel
|
83,478,784 |
3,972 |
|
Excel
File Format Processing
Text & Documents
|
→ |
|
A Python module for creating Excel .xlsx files.
|
|
|
|
| 21 |
reportlab
PDF
|
77,178,030 |
External |
— |
PDF
File Format Processing
Text & Documents
|
→ |
|
Allowing Rapid creation of rich PDF documents.
|
|
|
|
| 22 |
python-slugify
Transliteration and Slugs
|
64,475,118 |
1,624 |
|
Transliteration and Slugs
Text Processing
Text & Documents
|
→ |
|
A Python slugify library that translates unicode to ASCII.
|
|
|
|
| 23 |
mistune
Markdown
|
59,752,370 |
3,075 |
|
Markdown
File Format Processing
Text & Documents
|
→ |
|
Fastest and full featured pure Python parsers of Markdown.
|
|
|
|
| 24 |
pdfminer.six
PDF
|
52,683,505 |
7,022 |
|
PDF
File Format Processing
Text & Documents
|
→ |
|
Pdfminer.six is a community maintained fork of the original PDFMiner.
|
|
|
|
| 25 |
python-pptx
PowerPoint
|
49,930,083 |
3,528 |
|
PowerPoint
File Format Processing
Text & Documents
|
→ |
|
Python library for creating and updating PowerPoint (.pptx) files.
|
|
|
|
| 26 |
phonenumbers
Parser
|
33,480,795 |
3,772 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
Parsing, formatting, storing and validating international phone numbers.
|
|
|
|
| 27 |
weasyprint
HTML-to-PDF
|
31,320,049 |
9,593 |
|
HTML-to-PDF
File Format Processing
Text & Documents
|
→ |
|
A visual rendering engine for HTML and CSS that can export to PDF.
|
|
|
|
| 28 |
python-magic
Text & Documents
|
29,873,915 |
2,918 |
|
File Manipulation
Text & Documents
|
→ |
|
A Python interface to the libmagic file type identification library.
|
|
|
|
| 29 |
unidecode
Transliteration and Slugs
|
26,080,572 |
610 |
|
Transliteration and Slugs
Text Processing
Text & Documents
|
→ |
|
ASCII transliterations of Unicode text.
|
|
|
|
| 30 |
shortuuid
Unique identifiers
|
18,227,781 |
2,199 |
|
Unique identifiers
Text Processing
Text & Documents
|
→ |
|
A generator library for concise, unambiguous and URL-safe UUIDs.
|
|
|
|
| 31 |
markitdown
File Conversion
|
15,170,649 |
183,994 |
|
File Conversion
File Format Processing
Text & Documents
|
→ |
|
Python tool for converting files and office documents to Markdown.
|
|
|
|
| 32 |
ftfy
Encoding and Unicode
|
14,934,602 |
4,063 |
|
Encoding and Unicode
Text Processing
Text & Documents
|
→ |
|
Makes Unicode text less broken and more consistent automagically.
|
|
|
|
| 33 |
pyelftools
General
|
11,987,940 |
2,279 |
|
General
File Format Processing
Text & Documents
|
→ |
|
Parsing and analyzing ELF files and DWARF debugging information.
|
|
|
|
| 34 |
pyfiglet
General
|
5,653,898 |
1,582 |
|
General
Text Processing
Text & Documents
|
→ |
|
An implementation of figlet written in Python.
|
|
|
|
| 35 |
docling
File Conversion
|
5,128,858 |
66,413 |
|
File Conversion
File Format Processing
Text & Documents
|
→ |
|
Library for converting documents into structured data.
|
|
|
|
| 36 |
tablib
General
|
4,615,146 |
4,757 |
|
General
File Format Processing
Text & Documents
|
→ |
|
A module for Tabular Datasets in XLS, CSV, JSON, YAML.
|
|
|
|
| 37 |
parsy
Parser
|
4,168,210 |
452 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
Easy, generic parser combinator library for creating parsers.
|
|
|
|
| 38 |
sqids
Unique identifiers
|
542,924 |
522 |
|
Unique identifiers
Text Processing
Text & Documents
|
→ |
|
A library for generating short unique IDs from numbers.
|
|
|
|
| 39 |
justhtml
Text & Documents
|
73,537 |
1,156 |
|
HTML Manipulation
Text & Documents
|
→ |
|
A pure Python HTML5 parser that just works.
|
|
|
|
| 40 |
difflib
General
|
Stdlib |
77,167 |
|
General
Text Processing
Text & Documents
Stdlib
|
→ |
|
(Python standard library) Helpers for computing deltas.
|
|
|
|
| 41 |
mimetypes
Text & Documents
|
Stdlib |
77,167 |
|
File Manipulation
Text & Documents
Stdlib
|
→ |
|
(Python standard library) Map filenames to MIME types.
|
|
|
|
| 42 |
pathlib
Text & Documents
|
Stdlib |
77,167 |
|
File Manipulation
Text & Documents
Stdlib
|
→ |
|
(Python standard library) A cross-platform, object-oriented path library.
|
|
|
|
| 43 |
tomllib
Data Formats
|
Stdlib |
77,167 |
|
Data Formats
File Format Processing
Text & Documents
Stdlib
|
→ |
|
(Python standard library) Parse TOML files.
|
|
|
|