| 1 |
charset-normalizer
Encoding and Unicode
|
1,513,041,225 |
792 |
|
Encoding and Unicode
Text Processing
Text & Documents
|
→ |
|
Universal character encoding detector, the default of the requests ecosystem.
|
|
|
|
| 2 |
pygments
Parser
|
1,107,477,999 |
2,205 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
A generic syntax highlighter.
|
|
|
|
| 3 |
pyyaml
Data Formats
|
1,067,064,724 |
2,941 |
|
Data Formats
File Format Processing
Text & Documents
|
→ |
|
YAML implementations for Python.
|
|
|
|
| 4 |
markupsafe
Text & Documents
|
748,176,132 |
697 |
|
HTML Manipulation
Text & Documents
|
→ |
|
Implements a XML/HTML/XHTML Markup safe string for Python.
|
|
|
|
| 5 |
markdown-it-py
Markdown
|
537,590,272 |
1,356 |
|
Markdown
File Format Processing
Text & Documents
|
→ |
|
Markdown parser with 100% CommonMark support, extensions, and syntax plugins.
|
|
|
|
| 6 |
pyparsing
Parser
|
376,975,748 |
2,487 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
A general purpose framework for generating parsers.
|
|
|
|
| 7 |
beautifulsoup4
Text & Documents
|
368,681,476 |
External |
— |
HTML Manipulation
Text & Documents
|
→ |
|
Providing Pythonic idioms for iterating, searching, and modifying HTML or XML.
|
|
|
|
| 8 |
lxml
Text & Documents
|
349,832,871 |
3,054 |
|
HTML Manipulation
Text & Documents
|
→ |
|
A very fast, easy-to-use and versatile library for handling HTML and XML.
|
|
|
|
| 9 |
watchfiles
Text & Documents
|
343,160,095 |
2,532 |
|
File Manipulation
Text & Documents
|
→ |
|
Simple, modern and fast file watching and code reload in python.
|
|
|
|
| 10 |
openpyxl
Excel
|
312,072,914 |
External |
— |
Excel
File Format Processing
Text & Documents
|
→ |
|
A library for reading and writing Excel 2010 xlsx/xlsm/xltx/xltm files.
|
|
|
|
| 11 |
chardet
Encoding and Unicode
|
200,686,388 |
2,666 |
|
Encoding and Unicode
Text Processing
Text & Documents
|
→ |
|
Python character encoding detector.
|
|
|
|
| 12 |
rapidfuzz
Fuzzy Matching
|
148,268,106 |
4,110 |
|
Fuzzy Matching
Text Processing
Text & Documents
|
→ |
|
Rapid fuzzy string matching using various string metrics, with a C++ core.
|
|
|
|
| 13 |
pypdf
PDF
|
140,243,072 |
10,191 |
|
PDF
File Format Processing
Text & Documents
|
→ |
|
A library capable of splitting, merging, cropping, and transforming PDF pages.
|
|
|
|
| 14 |
sqlparse
Parser
|
133,508,739 |
4,016 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
A non-validating SQL parser.
|
|
|
|
| 15 |
babel
Internationalization
|
119,301,600 |
1,464 |
|
Internationalization
Text Processing
Text & Documents
|
→ |
|
An internationalization library for Python.
|
|
|
|
| 16 |
markdown
Markdown
|
111,206,346 |
4,245 |
|
Markdown
File Format Processing
Text & Documents
|
→ |
|
A Python implementation of John Gruber’s Markdown.
|
|
|
|
| 17 |
xmltodict
Text & Documents
|
110,005,094 |
5,754 |
|
HTML Manipulation
Text & Documents
|
→ |
|
Working with XML feel like you are working with JSON.
|
|
|
|
| 18 |
python-docx
Word
|
101,997,302 |
5,708 |
|
Word
File Format Processing
Text & Documents
|
→ |
|
Reads, queries and modifies Microsoft Word 2007/2008 docx files.
|
|
|
|
| 19 |
watchdog
Text & Documents
|
101,605,344 |
7,405 |
|
File Manipulation
Text & Documents
|
→ |
|
API and shell utilities to monitor file system events.
|
|
|
|
| 20 |
xlsxwriter
Excel
|
93,051,501 |
3,971 |
|
Excel
File Format Processing
Text & Documents
|
→ |
|
A Python module for creating Excel .xlsx files.
|
|
|
|
| 21 |
reportlab
PDF
|
82,807,962 |
External |
— |
PDF
File Format Processing
Text & Documents
|
→ |
|
Allowing Rapid creation of rich PDF documents.
|
|
|
|
| 22 |
python-slugify
Transliteration and Slugs
|
74,512,847 |
1,624 |
|
Transliteration and Slugs
Text Processing
Text & Documents
|
→ |
|
A Python slugify library that translates unicode to ASCII.
|
|
|
|
| 23 |
mistune
Markdown
|
66,884,727 |
3,070 |
|
Markdown
File Format Processing
Text & Documents
|
→ |
|
Fastest and full featured pure Python parsers of Markdown.
|
|
|
|
| 24 |
pdfminer.six
PDF
|
59,016,073 |
7,020 |
|
PDF
File Format Processing
Text & Documents
|
→ |
|
Pdfminer.six is a community maintained fork of the original PDFMiner.
|
|
|
|
| 25 |
python-pptx
PowerPoint
|
55,788,617 |
3,510 |
|
PowerPoint
File Format Processing
Text & Documents
|
→ |
|
Python library for creating and updating PowerPoint (.pptx) files.
|
|
|
|
| 26 |
phonenumbers
Parser
|
36,551,629 |
3,770 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
Parsing, formatting, storing and validating international phone numbers.
|
|
|
|
| 27 |
weasyprint
HTML-to-PDF
|
35,568,809 |
9,555 |
|
HTML-to-PDF
File Format Processing
Text & Documents
|
→ |
|
A visual rendering engine for HTML and CSS that can export to PDF.
|
|
|
|
| 28 |
python-magic
Text & Documents
|
33,159,776 |
2,917 |
|
File Manipulation
Text & Documents
|
→ |
|
A Python interface to the libmagic file type identification library.
|
|
|
|
| 29 |
unidecode
Transliteration and Slugs
|
28,939,630 |
611 |
|
Transliteration and Slugs
Text Processing
Text & Documents
|
→ |
|
ASCII transliterations of Unicode text.
|
|
|
|
| 30 |
shortuuid
Unique identifiers
|
19,848,511 |
2,197 |
|
Unique identifiers
Text Processing
Text & Documents
|
→ |
|
A generator library for concise, unambiguous and URL-safe UUIDs.
|
|
|
|
| 31 |
ftfy
Encoding and Unicode
|
15,807,077 |
4,062 |
|
Encoding and Unicode
Text Processing
Text & Documents
|
→ |
|
Makes Unicode text less broken and more consistent automagically.
|
|
|
|
| 32 |
pyelftools
General
|
13,927,229 |
2,277 |
|
General
File Format Processing
Text & Documents
|
→ |
|
Parsing and analyzing ELF files and DWARF debugging information.
|
|
|
|
| 33 |
markitdown
File Conversion
|
13,507,082 |
178,168 |
|
File Conversion
File Format Processing
Text & Documents
|
→ |
|
Python tool for converting files and office documents to Markdown.
|
|
|
|
| 34 |
docling
File Conversion
|
7,142,093 |
66,006 |
|
File Conversion
File Format Processing
Text & Documents
|
→ |
|
Library for converting documents into structured data.
|
|
|
|
| 35 |
pyfiglet
General
|
6,042,042 |
1,582 |
|
General
Text Processing
Text & Documents
|
→ |
|
An implementation of figlet written in Python.
|
|
|
|
| 36 |
tablib
General
|
5,110,967 |
4,755 |
|
General
File Format Processing
Text & Documents
|
→ |
|
A module for Tabular Datasets in XLS, CSV, JSON, YAML.
|
|
|
|
| 37 |
parsy
Parser
|
4,366,052 |
451 |
|
Parser
Text Processing
Text & Documents
|
→ |
|
Easy, generic parser combinator library for creating parsers.
|
|
|
|
| 38 |
sqids
Unique identifiers
|
698,987 |
521 |
|
Unique identifiers
Text Processing
Text & Documents
|
→ |
|
A library for generating short unique IDs from numbers.
|
|
|
|
| 39 |
justhtml
Text & Documents
|
77,154 |
1,154 |
|
HTML Manipulation
Text & Documents
|
→ |
|
A pure Python HTML5 parser that just works.
|
|
|
|
| 40 |
difflib
General
|
Stdlib |
75,981 |
|
General
Text Processing
Text & Documents
Stdlib
|
→ |
|
(Python standard library) Helpers for computing deltas.
|
|
|
|
| 41 |
mimetypes
Text & Documents
|
Stdlib |
75,981 |
|
File Manipulation
Text & Documents
Stdlib
|
→ |
|
(Python standard library) Map filenames to MIME types.
|
|
|
|
| 42 |
pathlib
Text & Documents
|
Stdlib |
75,981 |
|
File Manipulation
Text & Documents
Stdlib
|
→ |
|
(Python standard library) A cross-platform, object-oriented path library.
|
|
|
|
| 43 |
tomllib
Data Formats
|
Stdlib |
75,981 |
|
Data Formats
File Format Processing
Text & Documents
Stdlib
|
→ |
|
(Python standard library) Parse TOML files.
|
|
|
|