Last modified: Oct 06, 2026

Extract Tables from PDF with pdfplumber

PDF files often contain valuable tables. Extracting that data manually is slow and error-prone. Python's pdfplumber library makes this task easy and accurate.

This guide shows you how to extract tables from PDFs using pdfplumber. You will learn basic extraction, handling multiple tables, and exporting data to CSV.

What is pdfplumber?

pdfplumber is a Python library for reading PDF files. It can extract text, lines, rectangles, and tables. It is built on top of pdfminer.six.

The library is designed for machine-generated PDFs. It gives you detailed layout information. That makes table extraction more reliable than simple text parsing.

Why Extract Tables from PDFs?

Tables in PDFs hold structured data. That data is often needed in spreadsheets or databases. Manual copying is tedious.

Automated extraction saves time and reduces errors. You can process hundreds of PDFs in minutes. This is useful for invoices, financial reports, and research papers.

Prerequisites

You need Python 3.10 or higher. You also need the pdfplumber library installed. If you have not installed it yet, run this command:


pip install pdfplumber

After installation, you are ready to extract tables.

Basic Table Extraction with extract_table

The simplest way to get a table is the extract_table method. It works on a single page and returns the first table it finds.

First, open your PDF using the open method. Then access a page with the pages attribute. Finally, call extract_table.

Here is a complete example:


import pdfplumber

# Open the PDF file
with pdfplumber.open("data.pdf") as pdf:
    # Get the first page
    page = pdf.pages[0]
    # Extract the first table
    table = page.extract_table()
    # Print each row
    for row in table:
        print(row)

The output is a list of lists. Each inner list is a row. Here is an example output:


['Product', 'Price', 'Quantity']
['Apples', '$1.20', '10']
['Bananas', '$0.50', '20']
['Oranges', '$0.80', '15']

This method is quick. But it only returns one table per page. If a page has multiple tables, you need a different approach.

Extracting Multiple Tables with extract_tables

Use the extract_tables method to get all tables on a page. It returns a list of tables. Each table is a list of rows.

This is useful when a page contains several separate tables. The method automatically detects them.


import pdfplumber

with pdfplumber.open("report.pdf") as pdf:
    page = pdf.pages[0]
    # Get all tables on the page
    tables = page.extract_tables()
    # Loop through each table
    for i, table in enumerate(tables):
        print(f"Table {i+1}:")
        for row in table:
            print(row)
        print()

Suppose the page has two tables. The output might look like this:


Table 1:
['Name', 'Age']
['Alice', '30']
['Bob', '25']

Table 2:
['City', 'Population']
['New York', '8M']
['London', '9M']

Always check the number of tables found. Sometimes a single visual table is split into multiple tables by the detector.

Understanding Table Settings

pdfplumber uses table settings to decide what counts as a table. You can customize these settings for better results.

Common settings include vertical_strategy and horizontal_strategy. They control how lines are detected.

The default strategy is "lines". It looks for actual lines in the PDF. If your table has no lines, use "text" instead.

Here is how to pass table settings:


import pdfplumber

with pdfplumber.open("borderless.pdf") as pdf:
    page = pdf.pages[0]
    # Use text-based detection for tables without borders
    table = page.extract_table({
        "vertical_strategy": "text",
        "horizontal_strategy": "text"
    })
    for row in table:
        print(row)

This is very helpful for PDFs with invisible table borders. Experiment with different strategies to get the best result.

Handling Complex Tables

Real-world tables can be messy. They may have merged cells, multiple headers, or nested tables. pdfplumber handles many of these cases.

For merged cells, the extracted output may contain None values. You can clean these up after extraction.

For tables that span multiple pages, you need to extract from each page and combine the rows. Do not assume one table equals one page.

Here is an example of cleaning up None values:


import pdfplumber

with pdfplumber.open("complex.pdf") as pdf:
    page = pdf.pages[0]
    table = page.extract_table()
    cleaned = []
    for row in table:
        # Replace None with empty string
        cleaned_row = [cell if cell is not None else "" for cell in row]
        cleaned.append(cleaned_row)
    for row in cleaned:
        print(row)

This simple step makes the data easier to work with.

Exporting Tables to CSV

Once you have the table data, you often want to save it. The most common format is CSV. You can use Python's built-in csv module.

Here is how to write a table to a CSV file:


import pdfplumber
import csv

with pdfplumber.open("data.pdf") as pdf:
    page = pdf.pages[0]
    table = page.extract_table()
    # Open a CSV file for writing
    with open("output.csv", "w", newline="") as f:
        writer = csv.writer(f)
        writer.writerows(table)

This creates a file named output.csv with your table data. You can open it in Excel or Google Sheets.

If you prefer pandas, you can convert the table to a DataFrame and use to_csv. But that requires installing pandas separately.

Troubleshooting Common Issues

Sometimes extraction does not work as expected. Here are common problems and fixes.

Issue: Table is empty. The PDF may have no lines. Try setting both strategies to "text".

Issue: Extra columns. The detector may split one column into two. Adjust the vertical strategy or use explicit line positions.

Issue: Missing rows. Some rows may be skipped if they lack clear borders. You can use extract_text to manually parse those rows.

Issue: Scanned PDFs. pdfplumber cannot read scanned images. You need an OCR tool like Tesseract first.

Always test on a small sample before processing many files.

Conclusion

Extracting tables from PDFs with pdfplumber is straightforward. Start with extract_table for single tables. Use extract_tables for multiple tables. Customize table settings for complex layouts.

Remember to handle None values and export to CSV for further analysis. With a few lines of code, you can turn static PDF tables into usable data.

Practice on different PDFs to understand the settings. Soon you will extract tables quickly and accurately.