Skip to content
PDFToolz

PDF to Markdown Converter

Convert a PDF to clean Markdown with headings, lists, tables and links, ready to paste into ChatGPT or a docs site.

Select a PDF file

or drop it anywhere on this page, or paste

Opens in your browser. Nothing is uploaded.

PDF to Markdown Converter, free and private

Convert PDF to Markdown for free and get a clean .md file with real structure. Larger text becomes # headings, wrapped lines are joined back into paragraphs, bullets and numbers become Markdown lists, aligned columns become pipe tables, and clickable links in the PDF become Markdown links.

Use it to paste a report into ChatGPT, Claude or another LLM, or to move a manual into a docs site or an Obsidian vault. Everything runs in your browser, so the PDF is never uploaded. A preview of the Markdown appears as soon as you add the file, and you can copy it straight to the clipboard.

How to convert a PDF to Markdown

  1. 1

    Add your PDF

    Press Select a PDF file or drop the PDF onto the page. The tool reads one document at a time and starts reading its pages right away.

  2. 2

    Check the Markdown preview

    The preview shows the Markdown source of the first 30 pages. Switch to Rendered to see the headings, lists and tables laid out. Copy Markdown puts the whole document on the clipboard once every page is read.

  3. 3

    Choose images, page breaks and pages

    By default you get one .md file. Turn on Save images too to get a ZIP with an images folder, linked from the Markdown. Page markers put a hidden page number comment before every page and a rule between pages. Choose pages limits the output to a range like 2-5, 9.

  4. 4

    Convert PDF to Markdown and download

    Press Convert to Markdown, then Download MD or Download ZIP. The result lists how many headings, tables and words were found, and names any pages that had no text.

Why use PDFToolz for this

Headings from font sizes

The tool measures the body text size of the whole document. Short lines set at least 15% larger become headings, and each larger size gets a higher level, from # down to #####. A short line set fully in bold at body size becomes the next level down.

Paragraphs without broken lines

Lines that wrapped in the PDF are joined into one paragraph. A word split by a hyphen at the end of a line is joined again, and a paragraph cut by a page break is joined with the rest on the next page.

Lists, tables and links

Bullets such as • and ▪ or a leading hyphen, numbers like 1. or 2) and letters like a) become Markdown lists, nested by indent up to four levels. Text in aligned columns becomes a GitHub Flavored Markdown table, with number columns aligned right. Link annotations become [text](url) links.

Bold and italic kept

When the font name says bold or italic, the text is written as **bold** or *italic*. Characters such as * and _ that Markdown would read as formatting are escaped, so the text shows exactly as it did in the PDF.

Headers and footers removed

Page numbers and lines repeated at the top or bottom of at least half the pages, such as a running title or a confidentiality notice, are left out by default. They waste tokens in an LLM prompt and break up paragraphs. Pick Keep to leave every line in.

Free, private and without Acrobat

There is no account, watermark or page limit, and you do not need Adobe Acrobat. The converter runs in Chrome, Edge, Firefox and Safari on Windows, Mac, Linux, iPhone, iPad and Android. Your PDF stays on your device.

PDF to Markdown for ChatGPT and other LLMs

Language models read Markdown well. Headings show where a section starts, lists keep their items apart, and a pipe table keeps each value in its row and column. Pasting raw text copied from a PDF loses all of that and often brings page numbers, running titles and words split by hyphens with it.

Convert the PDF here, press Copy Markdown and paste the result into the chat. Turn on Page markers when you want the model to cite page numbers. Each page then starts with a comment such as <!-- Page 12 -->, which the model can see but a Markdown viewer hides. For long documents, choose only the pages you need so the prompt stays within the model's context window.

What the Markdown keeps and what it drops

The converter reads the text layer of the PDF with pdf.js and rebuilds the structure from the position, size and font of each piece of text. It does not copy the look of the page.

  • Kept: headings, paragraphs, bulleted and numbered lists, tables, bold, italic, web and email links, and the reading order of two-column pages.
  • Optional: images, saved as JPG for photos and PNG for graphics in an images folder. Images under 24 points wide or tall, such as icons and rules, are skipped.
  • Not kept: fonts, colours, exact positions, page backgrounds, form field values, comments and links that jump to another page in the same PDF.
  • Written as plain text: footnotes, captions, math and code, since a PDF does not mark them as such.

How tables are found

A PDF stores a table as separate pieces of text placed in a grid. The tool looks for runs of lines that split into two or more cells at wide gaps, finds the columns across those lines, and writes a pipe table with the first row as the header. A cell that wraps onto a second line is joined back into its row.

Two columns of running text, or a label beside a long sentence, are not treated as tables. Tables with merged cells, or rows whose cells are centred vertically across several lines, can come out with extra rows. Check those tables in the preview before you rely on them.

Scanned PDFs need OCR first

A scanned PDF holds pictures of pages, so there is no text to convert. If none of the chosen pages has text, the tool stops and points you to OCR PDF. If only some pages are scans, the Markdown is still made and the result lists the pages that had no text.

Run OCR PDF on the file first. It adds a hidden text layer in your browser, and the OCR result converts here like any other PDF. Headings and tables from OCR text are less reliable, because every line comes out at a similar size.

Limits to know about

Layouts with three or more text columns, such as newsletters and cheat sheets, are read across the page row by row. Two-column layouts read the left column first. A PDF that needs a password to open has to go through Unlock PDF first, and you need to know the password. PDFs with only an owner password, which blocks copying or printing, convert without one.

Very long PDFs work, but the preview stops at 30 pages and Copy Markdown waits until every page is read. The download always has every chosen page.

Frequently asked questions

Related tools and guides