PDF to Markdown Converter
Convert a PDF to clean Markdown with headings, lists, tables and links, ready to paste into ChatGPT or a docs site.
or drop it anywhere on this page, or paste
Opens in your browser. Nothing is uploaded.
PDF to Markdown Converter, free and private
Convert PDF to Markdown for free and get a clean .md file with real structure. Larger text becomes # headings, wrapped lines are joined back into paragraphs, bullets and numbers become Markdown lists, aligned columns become pipe tables, and clickable links in the PDF become Markdown links.
Use it to paste a report into ChatGPT, Claude or another LLM, or to move a manual into a docs site or an Obsidian vault. Everything runs in your browser, so the PDF is never uploaded. A preview of the Markdown appears as soon as you add the file, and you can copy it straight to the clipboard.
How to convert a PDF to Markdown
- 1
Add your PDF
Press Select a PDF file or drop the PDF onto the page. The tool reads one document at a time and starts reading its pages right away.
- 2
Check the Markdown preview
The preview shows the Markdown source of the first 30 pages. Switch to Rendered to see the headings, lists and tables laid out. Copy Markdown puts the whole document on the clipboard once every page is read.
- 3
Choose images, page breaks and pages
By default you get one .md file. Turn on Save images too to get a ZIP with an images folder, linked from the Markdown. Page markers put a hidden page number comment before every page and a rule between pages. Choose pages limits the output to a range like 2-5, 9.
- 4
Convert PDF to Markdown and download
Press Convert to Markdown, then Download MD or Download ZIP. The result lists how many headings, tables and words were found, and names any pages that had no text.
Why use PDFToolz for this
Headings from font sizes
The tool measures the body text size of the whole document. Short lines set at least 15% larger become headings, and each larger size gets a higher level, from # down to #####. A short line set fully in bold at body size becomes the next level down.
Paragraphs without broken lines
Lines that wrapped in the PDF are joined into one paragraph. A word split by a hyphen at the end of a line is joined again, and a paragraph cut by a page break is joined with the rest on the next page.
Lists, tables and links
Bullets such as • and ▪ or a leading hyphen, numbers like 1. or 2) and letters like a) become Markdown lists, nested by indent up to four levels. Text in aligned columns becomes a GitHub Flavored Markdown table, with number columns aligned right. Link annotations become [text](url) links.
Bold and italic kept
When the font name says bold or italic, the text is written as **bold** or *italic*. Characters such as * and _ that Markdown would read as formatting are escaped, so the text shows exactly as it did in the PDF.
Headers and footers removed
Page numbers and lines repeated at the top or bottom of at least half the pages, such as a running title or a confidentiality notice, are left out by default. They waste tokens in an LLM prompt and break up paragraphs. Pick Keep to leave every line in.
Free, private and without Acrobat
There is no account, watermark or page limit, and you do not need Adobe Acrobat. The converter runs in Chrome, Edge, Firefox and Safari on Windows, Mac, Linux, iPhone, iPad and Android. Your PDF stays on your device.
PDF to Markdown for ChatGPT and other LLMs
Language models read Markdown well. Headings show where a section starts, lists keep their items apart, and a pipe table keeps each value in its row and column. Pasting raw text copied from a PDF loses all of that and often brings page numbers, running titles and words split by hyphens with it.
Convert the PDF here, press Copy Markdown and paste the result into the chat. Turn on Page markers when you want the model to cite page numbers. Each page then starts with a comment such as <!-- Page 12 -->, which the model can see but a Markdown viewer hides. For long documents, choose only the pages you need so the prompt stays within the model's context window.
What the Markdown keeps and what it drops
The converter reads the text layer of the PDF with pdf.js and rebuilds the structure from the position, size and font of each piece of text. It does not copy the look of the page.
- Kept: headings, paragraphs, bulleted and numbered lists, tables, bold, italic, web and email links, and the reading order of two-column pages.
- Optional: images, saved as JPG for photos and PNG for graphics in an images folder. Images under 24 points wide or tall, such as icons and rules, are skipped.
- Not kept: fonts, colours, exact positions, page backgrounds, form field values, comments and links that jump to another page in the same PDF.
- Written as plain text: footnotes, captions, math and code, since a PDF does not mark them as such.
How tables are found
A PDF stores a table as separate pieces of text placed in a grid. The tool looks for runs of lines that split into two or more cells at wide gaps, finds the columns across those lines, and writes a pipe table with the first row as the header. A cell that wraps onto a second line is joined back into its row.
Two columns of running text, or a label beside a long sentence, are not treated as tables. Tables with merged cells, or rows whose cells are centred vertically across several lines, can come out with extra rows. Check those tables in the preview before you rely on them.
Scanned PDFs need OCR first
A scanned PDF holds pictures of pages, so there is no text to convert. If none of the chosen pages has text, the tool stops and points you to OCR PDF. If only some pages are scans, the Markdown is still made and the result lists the pages that had no text.
Run OCR PDF on the file first. It adds a hidden text layer in your browser, and the OCR result converts here like any other PDF. Headings and tables from OCR text are less reliable, because every line comes out at a similar size.
Limits to know about
Layouts with three or more text columns, such as newsletters and cheat sheets, are read across the page row by row. Two-column layouts read the left column first. A PDF that needs a password to open has to go through Unlock PDF first, and you need to know the password. PDFs with only an owner password, which blocks copying or printing, convert without one.
Very long PDFs work, but the preview stops at 30 pages and Copy Markdown waits until every page is read. The download always has every chosen page.
Frequently asked questions
Add the PDF, check the preview, choose whether to save images, and press Convert to Markdown. Then press Download MD, or use Copy Markdown to put the text on your clipboard.
It is free, with no sign-up, watermark or page limit. The PDF is read and the Markdown is built in your browser, so the file never leaves your device.
Convert it to Markdown first and paste the result. Headings, lists and tables survive as Markdown, and page numbers and running headers are removed, so the model gets the content without the clutter. Turn on Page markers if you want answers that cite pages.
Yes. Text set in aligned columns becomes a GitHub Flavored Markdown pipe table, with number columns aligned right. Complex tables with merged cells may need a quick fix by hand.
Not directly. A scan has no text layer. Run OCR PDF first to add one, then convert the OCR result here.
By default they are left out and you get one .md file. Pick Save images too to get a ZIP with the .md file and an images folder, where every image is linked from the Markdown at the spot it had on the page.
Yes. The tool runs in Safari, Chrome, Edge and Firefox on Mac, Windows, Linux, iPhone, iPad and Android, with nothing to install. Large PDFs convert faster on a computer than on a phone.
Levels come from font sizes. A PDF that uses the same size for two kinds of heading gives them one level, and a bold line at body size always becomes the lowest level. Change the number of # signs in the file if you need a different outline.