PDF to Markdown
Drop in a PDF and get clean Markdown back — headings, lists, paragraphs, and links rebuilt from the page layout, ready for docs, notes, or an LLM prompt. Nothing is uploaded.
Drop a PDF here, or click to choose a file
Converted in your browser — the file is never uploaded
How this PDF to Markdown converter reads a page
A PDF doesn't contain paragraphs or headings — it contains positioned pieces of text: this string, in this font, at this size, at these coordinates. To convert PDF to Markdown, the tool reads those pieces with PDF.js (the same engine Firefox uses to display PDFs) and rebuilds the structure the way a person would when looking at the page:
- Lines and paragraphs come from baselines and spacing: text on the same baseline is a line, a bigger gap or an indented first line starts a new paragraph, and words hyphenated across a line break are joined back together.
- Headings are text set larger than the body size (the most common size in the document), ranked by size into #, ##, and ###. Short bold lines that stand on their own become headings too.
- Lists are recognized from bullet glyphs (•, ◦, ▪, Word's symbol bullets), numbering like 1. a) (iv), and — for browser-printed PDFs — bullets drawn as small shapes. Nesting follows the indent.
- Bold, italic, and code are read from font names ("Helvetica-Bold", "Times-Italic", Courier, Menlo, Consolas), so monospace blocks become fenced code and inline monospace becomes backticks.
- Links come from the PDF's link annotations and turn into [text](url).
- Columns and tables: two- and three-column pages are read column by column instead of straight across, and rows of aligned cells become Markdown tables.
Everything runs locally in your browser — there's no server step, so the PDF is never uploaded, and large documents are read page by page with a progress bar. Options apply instantly without re-reading the file.
What converts well — and what doesn't
PDF to MD conversion is only as good as the text inside the PDF. A quick guide to what to expect:
- Great: PDFs exported from Word, Google Docs, Pages, LaTeX, Markdown tools, and "Print to PDF" from a browser. Reports, papers, manuals, contracts, and resumes usually come out clean.
- Good, with some cleanup: multi-column academic papers, documents with footnotes (they land where they sit on the page), and tables with merged or multi-level header cells.
- Limited: math and equations (you get the symbols, not LaTeX), charts and diagrams (text labels only), and slide decks where text is scattered in boxes.
- Not possible without OCR: scanned documents and photos of pages. They contain images, not text — the converter detects them and says so instead of handing you an empty file.
PDF to Markdown for LLMs, ChatGPT, Claude, and RAG
The most common reason to convert PDF to Markdown today is to feed a document to a language model. Pasting a PDF's raw text loses its shape — headings run into paragraphs, tables collapse into a stream of numbers, and running headers repeat on every page. Markdown keeps that structure in plain text, which is exactly what models handle best, and it's cheaper in tokens than HTML.
- Chat: convert, copy, and paste into ChatGPT, Claude, or Gemini — or download the .md and attach it. The ~tokens figure is a rough size check (about four characters per token).
- RAG and embeddings: split the Markdown on headings to get chunks that follow the document's own sections. Choose "Page-number comments" to keep a
<!-- page 12 -->marker in each chunk so answers can cite the page. - Clean input: leave "Remove headers, footers & page numbers" on so repeated margin text doesn't pollute every chunk, and use the page range to skip covers, tables of contents, and appendices.
Other ways to convert PDF to Markdown (Python and CLI)
If you're converting thousands of files or building a pipeline, a library makes more sense than a web page. The honest landscape:
- pymupdf4llm — a Python package on top of PyMuPDF that outputs LLM-ready Markdown. Fast, handles headings and tables well; the usual first pick for PDF to Markdown in Python.
- MarkItDown— Microsoft's Python tool that converts PDF, Word, Excel, PowerPoint, and HTML to Markdown with one command. Its PDF path extracts text with little structure, so headings and tables are often flattened.
- Marker — deep-learning based PDF to Markdown with strong layout, table, and equation handling (equations become LaTeX). Heavier to install and best with a GPU.
- Docling— IBM's open-source document converter with layout and table-structure models, exporting Markdown or JSON; popular for RAG pipelines.
- Pandoc— the universal document converter, but it can't read PDF. Convert the PDF to Markdown first, then use Pandoc to go from Markdown to anything.
For a single document — a paper to summarize, a contract to paste into a chat, a manual to turn into docs — this page gets you there without installing anything, and your file stays on your machine.
Tips for a cleaner conversion
- If headings come out as plain paragraphs, the PDF probably uses the same font size for headings and body — they may still be caught as bold lines. Turn "Detect headings" off if a document ends up with too many.
- If a table comes out as separate lines, it likely has merged cells or ragged columns. Turn "Rebuild tables" off to get plain lines you can tidy by hand, or paste them into a spreadsheet.
- For a huge PDF, convert only the pages you need with the page range (e.g. 1-20, 45) — it's faster and keeps the output focused.
- A PDF that won't open or comes out garbled (odd symbols instead of letters) often has broken font encoding. Re-printing it to a new PDF from a viewer usually fixes it.
Frequently asked questions
How do I convert a PDF to Markdown?
Drop the PDF onto the box above (or click it to choose a file). The converter reads every page in your browser, rebuilds headings, paragraphs, lists, tables, code, and links from the layout, and shows the Markdown right away. Copy it, or download it as a .md file named after your PDF.
Is this PDF to Markdown converter free, and is my file uploaded?
It's free with no signup and no page limit, and nothing is uploaded. The PDF is opened with PDF.js inside your browser tab, so the file, its text, and any password you type never leave your device. You can even disconnect from the internet after the page loads and it keeps working.
Can it convert scanned PDFs?
Not directly. A scanned PDF is a picture of each page with no text inside, so there's nothing to extract — the converter detects this and tells you. Run OCR first (Adobe Acrobat's "Scan & OCR", Google Drive's "Open with Google Docs", or the free OCRmyPDF tool), then convert the OCR'd PDF here. PDFs that were scanned and already OCR'd work fine, because they carry an invisible text layer.
How do I convert PDF to Markdown for ChatGPT, Claude, or another LLM?
Convert the PDF here, then paste the Markdown into the chat or attach the .md file. Markdown keeps the document's structure — headings, lists, tables — in plain text, which models read more reliably than a raw PDF text dump, and the token estimate under the output tells you roughly how much of the context window it will use. Turn on page-number comments if you want the model to cite pages.
How do I convert PDF to Markdown in Python?
For scripts and pipelines, the common open-source options are pymupdf4llm (fast, layout-aware, built on PyMuPDF), Microsoft's MarkItDown (one converter for PDF, Word, Excel, and more), Marker (deep-learning based, strong on complex layouts and equations, benefits from a GPU), and IBM's Docling (layout and table models, exports Markdown or JSON). For a one-off document, this page is quicker: no install, no environment.
Can Pandoc convert PDF to Markdown?
No. Pandoc can write PDF (through LaTeX or another engine), but it can't read PDF as an input format. To go from PDF to Markdown you need a PDF text extractor — this page, or a library like pymupdf4llm, MarkItDown, Marker, or Docling — and then you can hand the Markdown to Pandoc to convert it anywhere else.
Does it keep tables, images, and math?
Simple tables — aligned columns of short cells, with or without borders — are rebuilt as GitHub-flavored Markdown tables, and cells that wrap onto several lines are merged. Merged cells and complex nested tables may come out as plain lines. Images aren't extracted (the output is text only), and math set as symbols comes through as its characters, not as LaTeX.
What does "Remove headers, footers & page numbers" do?
It finds lines that repeat in the top or bottom margin across pages — a running title, a chapter name, a confidentiality notice — plus page numbers that count up with the pages, and leaves them out so they don't interrupt your paragraphs. Paragraphs that continue across a page break are joined back together.
Can it open password-protected PDFs?
Yes, if you know the password. The converter asks for it on the page, uses it only to decrypt the file in your browser, and doesn't store it. PDFs that open without a password but restrict copying are converted normally.
Related free tools
See all free tools →Markdown to PDF
Turn Markdown into a clean, styled PDF in your browser
PDF to link
Upload a PDF, get a link that opens on any phone
Word to Markdown
Convert .docx Word documents to clean Markdown
HTML to Markdown
Turn HTML or pasted rich text into clean Markdown
Online Markdown editor
Write Markdown with live preview, autosave, and export
Built something? Put it online in seconds
host0 is the cloud for small software: bring any coding agent, build the tool only you need — like this one — and say "deploy to host0". Live at a shareable URL, no servers to run.