Don't have a PDF handy?

Everything runs in your browser. OCR (for scanned pages) is a best-effort reading of the page image, so proofread anything critical. Token counts are rough estimates.

Why use this PDF to Markdown converter

Correct reading order

Multi-column layouts, sidebars, and footnotes are put back into the order a person reads them, instead of columns interleaved into noise.

Real tables

Tables come through as real Markdown tables, so rows and columns keep their meaning. Ambiguous layouts fall back to clean text, never a broken grid.

Your files stay on your device

Everything runs locally in your browser, with no upload and no server in the path. Contracts, financials, and research never leave your machine.

How it works

How to convert a PDF to Markdown

Three steps, no account, and no file ever leaves your device.

  1. 1

    Drop in your PDF

    Select a PDF or drag it onto the converter. It loads straight into the browser tab, with no upload and no queue.

  2. 2

    Siteiz rebuilds the structure

    Reading order, headings, lists, and tables are reconstructed as Markdown. Scanned pages are detected and read with local OCR.

  3. 3

    Copy or download the .md

    Copy the Markdown to the clipboard or download the .md file. It is plain text, so it opens anywhere and drops into any AI tool.

The same converter handles PDF to MD, PDF to text, and scanned PDF to Markdown. Markdown is plain text, so the .md file opens in any editor and pastes straight into ChatGPT, Claude, or a RAG pipeline.

Pricing

Free for text PDFs. Pay only for OCR.

Text conversion is free forever. You only pay when a document is scanned and needs OCR. No account at any tier.

Standard
Free

For text-based PDFs.

  • Unlimited text-PDF conversions
  • 100% in-browser, nothing uploaded
  • Safe for confidential files
  • No account, no sign-up
Convert a PDF, free
24-hour passMost popular
99 kr/ 24 hours

For scanned documents and heavy reports.

  • Everything in Standard
  • OCR for scanned and image-only pages
  • Unlimited conversions for 24 hours
  • Tracked in your browser, no account
  • Nothing renews, no subscription
Business licence
20000 kr/ year

For organisations converting confidential document libraries.

  • Unlimited OCR and native processing
  • Commercial use across your organisation
  • Documents stay on your own machines
  • Shared team access link, no registration
  • Direct line to support
Request a corporate invoice

Prices in Swedish kronor (SEK): the 24-hour pass is 99 kr (about $9 USD), the business licence is 20000 kr per year (about $1,900 USD). VAT is handled at checkout. The pass is billed once through Stripe and stops after 24 hours; the business licence is billed annually by invoice.

Why it matters

Your PDF looks fine to you. AI sees something else.

Ordinary extractors read straight across a two-column page, so the columns interleave and the sentences stop making sense. The difference is not formatting, it is whether AI can understand the document at all.

A typical converter (columns interleaved)scrambled
The options for choosing a customised
interior design are 1. Brown/Light
practically limitless. To make your
choice a little easier, our 2. Brown/Dark
designers have put together a selection
of interiors with 3. Black/Dark
perfectly matched colours.
Siteiz (reading order fixed)clean
The options for choosing a customised
interior design are practically
limitless. To make your choice a little
easier, our designers have put together
a selection of interiors with perfectly
matched colours.

1. Brown/Light
2. Brown/Dark
3. Black/Dark

The output

Clean, compact, ready for your AI stack

Structured Markdown that stays small and drops straight into the tools you use.

Measured result

Original PDF14.3 MB
Siteiz Markdown848 KB
94.1%smaller, 232 pages in one pass

One real 232-page annual report, dense financial tables and scanned pages included. Measured, nothing modelled.

Works with your stack

The Markdown loads straight into the assistants and pipelines you already use.

Assistants

  • ChatGPT
  • Claude
  • Gemini
  • Microsoft Copilot

Pipelines & tools

  • RAG systems
  • Knowledge bases
  • LangChain
  • LlamaIndex
For developers: load it into LangChain or LlamaIndexCode
LangChain
from langchain_community.document_loaders import (
    UnstructuredMarkdownLoader,
)
from langchain_text_splitters import (
    MarkdownHeaderTextSplitter,
)

docs = UnstructuredMarkdownLoader("report.md").load()

# Headings give the splitter real boundaries
splitter = MarkdownHeaderTextSplitter(
    headers_to_split_on=[("#", "h1"), ("##", "h2")],
)
chunks = splitter.split_text(docs[0].page_content)
LlamaIndex
from llama_index.core import SimpleDirectoryReader
from llama_index.readers.file import MarkdownReader

reader = SimpleDirectoryReader(
    input_files=["report.md"],
    file_extractor={".md": MarkdownReader()},
)
documents = reader.load_data()
# Markdown structure carries into your nodes

FAQ

Questions, answered

What is PDF to Markdown?

PDF to Markdown converts a PDF into a clean .md file that keeps the document's headings, lists, and tables, so the content is easy to search, edit, and feed to AI systems. Siteiz does this entirely in your browser: it reconstructs multi-column reading order, keeps real tables, and reads scanned pages with OCR, so nothing is uploaded.

Is my PDF uploaded to a server?

No. Conversion runs entirely in your browser with pdf.js, and OCR runs locally as WebAssembly. The file never leaves your device and nothing is sent to Siteiz or any third party, so it is safe for confidential contracts, financial statements, and other regulated documents.

Does it handle scanned PDFs (OCR), and what does it cost?

Converting text-based PDFs is free and unlimited, with no account. Scanned or image-only pages are detected automatically and read with local OCR, which is PDF Pro: 99 kr for 24 hours of unlimited OCR on any file. It is a single charge that stops after 24 hours and never renews. Teams that need OCR continuously use the annual business licence.

Does it work with RAG, LangChain, and LlamaIndex?

Yes, that is what it is built for. Convert the PDF to Markdown here, then load the .md with your framework's Markdown loader instead of a PDF loader. Reconstructed reading order keeps two-column papers from interleaving mid-sentence, Markdown headings give splitters real chunk boundaries, and tables stay as tables so row and column relationships survive into the embedding.

Why Markdown instead of plain text for LLMs?

Markdown preserves the structure that plain text discards: headings, lists, and tables. That structure improves chunking and retrieval in RAG pipelines and gives an LLM cleaner boundaries to cite, while staying compact and token-efficient.

How accurate is OCR on scanned documents?

OCR is a best-effort reading of the page image, so proofread anything critical. Clean, high-resolution scans of printed text read well. Low-resolution faxes, handwriting, heavy background imagery, and unusual typefaces are harder, and pages that are mostly photographs can produce stray characters.

How do I convert a PDF to Markdown for free?

Open the converter on this page, drop in your PDF, and copy or download the .md file. Text-based PDFs are free and unlimited, with no account and no page cap. The conversion runs in your browser, so there is no upload step and no waiting in a queue.

What is the difference between PDF to MD and PDF to Markdown?

They are the same thing. Markdown files use the .md extension, so a PDF to MD converter and a PDF to Markdown converter describe one conversion: a PDF turned into a structured .md text file.

Can I use this as a PDF to text converter?

Yes. Markdown is plain text, so the .md output opens in any text editor and can be renamed .txt if you need it. The advantage over a raw text dump is that headings, lists, and tables survive instead of collapsing into one undifferentiated block.

Is there a page or file size limit?

Text-based PDFs have no page limit and no file cap beyond what your browser can hold in memory, since the work happens on your own machine. Very large scanned documents run slower because OCR reads each page image locally.

Which languages does the converter handle?

Text-based PDFs convert in any language the file already contains, including Swedish and other Latin-script languages, because the text is read directly from the document. OCR for scanned pages is tuned for English and Latin-script text, so results on other scripts vary.

Do you offer a business or team licence?

Yes. The business licence is 20000 kr per year, billed annually by invoice. It covers unlimited OCR and native processing for commercial use across your organisation and gives your team a shared access link with no user registration. Because conversion happens in the browser, your documents stay on your own machines. Email hello@siteiz.com for a corporate invoice.

Ready to convert?

Text-based PDFs are free and unlimited. Scanned documents need OCR, which is a 99 kr pass for 24 hours. No subscription, no account.

Convert a PDF for free

Guides

Get more out of the conversion

Also from Siteiz

Curious how AI reads your live pages?

Try our free AI audit to see what AI systems understand and cite on your website.

Run a free audit

Something converted badly?

PDFs are endlessly varied, and a layout that defeats the parser is a bug worth fixing. Send the file, or just describe it, and it becomes a test case.