Private by design
PDF to Markdown Converter
Make your PDF content readable by AI.
Siteiz turns reports, white papers, and case studies into a clean .md file that keeps headings, tables, and reading order, so AI systems can search, quote, and reuse the text. Scanned PDFs are read with local OCR. Nothing is uploaded.
Private by design · No signup · No upload
Don't have a PDF handy?
Everything runs in your browser. OCR (for scanned pages) is a best-effort reading of the page image, so proofread anything critical. Token counts are rough estimates.
Why use this PDF to Markdown converter
Correct reading order
Multi-column layouts, sidebars, and footnotes are put back into the order a person reads them, instead of columns interleaved into noise.
Real tables
Tables come through as real Markdown tables, so rows and columns keep their meaning. Ambiguous layouts fall back to clean text, never a broken grid.
Your files stay on your device
Everything runs locally in your browser, with no upload and no server in the path. Contracts, financials, and research never leave your machine.
How it works
How to convert a PDF to Markdown
Three steps, no account, and no file ever leaves your device.
- 1
Drop in your PDF
Select a PDF or drag it onto the converter. It loads straight into the browser tab, with no upload and no queue.
- 2
Siteiz rebuilds the structure
Reading order, headings, lists, and tables are reconstructed as Markdown. Scanned pages are detected and read with local OCR.
- 3
Copy or download the .md
Copy the Markdown to the clipboard or download the .md file. It is plain text, so it opens anywhere and drops into any AI tool.
The same converter handles PDF to MD, PDF to text, and scanned PDF to Markdown. Markdown is plain text, so the .md file opens in any editor and pastes straight into ChatGPT, Claude, or a RAG pipeline.
Pricing
Free for text PDFs. Pay only for OCR.
Text conversion is free forever. You only pay when a document is scanned and needs OCR. No account at any tier.
For text-based PDFs.
- Unlimited text-PDF conversions
- 100% in-browser, nothing uploaded
- Safe for confidential files
- No account, no sign-up
For scanned documents and heavy reports.
- Everything in Standard
- OCR for scanned and image-only pages
- Unlimited conversions for 24 hours
- Tracked in your browser, no account
- Nothing renews, no subscription
For organisations converting confidential document libraries.
- Unlimited OCR and native processing
- Commercial use across your organisation
- Documents stay on your own machines
- Shared team access link, no registration
- Direct line to support
Prices in Swedish kronor (SEK): the 24-hour pass is 99 kr (about $9 USD), the business licence is 20000 kr per year (about $1,900 USD). VAT is handled at checkout. The pass is billed once through Stripe and stops after 24 hours; the business licence is billed annually by invoice.
Why it matters
Your PDF looks fine to you. AI sees something else.
Ordinary extractors read straight across a two-column page, so the columns interleave and the sentences stop making sense. The difference is not formatting, it is whether AI can understand the document at all.
The options for choosing a customised interior design are 1. Brown/Light practically limitless. To make your choice a little easier, our 2. Brown/Dark designers have put together a selection of interiors with 3. Black/Dark perfectly matched colours.
The options for choosing a customised interior design are practically limitless. To make your choice a little easier, our designers have put together a selection of interiors with perfectly matched colours. 1. Brown/Light 2. Brown/Dark 3. Black/Dark
The output
Clean, compact, ready for your AI stack
Structured Markdown that stays small and drops straight into the tools you use.
Measured result
One real 232-page annual report, dense financial tables and scanned pages included. Measured, nothing modelled.
Works with your stack
The Markdown loads straight into the assistants and pipelines you already use.
Assistants
- ChatGPT
- Claude
- Gemini
- Microsoft Copilot
Pipelines & tools
- RAG systems
- Knowledge bases
- LangChain
- LlamaIndex
For developers: load it into LangChain or LlamaIndexCode
from langchain_community.document_loaders import (
UnstructuredMarkdownLoader,
)
from langchain_text_splitters import (
MarkdownHeaderTextSplitter,
)
docs = UnstructuredMarkdownLoader("report.md").load()
# Headings give the splitter real boundaries
splitter = MarkdownHeaderTextSplitter(
headers_to_split_on=[("#", "h1"), ("##", "h2")],
)
chunks = splitter.split_text(docs[0].page_content)from llama_index.core import SimpleDirectoryReader
from llama_index.readers.file import MarkdownReader
reader = SimpleDirectoryReader(
input_files=["report.md"],
file_extractor={".md": MarkdownReader()},
)
documents = reader.load_data()
# Markdown structure carries into your nodesFAQ
Questions, answered
What is PDF to Markdown?
PDF to Markdown converts a PDF into a clean .md file that keeps the document's headings, lists, and tables, so the content is easy to search, edit, and feed to AI systems. Siteiz does this entirely in your browser: it reconstructs multi-column reading order, keeps real tables, and reads scanned pages with OCR, so nothing is uploaded.
Is my PDF uploaded to a server?
No. Conversion runs entirely in your browser with pdf.js, and OCR runs locally as WebAssembly. The file never leaves your device and nothing is sent to Siteiz or any third party, so it is safe for confidential contracts, financial statements, and other regulated documents.
Does it handle scanned PDFs (OCR), and what does it cost?
Converting text-based PDFs is free and unlimited, with no account. Scanned or image-only pages are detected automatically and read with local OCR, which is PDF Pro: 99 kr for 24 hours of unlimited OCR on any file. It is a single charge that stops after 24 hours and never renews. Teams that need OCR continuously use the annual business licence.
Does it work with RAG, LangChain, and LlamaIndex?
Yes, that is what it is built for. Convert the PDF to Markdown here, then load the .md with your framework's Markdown loader instead of a PDF loader. Reconstructed reading order keeps two-column papers from interleaving mid-sentence, Markdown headings give splitters real chunk boundaries, and tables stay as tables so row and column relationships survive into the embedding.
Why Markdown instead of plain text for LLMs?
Markdown preserves the structure that plain text discards: headings, lists, and tables. That structure improves chunking and retrieval in RAG pipelines and gives an LLM cleaner boundaries to cite, while staying compact and token-efficient.
How accurate is OCR on scanned documents?
OCR is a best-effort reading of the page image, so proofread anything critical. Clean, high-resolution scans of printed text read well. Low-resolution faxes, handwriting, heavy background imagery, and unusual typefaces are harder, and pages that are mostly photographs can produce stray characters.
How do I convert a PDF to Markdown for free?
Open the converter on this page, drop in your PDF, and copy or download the .md file. Text-based PDFs are free and unlimited, with no account and no page cap. The conversion runs in your browser, so there is no upload step and no waiting in a queue.
What is the difference between PDF to MD and PDF to Markdown?
They are the same thing. Markdown files use the .md extension, so a PDF to MD converter and a PDF to Markdown converter describe one conversion: a PDF turned into a structured .md text file.
Can I use this as a PDF to text converter?
Yes. Markdown is plain text, so the .md output opens in any text editor and can be renamed .txt if you need it. The advantage over a raw text dump is that headings, lists, and tables survive instead of collapsing into one undifferentiated block.
Is there a page or file size limit?
Text-based PDFs have no page limit and no file cap beyond what your browser can hold in memory, since the work happens on your own machine. Very large scanned documents run slower because OCR reads each page image locally.
Which languages does the converter handle?
Text-based PDFs convert in any language the file already contains, including Swedish and other Latin-script languages, because the text is read directly from the document. OCR for scanned pages is tuned for English and Latin-script text, so results on other scripts vary.
Do you offer a business or team licence?
Yes. The business licence is 20000 kr per year, billed annually by invoice. It covers unlimited OCR and native processing for commercial use across your organisation and gives your team a shared access link with no user registration. Because conversion happens in the browser, your documents stay on your own machines. Email hello@siteiz.com for a corporate invoice.
Ready to convert?
Text-based PDFs are free and unlimited. Scanned documents need OCR, which is a 99 kr pass for 24 hours. No subscription, no account.
Guides
Get more out of the conversion
- How to convert PDF to Markdown for RAGWhy PDF loaders break retrieval, and the LangChain and LlamaIndex swap that fixes it.
- How to OCR a scanned PDF into MarkdownTell a scan from a text PDF, know what OCR can read, and get better results from a bad scan.
- PDF to Markdown vs PDF to textBoth strip the packaging. Only one keeps the headings, lists, and tables AI needs.
Also from Siteiz
Curious how AI reads your live pages?
Try our free AI audit to see what AI systems understand and cite on your website.
Something converted badly?
PDFs are endlessly varied, and a layout that defeats the parser is a bug worth fixing. Send the file, or just describe it, and it becomes a test case.