Siteiz.
PDF to .md · runs in your browser · nothing uploaded

Turn messy PDFs into clean Markdown your AI can actually read.

Drop in a PDF and get clean Markdown back. Columns land in the right reading order and tables stay tables, so nothing reaches your model scrambled. It all runs inside your browser, which means your file never leaves your computer. Text PDFs are free.

Convert a PDF for free
  • Nothing is uploaded
  • Handles scanned pages too
  • No account, no sign-up
Correct reading orderReal Markdown tablesFree OCR for scans

FreeText-based PDFsUnlimited. No account. No sign-up.

PassScanned PDFs (OCR)99 kr for 24 hours, unlimited pages. Nothing renews.

Both run entirely in your browser. Business licence and batch conversion

Reconstructs multi-column reading order, extracts tables, heals hyphenation, and strips repeated headers, all 100% in your browser. Token counts are rough estimates. Converting text-based PDFs is free and unlimited. Scanned pages carry no text layer and need OCR, which is the PDF Pro pass (99 kr for 24 hours, unlimited pages, nothing renews): the engine (Tesseract, open source) downloads on first use and runs locally, so your document still never leaves your browser. OCR output is a best-effort reading of the page image, so proofread anything critical.

Siteiz PDF Pro, OCR to Markdown
PDF Pro · 99 kr / 24-hour pass

Got a scanned PDF?

If your pages are pictures rather than selectable text, OCR reads them. One payment covers unlimited pages on any file for 24 hours. Nothing renews and there is no account.

How it works

Three steps, about ten seconds.

  1. Step 1

    Drop your PDF

    Drag the file onto the box above, or click to choose one. It stays on your computer. Nothing is uploaded, so there is no waiting for a server.

  2. Step 2

    It rebuilds the document

    Columns are put back into reading order, tables are kept as tables, headings stay headings, and repeated page furniture is stripped out.

  3. Step 3

    Copy or download the .md

    Take the clean Markdown straight into ChatGPT, Claude, a RAG pipeline, or your notes. You also see roughly how many tokens you saved.

Two kinds of PDF

Is your PDF text, or a photo of text?

This one difference decides whether the tool is free for you. You can check in two seconds: open your PDF and try to select a sentence with your mouse.

Free

Text you can select

Exported from Word, Google Docs, InDesign, or almost any accounting or reporting system. The words are stored inside the file, so the text can be read directly.

Unlimited conversions. No account, no sign-up, no payment.

Siteiz PDF Pro, OCR to MarkdownPDF Pro · 99 kr / 24-hour pass

Text you cannot select

Scans, photographed pages, and faxes. The page is just a picture, so the words have to be recognised from the image first. That recognition step is called OCR.

99 kr opens a 24-hour pass with OCR for unlimited pages on any file. Nothing renews.

Before & after

What your model actually receives.

The same two-column page, converted two ways. Ordinary extractors read straight across the page, so the columns interleave and the sentences stop making sense.

Typical converter

The options for choosing a customised
interior design are 1. Brown/Light
practically limitless. To make your
choice a little easier, our 2. Brown/Dark
designers have put together a selection
of interiors with 3. Black/Dark
perfectly matched colours.

Siteiz

The options for choosing a customised
interior design are practically
limitless. To make your choice a little
easier, our designers have put together
a selection of interiors with perfectly
matched colours.

1. Brown/Light
2. Brown/Dark
3. Black/Dark

Left is what a scrambled read produces: option labels dropped into the middle of sentences. Right is the same page after reading order is reconstructed. Feed the left one to a model and it will answer from nonsense.

Benchmark

One 232-page annual report, measured.

We ran a real corporate annual report through the converter: dense financial tables, native text, and scanned pages that needed OCR. These are the file sizes we recorded, nothing modelled or rounded up.

MetricOriginal PDFSiteiz MarkdownChange
File size14.3 MB848 KB94.1% smaller
Per page61.6 KB3.6 KB17x lighter
Pages232, tables and scans includedText preserved

Financial tables held together

Multi-column financial statements came through as valid Markdown pipe tables rather than a flattened wall of numbers, so the rows and columns stay queryable.

Characters kept clean

Currency symbols, percentages, and punctuation survived intact, and heading levels mapped to the document's own hierarchy instead of being guessed.

Native text and OCR in one pass

The converter used the embedded text where the pages had it and OCR where they were scans, producing one continuous Markdown file with no seam to stitch.

This is a single real document, not an averaged benchmark suite. Your own reduction depends on how much of your PDF is layout overhead versus text. The tool shows you the before and after on every file, so you never take our word for it.

Why it is different

Structure in, not noise in.

Standard extractors read straight across the page and hand your model scrambled columns and flattened tables. Siteiz reconstructs the document first, so what reaches your pipeline is faithful to the original.

Correct reading order

Multi-column layouts, sidebars, and footnotes are reconstructed with recursive XY-cut segmentation before a single character is written. Text flows in true reading order instead of columns interleaved into noise, which is the difference between a grounded answer and a hallucination when the document reaches your retrieval step.

Real tables

Tabular data is detected by clustering glyph coordinates into rows and columns, then rendered as GitHub-flavored Markdown tables. A confidence threshold guards every table: when the grid is ambiguous, it falls back to clean text rather than emit a broken structure. Your embeddings keep the row and column relationships that give a table its meaning.

Your file never leaves the browser

Parsing runs entirely client-side with pdf.js, and OCR runs locally as WebAssembly. There is no upload, no server round trip, and no third party in the path. Contracts, financials, and research never leave the device. That is what makes this usable where documents simply cannot be sent to a SaaS backend.

For AI search and RAG

If an AI can't read it, it can't cite you.

Your white papers, reports and case studies are the content most worth quoting, and they are the content AI reads worst. Straight across the page, columns interleave and tables flatten. Whatever answer gets generated is built on that garbled version, and your brand does not get the citation.

Quotable passages

Headings and reading order stay intact, so an assistant can lift a coherent paragraph and attribute it. Scrambled text gets summarised vaguely or skipped, and vague summaries do not carry your name.

Your data stays data

Pricing tables, benchmarks and comparison grids come through as real Markdown tables. The numbers keep their rows and columns, so a model can answer questions about them instead of guessing.

Fewer tokens, lower cost

Repeated headers, footers and page numbers are stripped out. You see the token count before and after, so a 40-page report stops paying to send the same boilerplate forty times.

Pricing

Three plans. No sign-ups.

Text conversion is free forever. You only pay when you need OCR on scanned pages, and there is no account to create at any tier.

Standard
Free

For fast parsing of text-based PDFs.

  • Unlimited text-PDF conversions
  • 100% in-browser, nothing uploaded
  • Safe for confidential files
  • No account, no sign-up
Convert a PDF for free
24-hour passMost popular
99 kr/ 24 hours

For scanned documents and heavy reports.

  • Everything in Standard
  • OCR for scanned and image-only pages
  • Unlimited conversions for 24 hours
  • Tracked in your browser, no account
  • Nothing renews, no subscription
Business licence
20 000 kr/ year

For teams automating RAG and enterprise archives.

  • Unlimited OCR and native processing
  • Commercial use across your organisation
  • Documents stay on your own machines
  • Shared team access link, no user registration
  • Direct line to support and formatting roadmap
Request a corporate invoice

Prices in Swedish kronor, VAT handled at checkout. The 24-hour pass is billed once through Stripe and simply stops after 24 hours. The business licence is billed annually by invoice.

FAQ

Questions, answered

Is my document uploaded to a server?
No. Conversion runs entirely in your browser using pdf.js. The file never leaves your device, and nothing is sent to Siteiz or any third party. That makes it safe for confidential contracts, financial statements, and other regulated documents.
Does it support scanned PDFs or images (OCR)?
Yes. Pages with no text layer are detected automatically and read with built-in OCR that runs in your browser using Tesseract, an open-source engine. The engine downloads on first use, and your document still never leaves your device. OCR is part of PDF Pro, a 99 kr pass that covers unlimited scanned pages on any file for 24 hours.
What does it cost?
Converting text-based PDFs is free and unlimited, with no account and no sign-up. Scanned PDFs need OCR, which is the PDF Pro pass: 99 kr for 24 hours of unlimited OCR on any file. It is a single charge that simply stops after 24 hours. Nothing renews and you are not charged again.
How long does the 99 kr pass last?
24 hours from the moment your payment clears, not from first use. During that window you can convert unlimited scanned pages on any number of files. When the 24 hours are up, OCR stops and free text conversion continues as normal. If you need OCR again later, you buy another pass. Teams that need OCR continuously are better served by the annual business licence.
I bought a pass. What happens if I clear my browser data?
Your pass is stored in your browser rather than in an account, so clearing site data removes it from that browser. When you buy, the page shows a pass link: save it, and opening that link restores your pass on any browser or device for as long as the 24 hours are still running. Lost it? Email hello@siteiz.com from the address you paid with.
Do you offer a business or team licence?
Yes. The business licence is 20 000 kr per year, billed annually by invoice. It covers unlimited OCR and native processing for commercial use across your organisation, gives your team a shared access link with no user registration, and includes a direct line to support for custom formatting or roadmap requests. Because conversion happens entirely in the browser, your documents stay on your own machines, which is usually the reason regulated teams get in touch. Email hello@siteiz.com to request a corporate invoice.
How accurate is OCR on scanned documents?
OCR is a best-effort reading of the page image, so proofread anything critical. Clean, high-resolution scans of printed text read well. Low-resolution faxes, handwriting, heavy background imagery, and unusual typefaces are harder, and pages that are mostly photographs can produce stray characters.
Is there a file size limit?
Because processing is local, the practical limit is your device memory, not a server quota. Text PDFs of several hundred pages convert comfortably on a typical laptop. Very large or image-heavy files depend on available RAM.
How do I parse multi-column PDFs for LangChain or LlamaIndex?
Convert the PDF to Markdown here first, then load the .md with your framework's Markdown loader instead of a PDF loader. Most PDF loaders read straight across the page, so a two-column paper arrives with its columns interleaved mid-sentence and every chunk is corrupted. Reconstructing reading order before chunking is what makes retrieval return coherent passages.
Does this work for RAG pipelines and embeddings?
That is what it is built for. Markdown headings give recursive character or Markdown-header splitters real boundaries to chunk on, so a chunk ends at a section rather than mid-sentence. Tables stay as tables, so row and column relationships survive into the embedding. Repeated page headers and footers are stripped, so the same boilerplate is not embedded hundreds of times.
Can I convert PDF tables to Markdown without uploading data?
Yes. Table detection and the whole conversion run inside your browser, so no data is transmitted. Tables are found by clustering glyph coordinates into rows and columns and are only emitted as a Markdown grid when the structure is confident; ambiguous layouts fall back to clean text rather than a broken table.
Why Markdown instead of plain text for LLMs and RAG?
Markdown preserves the structure that plain text discards: headings, lists, and tables. That structure improves chunking, retrieval accuracy, and citation quality in RAG pipelines, while staying compact and token-efficient.
How is this different from other PDF to Markdown converters?
Most extractors read straight across the page, scrambling columns and flattening tables. Siteiz reconstructs the document with reading-order segmentation and confidence-gated table detection, and it does all of it locally, so structured output and data privacy come together instead of as a trade-off.

Ready to convert?

Text-based PDFs are free and unlimited. Scanned documents need OCR, which is a 99 kr pass for 24 hours. No subscription, no account.

Convert a PDF for free

For engineering and data teams

Stop paying to ingest bloated PDFs.

Every scrambled column and flattened table you feed a model is tokens spent on noise, and every raw PDF you embed is vector storage spent on layout overhead. Clean, structured Markdown is smaller going in and more accurate coming out. In our own test a 14.3 MB report became 848 KB of Markdown, and the tool shows the token count on every file so your team can see the ingestion saving per document rather than trust a headline figure.

The business licence covers unlimited OCR and native processing across your organisation, keeps every document on your own machines, and gives your team a direct line to support for custom formatting or roadmap requests. Access is a single shared link, with no IT onboarding and no user registration.

Something converted badly?

PDFs are endlessly varied, and a layout that defeats the parser is a bug worth fixing. Send the file, or just describe it, and it becomes a test case.

Built for engineers and teams preparing documents for LLMs and AI search: turn contracts, datasheets, research, and knowledge bases into content a model can actually read.