Paperless-ngx

docs.paperless-ngx.comcontributed by Samuele Ongaro

YES

Replaceable in one session with an AI coding agent.

A document manager for scanned paper: it reads the text, tags documents automatically, and makes years of correspondence searchable.

Promptfree, for everyone, and the only version there is
Build me a document archive in the shape Paperless-ngx has: scanned paper, read, tagged and searchable, on my own machine — and read the warning first.

Read this first: Paperless-ngx is free and excellent, so nothing is being saved in fees. What you take on is real: the server, the storage, and the certainty that the only copy of a decade of correspondence does not vanish. If you are not going to run a verified off-site backup, do not put your only copies here.

STACK
- Node 20+ with Fastify
- SQLite through better-sqlite3, WAL mode, with FTS5
- tesseract for text recognition, ghostscript and qpdf for PDF handling, ImageMagick for images
- A worker for the consumption pipeline
- Caddy in front, on a private network

THE DATA MODEL
- documents: id, title, original_path, archive_path, sha256_original, sha256_archive, mime, bytes, page_count, created_date, added_at, modified_at, correspondent_id, document_type_id, storage_path_id, asn, content_text, language, ocr_confidence, deleted_at
- correspondents, document_types, tags, storage_paths — the four dimensions everything is filed by
- document_tags: document_id, tag_id
- matching_rules: id, target_kind, target_id, kind, pattern, is_insensitive — how a document gets filed automatically
- notes, custom_fields, custom_field_values
- tasks: id, kind, document_id, status, attempts, error, at — the consumption queue
- audit: id, document_id, action, actor, before_json, after_json, at — append-only
- The original file is never modified. Ever. The archive version with its text layer is a separate file, and both hashes are recorded

CONSUMPTION, WHICH IS THE PIPELINE THAT MATTERS
- A watched folder, an email address, and an upload endpoint. A scanner drops a file and everything else happens by itself
- Deduplicate by hash before anything else: scanning the same page twice is the most common event in this system
- Split multi-document scans on a separator page or a barcode, which is what makes a stack of post into a stack of documents
- Rotate by detected orientation, deskew, and remove blank pages, with the original untouched
- Text recognition with the right language, producing a searchable PDF that keeps the original image and adds an invisible text layer. Never replace the image with recognised text — the recognition is imperfect and the scan is the evidence
- Skip recognition where the PDF already has real text, and record which path was taken
- Record the confidence. A document recognised badly is a document you will never find again, and a low-confidence flag is what gets it re-scanned

FILING AUTOMATICALLY
- Matching rules per correspondent, type and tag: a literal phrase, any word, all words, a regular expression, or a fuzzy match
- Rules run against the recognised text and the filename; every match recorded with which rule fired, so a wrong assignment is diagnosable
- Optionally a simple learned classifier trained on what you have already filed by hand — but always as a suggestion the first time, never as a silent decision
- The date parsed from the document's own content in the formats your correspondence uses, falling back to the filename and then the file time, with the source recorded

FINDING THINGS, WHICH IS THE WHOLE POINT
- FTS5 over recognised text, titles, correspondents and notes, with stemming for your language
- Filters that compose: correspondent, type, tag, date range, and whether it has been read
- The matching text shown in the result with the page number, and the viewer opening at that page
- Saved searches as views
- An archive serial number, printed on the paper you keep, so a physical document and its scan can find each other in both directions

STORAGE
- Files laid out on disk by a template — year, correspondent, title — so they are usable with no software at all. That property is worth more than any feature here
- The database can be rebuilt from the files and their sidecars; write the rebuild command and test it
- Optional encryption at rest, with the honest note that server-side encryption protects against a stolen disk and nothing else

KEEPING IT
- Nightly backup of the database and the documents, off the machine, encrypted, with retention
- A verification job that re-hashes a sample of files and reports any that have changed. Silent corruption over ten years is exactly the failure this archive is meant to prevent
- A restore rehearsed onto a clean machine with the time it took written down
- Export everything — files, metadata as JSON, and the tag structure — in one command

THE INTERFACE
- A list with thumbnails, a viewer with the text layer selectable, and inline editing of the four dimensions
- Bulk editing, because filing happens in batches
- Fast on a phone, since that is where a photograph of a letter starts
- Dark and light

OPERATIONS
- .env: DATABASE_PATH, MEDIA_PATH, CONSUME_PATH, BASE_URL, SESSION_SECRET, OCR_LANGUAGES, IMAP_URL
- Migrations on boot, each once
- Behind a VPN rather than exposed; this is every piece of paper you own
- Health endpoint reporting queue depth, free space and the age of the last verified backup

WHAT MATTERS MOST
The consumption pipeline and the backup. Get deduplication, recognition and filing working on a hundred real documents, then set up the off-site copy and restore from it before you shred anything. The entire value of this system is that it is still there in ten years.

Give me the repository, the consumption worker, migrations, .env.example, the backup, restore and rebuild scripts, and a README that opens with the backup rehearsal.

What you lose

  • Nothing in fees; what you take on is the server, the storage and the backups
  • The certainty that a large provider will not lose the only copy of a document
  • Mobile scanning apps as polished as the commercial ones

If you would rather not build

  • A folder with a naming convention, which genuinely works

What it costs

as published on their pricing page

PlanBilled monthlyBilled yearlyLast read
—$0/mo——

Their pricing page is where these came from. Seeing a different price? Tell us.

The escape hatch

open source · no votes, no paid placement

Paperless-ngx

$0

The product itself: OCR, tagging and full-text search over scans.

paperless-ngx/paperless-ngxfree · open source

Docspell

$0

Self-hosted document organisation with automatic metadata extraction.

eikek/docspellfree · open source

Why this verdict

our own opinion · changed only by a person

87/100

Verdict yes at 87, and it costs nothing. The honest work is designing an intake you will actually use, and backing up files that may be the only copy.

History

tracked since 10 Aug 2026 · nothing is ever overwritten

Interest · last 30 dayspeak 2/day
views01230 Aug4 Sept9 Sept14 Sept19 Sept24 Sept28 Sept
— views— prompt copies none yet— votes none yet

Questions about Paperless-ngx

answered from the record above

Is Paperless-ngx free?

No — the plan we track is $0 a month. Free and open source; the cost is the server it runs on.

Can you replace Paperless-ngx by building your own?

YES. Replaceable in one session with an AI coding agent. Replacement score 87 out of 100, build time one session. Read what you lose before you decide.

How much does Paperless-ngx cost?

$0 a month on Free — $0 a year. Recorded 10 Aug 2026.

What do you lose by replacing Paperless-ngx?

Nothing in fees; what you take on is the server, the storage and the backups; The certainty that a large provider will not lose the only copy of a document; Mobile scanning apps as polished as the commercial ones. If any of those carry weight for you, keep paying.

Is there an open-source alternative to Paperless-ngx?

Yes: Paperless-ngx, Docspell. The prompt on this page is for when you want it your way instead.

Related entries

same category first, most replaced first

All 40 in Notes, docs & writing

Not sending yet

Every week, something stops being worth paying for.

New verdicts, prices that moved, entries added. One email a week. Unsubscribe in one click. Nothing is being sent yet — your address is kept here, and the first issue is the first thing it is used for.

free forever · no tracking pixel · stored here, never passed to anyone

Esc