--- snapshot-1789251323+++ snapshot-1789727250@@ -2,7 +2,11 @@
Release notes from paperless-gpt
-2026-07-21T10:06:41Z tag:github.com,2008:Repository/861746639/v0.27.0 2026-07-21T10:13:57Z
+2026-09-18T09:55:36Z tag:github.com,2008:Repository/861746639/v0.28.0 2026-09-18T10:09:19Z
+
+v0.28.0 โ The Follow-Through Expansion
+
+
Automation you can leave alone. This release is almost entirely about paperless-gpt doing what its logs already claimed it was doing โ tags that were computed and then silently dropped, retry loops that never concluded, and settings that quietly had no effect.
The theme is follow-through: if the app says it added a tag, the tag is there; if a document can't be processed, it stops being retried and says so; if you set an option, it reaches the model.
Heads-up โ two behaviour changes worth reading before you upgrade. See Notable changes at the bottom.
๐ท๏ธ Tags that actually get applied
Five separate defects conspired to make completion tags vanish between the log line and paperless-ngx. Every one of them is fixed.
AUTO_TAG_COMPLETE is applied instead of being dropped. The tag was written to an AddTags field that nothing ever read, so it was computed, logged as "adding", and discarded. A second tag path โ taken when no other field changed, which is the likeliest case for an already-tidy document โ ignored it too. Thanks @skort-90 for the four-document reproduction and @Denim5660 for the original root-cause analysis. (#1059, closes #1006, #1074) PDF_OCR_COMPLETE_TAG is created at startup, like FAIL_TAG already was. Without it, the nameโID resolution silently skipped a tag that didn't exist yet โ which is why this worked for some people and not others, with nothing in the log to explain the difference. (#1064, closes #854) - Tag names now resolve case-insensitively. The rest of the tag handling already ignored case; the nameโID lookup was the odd one out, so a configured
Paperless-GPT-Auto-Complete couldn't find a stored paperless-gpt-auto-complete. Where a database holds several case variants, the original (lowest id) wins deterministically and the ambiguity is logged. Thanks @ivanzud! (#1075) - An added tag now wins a name collision with a removed one. Configuring the completion tag to the same name as the trigger tag used to leave the document with neither. (#1059, closes #458)
- paperless-gpt's own tags are never offered to the LLM.
FAIL_TAG, AUTO_TAG_COMPLETE and PDF_OCR_COMPLETE_TAG were missing from the exclusion list, so the model suggested them as ordinary document tags. With an OCRโtagging workflow that produced an unbounded loop, re-billing every LLM call on each pass. Thanks @hakehardware and @fr0der1c for reporting it, twice. (#1061, closes #877, #1015)
๐ Loops that end
Three ways the auto pipeline can fail. Until now, only two of them could stop.
- Repeated suggestion failures now break the loop. A permanent LLM error โ a prompt that can't fit the context, a model that no longer exists โ used to re-bill the same request forever: one report measured ~12,000 rejected requests over seven days. Worse, because the poll fetches a single unordered page of 25, enough stuck documents starved everything behind them (67 documents, with the container still reporting healthy). After
AUTO_TAG_MAX_RETRIES attempts the auto tag comes off and FAIL_TAG goes on. Thanks @JoGres-DE for an exceptionally precise report. (#1076, closes #1071)
This should also end the "worker goes permanently silent" reports: the poll loop doubles its backoff on every failing cycle, capped at one hour, so a single permanently failing document kept the whole worker in hour-long sleeps until a container restart. With the document leaving the queue, the cycle succeeds and the backoff resets. (likely #1016) - Ollama requests can no longer hang forever. Both Ollama paths ran on a timeout-less HTTP client, so a single stalled generation blocked the background loop indefinitely โ recoverable only by restarting the container.
OLLAMA_TIMEOUT_SECONDS (default 300s) bounds it. Thanks @swiception! (#1070)
Note this covers the Ollama paths only โ other providers still share an unbounded HTTP client, so #937 stays open.
๐ง Ollama: OLLAMA_THINK finally does something
OLLAMA_THINK had no effect at all, in either direction. The underlying library nested think inside the request's options object, where the Ollama server doesn't look for it, and dropped the value entirely when it was false.
On thinking-capable models (Qwen 3 / 3.5, Gemma 3) that meant unconditional reasoning: the model spent its output budget on a reasoning trace and returned results that were glacial or empty โ worst exactly where it matters most, on closed-list classification and strict JSON.
The metadata path now uses Ollama's official client, where think is a top-level field. Reasoning levels (low / medium / high) work too. Thanks @lunetics for the diagnosis and the implementation, and @andglaser for the independent report. (#1063, closes #1024, #1056)
The same change brings LLM_TEMPERATURE for the text LLM (closes #1033), plus LLM_MAX_TOKENS and OLLAMA_KEEP_ALIVE. Ollama-only variables now warn instead of being silently inert when another provider is configured, and OLLAMA_HEADERS is treated as a secret so credential-bearing headers stay out of the configuration view.
๐ฌ OCR
OCR_LIMIT_PAGES=0 means all pages again โ an explicit zero was being mistaken for "unset" and silently replaced by the 5-page default. Thanks @CBOSSX! (#1047, closes #1037) - Transient provider errors are retried per page (HTTP 429/5xx, exponential backoff) instead of forfeiting the whole document. Deliberately more patient than the suggestion path: a failed page throws away every page before it. Thanks @MarcvsTvllivs! (#1004)
- Uploaded files are deleted from Mistral after OCR. They were being left in Mistral's cloud storage indefinitely โ a genuine data-retention problem for a self-hosted document tool, and invisible from the outside. Thanks @lunetics! (#1007)
- Oversized page images compress correctly. After stepping quality down to fit a size limit, the last-resort resize re-encoded at a higher quality than it had just settled on, partly undoing the reduction. Thanks @maksyms! (#946)
image_limit and image_min_size are configurable for Mistral OCR โ raise the minimum to stop small boxed fields (handwritten form entries) being skipped as images. Thanks @jipe-b! (#1050, closes #1048)
๐ Using a different AI provider
Five separate requests โ issues and pull requests โ asked for providers that already shipped, one of them implementing a whole provider branch that did nothing OPENAI_BASE_URL didn't already do. That was a documentation failure on our side, not missing features.
OpenAI-compatible providers is a new guide with copy-pasteable configuration for OpenRouter, LM Studio, vLLM, LiteLLM, llama.cpp and Azure โ plus the three things everyone gets wrong (a base URL missing /v1, vendor-specific model names, local servers rejecting an empty API key) and a troubleshooting section mapping the recurring errors to their causes. (#1060, closes #864, #908, #1041)
Anything that documents an "OpenAI-compatible endpoint" works today via LLM_PROVIDER=openai plus OPENAI_BASE_URL. No new release required.
โ๏ธ New options
| Variable | What it does |
AUTO_TAG_MAX_RETRIES | Give up on a document after N failed suggestion attempts (default 3, 0 retries forever) |
OLLAMA_TIMEOUT_SECONDS | Per-request timeout for Ollama (default 300) |
PRESERVE_EXISTING_METADATA | Keep a correspondent or document type that is already set, so paperless-ngx' own classifier or a manual correction stays in charge. Thanks @keefar! (#1065, closes #1032) |
CORRESPONDENT_PROMPT_LIMIT | Cap how many correspondents go into the prompt. On a 13k-document instance with 628 correspondents, the correspondent step went from a >10-minute timeout per document to seconds. Thanks @interruptor! (#1043) |
LLM_TEMPERATURE, LLM_MAX_TOKENS, OLLAMA_KEEP_ALIVE | Ollama metadata generation |
MISTRAL_OCR_IMAGE_LIMIT, MISTRAL_OCR_IMAGE_MIN_SIZE | Mistral OCR image extraction |
๐ Privacy & deployment
- Document content is no longer logged at info level. Full document text was landing in
docker logs and any log shipper. Thanks @mrab54! (#922, closes #921) - Non-root containers start correctly โ Kubernetes
runAsNonRoot and docker run --user now skip the privilege-drop path instead of failing. Thanks @vistalba! (#1002)
๐ฆ Under the hood
- The end-to-end test suite works against paperless-ngx 3.x again (it refuses to start on the documented placeholder secret key, so the container died during init) and is pinned by digest instead of tracking
:latest. Extracted from @lunetics' work in #1058. (#1062) - Container images are assembled without a redundant cross-registry copy, which was exhausting Docker Hub pull quota on busy days. (#1068)
- Dependencies refreshed, including Node 24 and Alpine 3.24.
โ ๏ธ Notable changes
1. Documents whose suggestions keep failing now leave the queue. Previously they were retried forever. After AUTO_TAG_MAX_RETRIES attempts (default 3) the auto tag is removed and FAIL_TAG applied. If you relied on indefinite retries, set AUTO_TAG_MAX_RETRIES=0. This mirrors what OCR_MAX_RETRIES already did for OCR in v0.27.0.
2. Ollama requests now time out after 300 seconds. Previously they could hang indefinitely. If you run very large models on slow hardware and a legitimate generation exceeds five minutes, raise OLLAMA_TIMEOUT_SECONDS or set it to 0 to restore the old behaviour.
Both defaults were chosen so the failure mode is "this document is marked for review" rather than "the worker is silently wedged".
๐ Credits
This release came almost entirely from the community โ reports, diagnoses and code.
- @lunetics (Matthias Breddin) โ the
OLLAMA_THINK diagnosis and native-client implementation (#1063), the Mistral data-retention fix (#1007), and the E2E fixes that unblocked CI for every open PR (#1062) - @swiception โ the Ollama request timeout (#1070)
- @ivanzud (Ivan) โ case-insensitive tag resolution, including the determinism argument that stopped it being subtly wrong (#1075)
- @MarcvsTvllivs โ per-page OCR retry with backoff (#1004)
- @interruptor โ
CORRESPONDENT_PROMPT_LIMIT, with measurements (#1043) - @keefar โ
PRESERVE_EXISTING_METADATA (#1065) - @CBOSSX โ the explicit-zero page limit fix (#1047)
- @jipe-b โ Mistral image parameters (#1050)
- @maksyms โ the image re-encoding fix (#946)
- @vistalba โ non-root container support (#1002)
- @mrab54 โ content logging at debug level (#922)
- @JoGres-DE, @skort-90, @Denim5660, @hakehardware, @fr0der1c, @andglaser, @nmeden, @vanderfran, @embdMan, @jacobhausler, @Christoph274 โ reports precise enough to fix from, several with the root cause already correctly identified
Full Changelog: v0.27.0...v0.28.0
icereed tag:github.com,2008:Repository/861746639/v0.27.0 2026-07-21T10:13:57Z
v0.27.0 โ The Clarity Expansion
@@ -38,8 +42,4 @@
v0.21.0
-Release Highlights ๐
New Features
๐ฎ Mistral OCR Integration with Advanced PDF Processing
- Extended PDF processing support - Mistral OCR now joins Google Document AI in supporting all processing modes:
image, pdf, and whole_pdf - Cost-effective OCR - Purpose-built OCR endpoint optimized for document processing with competitive pricing
- Markdown-formatted output - Returns well-structured markdown text that preserves document formatting and layout
- Large document support - Handles files up to 50MB and 1,000 pages efficiently
- Set
OCR_PROVIDER: "mistral_ocr" and configure your Mistral API key to get started
๐ท๏ธ Enhanced Title Generation with Context
- Original title context - Title generation now includes the existing document title as contextual information
- Improved relevance - Language models can use the original title to generate more accurate and contextually appropriate suggestions
- Better continuity - Maintains document naming consistency while enhancing title quality
- Smart fallbacks - Handles cases where original titles are missing or incomplete
Improvements & Refinements
๐ก๏ธ Configuration Validation
- OCR provider compatibility checks - Prevents invalid combinations of OCR providers and processing modes
- Clear error messages - Detailed feedback when unsupported mode combinations are detected
- Startup validation - Early detection of configuration issues before processing begins
- Provider-specific guidance - Helpful error messages explain which modes are supported by each provider
๐ Enhanced PDF Processing Architecture
- Hybrid file naming - Improved PDF splitting with standardized naming conventions that maintain backward compatibility
- More provider choice - Users can now choose between Google Document AI and Mistral OCR for advanced PDF processing
- Consistent behavior - Both advanced providers support
pdf and whole_pdf modes with similar performance characteristics
๐งช Comprehensive E2E Testing
- Mistral OCR test suite - Full end-to-end testing of Mistral OCR integration with real PDF documents
- Processing mode validation - Tests verify
whole_pdf mode works correctly with multi-page documents - Performance metrics - Test output includes detailed comparison of original vs. enhanced OCR content
- Cross-provider compatibility - Tests ensure consistent behavior across different OCR providers
Documentation Updates
๐ OCR Provider Comparison
- Updated provider documentation - Clear explanation of which providers support which processing modes
- Mode compatibility matrix - Easy reference for choosing the right provider and mode combination
- Mistral-specific guidance - Detailed setup instructions and best practices for Mistral OCR
- Configuration examples - Complete docker-compose examples for all supported configurations
Technical Details
Provider Mode Support Matrix
| Provider | image | pdf | whole_pdf |
| LLM (OpenAI/Ollama) | โ
| โ | โ |
| Azure Document Intelligence | โ
| โ | โ |
| Google Document AI | โ
| โ
| โ
|
| Mistral OCR (New!) | โ
| โ
| โ
|
| Docling | โ
| โ | โ |
What's Changed
- feat: Add Mistral OCR provider with advanced PDF processing support - Extends
pdf and whole_pdf mode support to a second provider - feat: Add OCR provider and processing mode validation - Prevents misconfigurations and provides helpful error messages
- feat: Pass original document title to title generation prompt - Improves context and relevance of AI-generated titles #453
- feat: Implement hybrid PDF naming strategy - Improved file naming with backward compatibility
- test: Add comprehensive Mistral OCR E2E tests - Full test coverage including diff comparison utilities
- docs: Update OCR processing modes documentation - Clear provider compatibility information
Configuration Example
environment: # Mistral OCR (new advanced PDF support) OCR_PROVIDER: "mistral_ocr" MISTRAL_API_KEY: "your_mistral_api_key" MISTRAL_MODEL: "mistral-ocr-latest" # Optional OCR_PROCESS_MODE: "whole_pdf" # Now supported!
Migration Notes
- No breaking changes - Existing configurations continue to work as expected
- More provider choice - Users now have two options for advanced PDF processing (
pdf and whole_pdf modes)
Performance Benefits
- Provider flexibility - Choose between Google Document AI and Mistral OCR based on your needs and pricing preferences
- Reduced API calls -
whole_pdf mode processes entire documents in one request (now available with both advanced providers) - Better accuracy - Direct PDF processing maintains document structure and formatting
- Smarter title generation - Original title context leads to more relevant AI suggestions
PRs
- fix(deps): update module github.com/pdfcpu/pdfcpu to v0.11.0 by @renovate in #434
- fix: mislabeled data types in azure types by @moarsmokes in #455
- chore(deps): update react monorepo to v19.1.7 by @renovate in #429
- chore(deps): update dependency @vitejs/plugin-react-swc to v3.10.2 by @renovate in #424
- Enhance title suggestions with original title by @icereed in #466
- [mistral-ocr] Add MIME type detection, structured logging, and improvโฆ by @icereed in #468
Full Changelog: v0.20.0...v0.21.0
icereed tag:github.com,2008:Repository/861746639/v0.20.0 2025-05-30T14:39:51Z
-
-v0.20.0
-
-Release Highlights ๐
New Features
๐ง Google Gemini AI Integration
- Added Google Gemini AI support - Paperless-GPT now supports Google's Gemini AI models as a new LLM provider option
- Thinking budget support - Leverages Gemini's new thinking capabilities for enhanced document processing
- Enhanced error handling - Improved API response validation and error management for Google AI services
- Set
LLM_PROVIDER: "googleai" and configure your Google AI API credentials to get started
Improvements & Refinements
๐ง LLM Prompt Optimization
- Enhanced prompt structure - Added XML-like separators to LLM prompts for improved parsing accuracy and consistency
- Better data organization - Input data is now enclosed in structured tags for clearer LLM interpretation
- Improved reliability - More consistent results across different document types and LLM providers
๐ PDF Processing Fixes
- Fixed split PDF logic - Updated file naming pattern to match
pdfcpu output format (removes zero-padding) - Consistent naming - Split files now use simplified naming:
original_1.pdf, original_2.pdf, etc. - Better workflow integration - Improved compatibility with existing PDF processing pipelines
Documentation Updates
๐ Model Recommendations
- Updated model suggestions - Documentation now recommends
qwen3:8b instead of deepseek-r1:8b for Ollama users - Better performance -
qwen3:8b offers more recent and powerful reasoning capabilities - Improved examples - Updated configuration examples throughout the documentation
Dependencies & Maintenance
๐ Dependency Updates
- testcontainers updated to v10.28.0 - Latest testing framework improvements
- globals updated to v16.2.0 - Enhanced JavaScript globals definitions
- Automated maintenance - Renovate bot ensures dependencies stay current and secure
Technical Details
What's Changed
- adds support for new gemini models with thinking budget #441
- refactor: Add XML-like separators to LLM prompts for improved parsing #442
- fix: the split pdf logic to be consistent with the output from pdfcpu #435
- doc: mention qwen3:8b instead of deepseek-r1:8b #439
- chore(deps): update dependency testcontainers to v10.28.0 #425
- chore(deps): update dependency globals to v16.2.0 #428
Contributors
Special thanks to @thiswillbeyourgithub, @dawidkulpa, and @moarsmokes for their contributions to this release!
Configuration Notes
- For Google Gemini AI: Set environment variables for
GOOGLE_AI_API_KEY and configure LLM_PROVIDER: "googleai" - No breaking changes - existing configurations continue to work as expected
- Consider updating to
qwen3:8b model if using Ollama for better performance
New Contributors
Full Changelog: v0.19.0...v0.20.0
icereed+Release Highlights ๐
New Features
๐ฎ Mistral OCR Integration with Advanced PDF Processing
- Extended PDF processing support - Mistral OCR now joins Google Document AI in supporting all processing modes:
image, pdf, and whole_pdf - Cost-effective OCR - Purpose-built OCR endpoint optimized for document processing with competitive pricing
- Markdown-formatted output - Returns well-structured markdown text that preserves document formatting and layout
- Large document support - Handles files up to 50MB and 1,000 pages efficiently
- Set
OCR_PROVIDER: "mistral_ocr" and configure your Mistral API key to get started
๐ท๏ธ Enhanced Title Generation with Context
- Original title context - Title generation now includes the existing document title as contextual information
- Improved relevance - Language models can use the original title to generate more accurate and contextually appropriate suggestions
- Better continuity - Maintains document naming consistency while enhancing title quality
- Smart fallbacks - Handles cases where original titles are missing or incomplete
Improvements & Refinements
๐ก๏ธ Configuration Validation
- OCR provider compatibility checks - Prevents invalid combinations of OCR providers and processing modes
- Clear error messages - Detailed feedback when unsupported mode combinations are detected
- Startup validation - Early detection of configuration issues before processing begins
- Provider-specific guidance - Helpful error messages explain which modes are supported by each provider
๐ Enhanced PDF Processing Architecture
- Hybrid file naming - Improved PDF splitting with standardized naming conventions that maintain backward compatibility
- More provider choice - Users can now choose between Google Document AI and Mistral OCR for advanced PDF processing
- Consistent behavior - Both advanced providers support
pdf and whole_pdf modes with similar performance characteristics
๐งช Comprehensive E2E Testing
- Mistral OCR test suite - Full end-to-end testing of Mistral OCR integration with real PDF documents
- Processing mode validation - Tests verify
whole_pdf mode works correctly with multi-page documents - Performance metrics - Test output includes detailed comparison of original vs. enhanced OCR content
- Cross-provider compatibility - Tests ensure consistent behavior across different OCR providers
Documentation Updates
๐ OCR Provider Comparison
- Updated provider documentation - Clear explanation of which providers support which processing modes
- Mode compatibility matrix - Easy reference for choosing the right provider and mode combination
- Mistral-specific guidance - Detailed setup instructions and best practices for Mistral OCR
- Configuration examples - Complete docker-compose examples for all supported configurations
Technical Details
Provider Mode Support Matrix
| Provider | image | pdf | whole_pdf |
| LLM (OpenAI/Ollama) | โ
| โ | โ |
| Azure Document Intelligence | โ
| โ | โ |
| Google Document AI | โ
| โ
| โ
|
| Mistral OCR (New!) | โ
| โ
| โ
|
| Docling | โ
| โ | โ |
What's Changed
- feat: Add Mistral OCR provider with advanced PDF processing support - Extends
pdf and whole_pdf mode support to a second provider - feat: Add OCR provider and processing mode validation - Prevents misconfigurations and provides helpful error messages
- feat: Pass original document title to title generation prompt - Improves context and relevance of AI-generated titles #453
- feat: Implement hybrid PDF naming strategy - Improved file naming with backward compatibility
- test: Add comprehensive Mistral OCR E2E tests - Full test coverage including diff comparison utilities
- docs: Update OCR processing modes documentation - Clear provider compatibility information
Configuration Example
environment: # Mistral OCR (new advanced PDF support) OCR_PROVIDER: "mistral_ocr" MISTRAL_API_KEY: "your_mistral_api_key" MISTRAL_MODEL: "mistral-ocr-latest" # Optional OCR_PROCESS_MODE: "whole_pdf" # Now supported!
Migration Notes
- No breaking changes - Existing configurations continue to work as expected
- More provider choice - Users now have two options for advanced PDF processing (
pdf and whole_pdf modes)
Performance Benefits
- Provider flexibility - Choose between Google Document AI and Mistral OCR based on your needs and pricing preferences
- Reduced API calls -
whole_pdf mode processes entire documents in one request (now available with both advanced providers) - Better accuracy - Direct PDF processing maintains document structure and formatting
- Smarter title generation - Original title context leads to more relevant AI suggestions
PRs
- fix(deps): update module github.com/pdfcpu/pdfcpu to v0.11.0 by @renovate in #434
- fix: mislabeled data types in azure types by @moarsmokes in #455
- chore(deps): update react monorepo to v19.1.7 by @renovate in #429
- chore(deps): update dependency @vitejs/plugin-react-swc to v3.10.2 by @renovate in #424
- Enhance title suggestions with original title by @icereed in #466
- [mistral-ocr] Add MIME type detection, structured logging, and improvโฆ by @icereed in #468
Full Changelog: v0.20.0...v0.21.0
icereed