--- snapshot-1789251323+++ snapshot-1789727250@@ -2,7 +2,11 @@ Release notes from paperless-gpt -2026-07-21T10:06:41Z tag:github.com,2008:Repository/861746639/v0.27.0 2026-07-21T10:13:57Z +2026-09-18T09:55:36Z tag:github.com,2008:Repository/861746639/v0.28.0 2026-09-18T10:09:19Z + +v0.28.0 โ€” The Follow-Through Expansion + +

Automation you can leave alone. This release is almost entirely about paperless-gpt doing what its logs already claimed it was doing โ€” tags that were computed and then silently dropped, retry loops that never concluded, and settings that quietly had no effect.

The theme is follow-through: if the app says it added a tag, the tag is there; if a document can't be processed, it stops being retried and says so; if you set an option, it reaches the model.

Heads-up โ€” two behaviour changes worth reading before you upgrade. See Notable changes at the bottom.


๐Ÿท๏ธ Tags that actually get applied

Five separate defects conspired to make completion tags vanish between the log line and paperless-ngx. Every one of them is fixed.

๐Ÿ” Loops that end

Three ways the auto pipeline can fail. Until now, only two of them could stop.

๐Ÿง  Ollama: OLLAMA_THINK finally does something

OLLAMA_THINK had no effect at all, in either direction. The underlying library nested think inside the request's options object, where the Ollama server doesn't look for it, and dropped the value entirely when it was false.

On thinking-capable models (Qwen 3 / 3.5, Gemma 3) that meant unconditional reasoning: the model spent its output budget on a reasoning trace and returned results that were glacial or empty โ€” worst exactly where it matters most, on closed-list classification and strict JSON.

The metadata path now uses Ollama's official client, where think is a top-level field. Reasoning levels (low / medium / high) work too. Thanks @lunetics for the diagnosis and the implementation, and @andglaser for the independent report. (#1063, closes #1024, #1056)

The same change brings LLM_TEMPERATURE for the text LLM (closes #1033), plus LLM_MAX_TOKENS and OLLAMA_KEEP_ALIVE. Ollama-only variables now warn instead of being silently inert when another provider is configured, and OLLAMA_HEADERS is treated as a secret so credential-bearing headers stay out of the configuration view.

๐Ÿ”ฌ OCR

๐Ÿ”Œ Using a different AI provider

Five separate requests โ€” issues and pull requests โ€” asked for providers that already shipped, one of them implementing a whole provider branch that did nothing OPENAI_BASE_URL didn't already do. That was a documentation failure on our side, not missing features.

OpenAI-compatible providers is a new guide with copy-pasteable configuration for OpenRouter, LM Studio, vLLM, LiteLLM, llama.cpp and Azure โ€” plus the three things everyone gets wrong (a base URL missing /v1, vendor-specific model names, local servers rejecting an empty API key) and a troubleshooting section mapping the recurring errors to their causes. (#1060, closes #864, #908, #1041)

Anything that documents an "OpenAI-compatible endpoint" works today via LLM_PROVIDER=openai plus OPENAI_BASE_URL. No new release required.

โš™๏ธ New options

Variable What it does
AUTO_TAG_MAX_RETRIES Give up on a document after N failed suggestion attempts (default 3, 0 retries forever)
OLLAMA_TIMEOUT_SECONDS Per-request timeout for Ollama (default 300)
PRESERVE_EXISTING_METADATA Keep a correspondent or document type that is already set, so paperless-ngx' own classifier or a manual correction stays in charge. Thanks @keefar! (#1065, closes #1032)
CORRESPONDENT_PROMPT_LIMIT Cap how many correspondents go into the prompt. On a 13k-document instance with 628 correspondents, the correspondent step went from a >10-minute timeout per document to seconds. Thanks @interruptor! (#1043)
LLM_TEMPERATURE, LLM_MAX_TOKENS, OLLAMA_KEEP_ALIVE Ollama metadata generation
MISTRAL_OCR_IMAGE_LIMIT, MISTRAL_OCR_IMAGE_MIN_SIZE Mistral OCR image extraction

๐Ÿ”’ Privacy & deployment

๐Ÿ“ฆ Under the hood


โš ๏ธ Notable changes

1. Documents whose suggestions keep failing now leave the queue. Previously they were retried forever. After AUTO_TAG_MAX_RETRIES attempts (default 3) the auto tag is removed and FAIL_TAG applied. If you relied on indefinite retries, set AUTO_TAG_MAX_RETRIES=0. This mirrors what OCR_MAX_RETRIES already did for OCR in v0.27.0.

2. Ollama requests now time out after 300 seconds. Previously they could hang indefinitely. If you run very large models on slow hardware and a legitimate generation exceeds five minutes, raise OLLAMA_TIMEOUT_SECONDS or set it to 0 to restore the old behaviour.

Both defaults were chosen so the failure mode is "this document is marked for review" rather than "the worker is silently wedged".


๐Ÿ™ Credits

This release came almost entirely from the community โ€” reports, diagnoses and code.

Full Changelog: v0.27.0...v0.28.0

icereed tag:github.com,2008:Repository/861746639/v0.27.0 2026-07-21T10:13:57Z v0.27.0 โ€” The Clarity Expansion @@ -38,8 +42,4 @@ v0.21.0 -

Release Highlights ๐Ÿš€

New Features

๐Ÿ”ฎ Mistral OCR Integration with Advanced PDF Processing

๐Ÿท๏ธ Enhanced Title Generation with Context

Improvements & Refinements

๐Ÿ›ก๏ธ Configuration Validation

๐Ÿ“„ Enhanced PDF Processing Architecture

๐Ÿงช Comprehensive E2E Testing

Documentation Updates

๐Ÿ“š OCR Provider Comparison

Technical Details

Provider Mode Support Matrix

Provider image pdf whole_pdf
LLM (OpenAI/Ollama) โœ… โŒ โŒ
Azure Document Intelligence โœ… โŒ โŒ
Google Document AI โœ… โœ… โœ…
Mistral OCR (New!) โœ… โœ… โœ…
Docling โœ… โŒ โŒ

What's Changed

Configuration Example

environment: # Mistral OCR (new advanced PDF support) OCR_PROVIDER: "mistral_ocr" MISTRAL_API_KEY: "your_mistral_api_key" MISTRAL_MODEL: "mistral-ocr-latest" # Optional OCR_PROCESS_MODE: "whole_pdf" # Now supported!

Migration Notes

Performance Benefits

PRs

Full Changelog: v0.20.0...v0.21.0

icereed tag:github.com,2008:Repository/861746639/v0.20.0 2025-05-30T14:39:51Z - -v0.20.0 - -

Release Highlights ๐Ÿš€

New Features

๐Ÿง  Google Gemini AI Integration

Improvements & Refinements

๐Ÿ”ง LLM Prompt Optimization

๐Ÿ“„ PDF Processing Fixes

Documentation Updates

๐Ÿ“š Model Recommendations

Dependencies & Maintenance

๐Ÿ”„ Dependency Updates

Technical Details

What's Changed

Contributors

Special thanks to @thiswillbeyourgithub, @dawidkulpa, and @moarsmokes for their contributions to this release!

Configuration Notes


New Contributors

Full Changelog: v0.19.0...v0.20.0

icereed+

Release Highlights ๐Ÿš€

New Features

๐Ÿ”ฎ Mistral OCR Integration with Advanced PDF Processing

๐Ÿท๏ธ Enhanced Title Generation with Context

Improvements & Refinements

๐Ÿ›ก๏ธ Configuration Validation

๐Ÿ“„ Enhanced PDF Processing Architecture

๐Ÿงช Comprehensive E2E Testing

Documentation Updates

๐Ÿ“š OCR Provider Comparison

Technical Details

Provider Mode Support Matrix

Provider image pdf whole_pdf
LLM (OpenAI/Ollama) โœ… โŒ โŒ
Azure Document Intelligence โœ… โŒ โŒ
Google Document AI โœ… โœ… โœ…
Mistral OCR (New!) โœ… โœ… โœ…
Docling โœ… โŒ โŒ

What's Changed

Configuration Example

environment: # Mistral OCR (new advanced PDF support) OCR_PROVIDER: "mistral_ocr" MISTRAL_API_KEY: "your_mistral_api_key" MISTRAL_MODEL: "mistral-ocr-latest" # Optional OCR_PROCESS_MODE: "whole_pdf" # Now supported!

Migration Notes

Performance Benefits

PRs

Full Changelog: v0.20.0...v0.21.0

icereed