AI Document Management
A client in the healthcare sector wanted to make large volumes of PDF documents accessible in a way that would allow relevant content to be found quickly and reliably incorporated into daily work.
To that end, we have developed an AI-based document management system—including a front end—that organizes complex content in a structured manner and operates entirely on-premises.
Key Areas of Cooperation
The main challenge was making a constantly growing collection of documents practically usable. This requires processing thousands of existing PDFs as well as new documents that are added on an ongoing basis and indexing their content.
A large portion of the files are scans of highly variable quality, which requires robust text recognition. At the same time, strict data protection requirements applied: The solution had to operate entirely on the local infrastructure without internet access. Limited system resources imposed tight technical constraints. The goal was a solution that, despite these limitations, would enable high-quality summaries, targeted information extraction, and a clear presentation of the results.
Custom applications offer greater flexibility, better scalability, and more efficient workflows—especially when unique business processes need to be digitized.
Our Project KPIs
High Scores Based on LLM-as-a-Judge Compared to Sample Summaries by Domain Experts
AI-Based Document Processing
![]()
Key Factors
- The Usability of Large Volumes of Documents in Everyday Work
- High quality despite varying document quality
- On-premises operation with limited resources
- Large Language Models
- Vision-Language Models
- Containerization
![]()
Key Performance Indicators
Text Recognition from PDFs (OCR):
- Classic OCR tool (Tesseract) with a character error rate (CER) of 0.1625
- Our optimized version using a Vision-Lange model (VLM) has a CER of 0.0884 (54 fewer % errors).
Summaries:
- High Scores Based on LLM-as-a-Judge Compared to Sample Summaries by Domain Experts
Extraction:
- 94.7 % extracts the „target values“ from the documents fully automatically.
Results and Outlook
AI-based document processing makes large volumes of PDFs accessible and, for the first time, enables their content to be systematically utilized. Relevant information can be found quickly, is consistently linked across multiple documents, and can be traced back to its source at any time. This significantly reduces the workload for professionals and enables them to reliably integrate even growing document collections into their daily work.
The modular architecture enables step-by-step expansion to include new extraction logic and continuous improvement. In the future, additional business processes can be integrated, and additional knowledge relationships can be automatically identified.
Our Services and Case Studies in Artificial Intelligence
- AI Consulting and Workshops,
- Implementations,
- AI-powered data classification,
- AI Document Management,
- AI-based process automation,
- AI agents,
- Machine Learning,
- LLMs
Artificial Intelligence Consulting
Sustainable know-how transfer for your company thanks to our AI consulting: workshops, use case evaluation, implementation consulting …
AI-powered Data Classification
Automated Classification of Structured Datasets in the Public Sector …
AI-Based Process Automation
AI-based process automation for corporate data sets and data mapping in the public sector …
We'd be happy to advise you
Dr. Danny Hucke
We understand that every project is unique. Dr. Danny Hucke looks forward to speaking with you to find customized solutions.