Template-free data extraction, from 2012 to generative AI

Template-free data extraction is the reason AIDA exists: our intelligent document processing (IDP) platform was born in 2012, ten years before ChatGPT, to read data from documents by learning from the user in a single click, instead of requiring hand-built templates. Today generative AI boosts that engine, which remains the foundation of the entire product.
Anyone choosing IDP software today is often faced with two kinds of offerings. On one side are the long-established platforms, built on template-based extraction. On the other are the startups of the last three years, which rely entirely on a language model and show convincing demos on the first invoice. AIDA comes from a different path, and understanding it helps you see what happens after the demo, when thousands of real documents come into play.
What was template-based data extraction?
Template-based extraction describes in advance where each piece of data sits on a given document layout: you draw zones or write rules that look for keywords and relative positions, then the software applies that description to incoming documents.
In 2012 the data capture market was led by ABBYY and Kofax (now Tungsten Automation), with powerful platforms built on this principle. In ABBYY FlexiCapture, the structure of variable-layout documents is described with a dedicated program, FlexiLayout Studio; in Kofax Transformation Modules, extraction projects are prepared and tested in Project Builder, which has its own dedicated training courses.
Anyone who worked on a document capture project in those years remembers the limitations well:
- configuration required specialized technicians, dedicated software and training, before the first document was even processed;
- every new supplier with a different layout meant a new template or new rules to write;
- with unstructured documents, such as letters or a law firm's court filings, the template had no fixed reference points to anchor to, because the data could appear in any sentence of the text;
- training and production were separate processes: the project was prepared in a design environment, published to production, and every change went through the same cycle again;
- even a slight change in layout could break extraction until someone updated the template.
How does AIDA's template-free data extraction work?
AIDA learns from users as they work: to teach it a field, you simply select the property and click the value in the document. One click on a single document, and from then on AIDA knows what to look for in subsequent documents of the same type.
In AIDA, the document type does not depend on visual appearance, only on the data you want to extract. A supplier invoice is still a supplier invoice, whether it comes from supplier A or supplier B with a completely different layout.
When you confirm the first document with a new layout, AIDA learns to read it; when another document with the same or a similar layout arrives, it recognizes it on its own. Nobody has to draw a template, name it or specify which one to use for each document, even when there are hundreds of suppliers.
Many solutions recognize a fixed list of standard fields, such as an invoice's supplier, number, date and amounts. AIDA's smart properties have always recognized dates, VAT numbers, IBANs and other common values without training, but the data that matters to a business is often something else: the job number, the cost center, the vehicle license plate, the POD (point of delivery) code on a utility bill. AIDA learns those too, with a click, on any type of document, and with AIDA Boost it finds them from the very first one.
Smart properties and one-click learning do not depend on a language model: they rest on a meticulous analysis of the document's structure, which we have been developing since 2012. AIDA's algorithm reproduces the way a person recognizes a document and identifies its key information. This approach is unique on the market, and it is the solid foundation that AIDA Boost enhances rather than replaces.
The deepest difference from the systems of that era is that training and production are one and the same. The operator opens a pending document, checks the proposed values, clicks the missing or incorrect ones and confirms. Confirmation archives the document or sends it to the chosen destination and, with the same action, consolidates what AIDA has learned. There is no need for a separate design environment or a consultant to republish the project.
In 2012 this was a disruptive approach, and the reason it works has not changed: the people who truly know the documents are those who handle them every day, not those who configure the software. AIDA puts learning in the hands of those people, with no code and no specialist training. It was human in the loop long before generative AI made the phrase commonplace.
Once results are stable, an automation rule can confirm documents on its own, while AIDA highlights unusual values in yellow for review. You can find the details on the page about automated data extraction with AIDA.
Why isn't generative AI alone enough for document management?
A language model reads even a never-before-seen document well, but on its own it does not guarantee repeatable results, often does not remember corrections and knows nothing about archives, duplicates or relationships between documents. These are the limitations that emerge when a platform built solely on generative AI moves from demo to production:
- the same document processed twice can produce different answers, and a model that cannot find a value tends to fill the gap with a plausible one;
- a user's correction often applies only to that document: to avoid repeating it, you have to rewrite the prompt or retrain a model;
- processing costs and times grow with every page sent to the model;
- extracting text is only the beginning, because a business also needs to classify, rename and archive documents, detect duplicates, link purchase order, delivery note and invoice, and retain everything according to precise rules.
That last point is the hardest gap to close. Document management know-how does not come bundled with a language model: it is built over years of working alongside the people who actually handle the documents.
What does AIDA Boost add to data extraction?
Almost everyone promises data extraction today, and one offering looks much like another until you look at what happens to real documents. AIDA Boost is how generative AI enhances AIDA's engine, and the difference shows in what you find ready and waiting when you open a document:
- table rows, such as item codes, quantities and batch numbers on a delivery note, arrive already structured, with no columns to configure;
- all properties, including those specific to your business, are extracted from the very first document: zero-shot extraction, with no training, which adds to what AIDA learns from users;
- AIDA understands what a property is asking for and can combine several related pieces of information from the document into a single value;
- AIDA Boost learns from corrections: when you fix a value with a click, it takes that into account in subsequent documents, combining AIDA's learning with the flexibility of generative AI;
- when documents change, the work of checking and correcting is kept to a minimum.
The result is extraction that goes beyond standard fields, on any type of document, with configuration done from the interface in just a few clicks.
The difference matters most with unstructured documents, such as the correspondence and court filings a law firm receives every day (we cover this in the article on legal innovation with AIDA). AIDA's Hybrid-AI engine reads the text and extracts the opposing party, the case number or the hearing date, even from documents it has never seen before. Unusual values stay highlighted for review, and if a value needs correcting, one click is enough for AIDA to learn it, even on a letter.
AIDA's AI agents start from here: they work on documents that are already classified, read, verified and linked to one another, and they calculate exact totals across the entire archive, so the model never has to guess at missing data.
Templates, generative AI alone and AIDA compared
| Aspect | Template-based extraction | Generative AI alone | AIDA |
|---|---|---|---|
| Initial setup | Templates and rules defined by specialists with dedicated software | Prompts or no configuration | Document types and properties chosen from the interface, with no code |
| New supplier layout | New template or new rules | Read by the model | Learned at the first confirmation and recognized automatically, at most one click on the values |
| Unstructured documents (letters, court filings) | Nearly impossible to describe | Read by the model, with results that may vary | Read from the very first document, with anomalies highlighted |
| Training | Separate from production | Corrections via agents, costly in tokens and often not retained | Happens while documents are being confirmed |
| Results | Stable until the layout changes | May vary between two runs | Consistent, with anomalies highlighted |
| Document management | Often left to other products | Often limited to extraction | End-to-end: personal and agentic tasks, archive, duplicates and document relationships included |
Thirty years of experience, the pace of a startup
AIDA grew out of our team's thirty years of experience in document management, from the ECM systems of the 1990s to digitization with multifunction printers in every office, a story that Giorgio, our CEO, tells in the interview on the evolution of document management. That experience shows in the features a business uses every day, beyond extraction:
- a document archive with unlimited storage on every plan and search across both text and extracted data;
- duplicate detection based on content or extracted fields, with the option to block their validation;
- document relationships, to find the purchase order, delivery note and invoice for the same transaction together;
- dozens of supported file formats, converted to PDF;
- AIDA Pay, which sends the customer a payment request with a link or QR code straight from the approved document and reconciles the payment automatically;
- automations and destinations that deliver data to the systems where it is needed.
Startups born with generative AI begin with a language model and try to build document management around it. AIDA took the opposite route: first it understood how businesses work with documents, then it built an engine that learns from people, and today it enhances that engine with generative AI.
With that experience behind it, AIDA evolves at the pace of a startup: over the last year its changelog has recorded more than 350 updates. These include AIDA 20.0, which introduced agentic AI, as well as AIDA 20.5, AIDA Match for document reconciliation and the new AIDA Mobile 7.0. In 2025 AIDA received the "One to Watch: Product of the Year" award at the Document Manager Awards.
Frequently asked questions
What is template-free data extraction?
It is an extraction method that does not require you to describe in advance where the data sits for each layout. In AIDA, the document type only defines which data to extract, and the engine learns where to find it from the documents users confirm, recognizing each new layout on its own.
Does AIDA require training before getting started?
No. Smart properties recognize dates, VAT numbers, IBANs and other common values without training, and with AIDA Boost all properties are extracted from the very first document. If a value needs correcting, one click during the normal document review is enough, and AIDA learns it.
Does AIDA use generative AI?
Yes. AIDA has a Hybrid-AI engine: document analysis and learning from users, which we have been developing since 2012, work together with the generative AI of AIDA Boost. AI agents use it to search, calculate and act on documents, always on data verified by users.
Do you need a programmer to configure AIDA?
No. Document types, properties, automations and destinations are set up from the interface in just a few clicks. APIs remain available for teams that want to integrate AIDA into their own systems, but they are not required.
To see how AIDA learns from your documents, explore AIDA's data extraction or book a demo and bring your real documents.





