Reference
Document digitization glossary
Definitions for the terms that come up when scoping a digitization project. Several of these are used interchangeably by suppliers when they should not be, so where two terms are commonly confused the entry says so explicitly and draws the distinction.
- Batch separation
Batch separation is the method that tells a scanner where one document ends and the next begins inside a continuous stack of paper, usually by inserting barcoded separator sheets, patch codes or blank pages between files.
Without it, a tray of 500 loose pages arrives as a single 500-page image file rather than the sixty separate records it actually contains, and the only remedy is splitting them by hand afterwards. Barcoded separators are the common choice because the barcode can also carry the index values for the document that follows, which removes a keying step. It is decided during preparation, and it is one of the quiet choices that determines whether indexing takes hours or days.
See alsoDocument scanningIndexingMetadataDocument classification
- Chain of custody
Chain of custody is the documented, unbroken record of who held a set of physical records at every point between collection and return, covering each transfer, storage location and any authorised destruction.
For regulated files such as patient records, personnel files or litigation material, the custody trail matters as much as image quality, because a gap in it can undermine the evidential standing of the whole batch. In practice it means box-level barcodes, a signature at every handover, and a log the client can inspect at any point. Ask to see the custody log format before the first box leaves your building, not after.
- Data extraction
Data extraction is the capture of specific field values from a document (invoice number, date, supplier, total) as structured data that can be validated and posted into a business system, rather than as undifferentiated page text.
OCR gives you the words on the page; extraction decides which of those words is the invoice total. Stable layouts are handled with templates, variable ones with trained models, and both need a validation rule set to catch the failures that would otherwise pass through quietly. The figure worth asking a vendor for is not character accuracy but how many documents clear validation with no human touching them.
See alsoOCR (optical character recognition)Intelligent document processing (IDP)Document classificationMetadata
- Digital archive
A digital archive is a managed long-term repository of digitized and born-digital records, held in preservation-grade formats with the metadata, access controls and integrity checks needed to keep them readable and trustworthy for decades.
A folder of PDFs on a shared drive is storage, not an archive, because nothing in it assures you the files will still open or still be unaltered in fifteen years. The difference shows up in format choice (PDF/A or TIFF rather than whatever a scanner wrote), in fixed descriptive metadata, and in periodic integrity verification. Where retention is measured in decades, archive requirements should shape the capture specification from the beginning; retrofitting them across millions of images is a second project.
See alsoPDF/ARecords retentionMetadataEnterprise content management (ECM)
- Document classification
Document classification is the identification of what type a document is (contract, invoice, passport copy, delivery note) so that it can be routed, indexed and retained under the rules that apply to that type.
It sits ahead of extraction in the processing chain, because you cannot pull the right fields until you know what you are looking at. Classification can be rule-based, keying off a barcode, a form number or a recognisable header, or model-based where a backfile arrives mixed and unlabelled. Getting it wrong is costly in a specific way: a misclassified record inherits the wrong retention period, so it is either destroyed early or kept long past the point policy allows.
See alsoIntelligent document processing (IDP)TaxonomyRecords retentionData extraction
- Document digitization
Document digitization is the full conversion of physical records into usable digital information, covering preparation, image capture, text recognition, indexing, quality control and delivery into the system where the records will actually be used.
Scanning is one step inside digitization, and treating the two as the same thing is the most common way a project gets under-scoped. A per-page price for images alone covers a fraction of the work, since indexing depth and the target system usually drive more of the cost than scanner time does. Decide early what you need to be able to search on, because that single decision moves the price more than page count.
See alsoDocument scanningIndexingOCR (optical character recognition)Document management system (DMS)
- Document management system (DMS)
A document management system is software for storing, versioning, securing and retrieving documents, giving every file a controlled location, an access rule and a history of who changed it and when.
The line between a DMS and an ECM platform is scope: a DMS manages documents, while ECM manages all enterprise content and the governance and processes wrapped around it. Many organisations that say they need ECM actually need a DMS with a retention module, and buying the larger platform to solve a filing problem is a familiar and expensive detour. Work out first whether the real requirement is retrieval, governance, or both, because that answer sizes the purchase.
See alsoEnterprise content management (ECM)IndexingRecords retentionDigital archive
- Document scanning
Document scanning is the imaging step that turns a physical page into a digital image file, using production scanners, flatbeds, book cradles or large-format equipment depending on what the material will tolerate.
On its own, scanning produces pictures of paper: files nobody can search, named by scanner sequence rather than by content. It becomes useful when paired with OCR and indexing, which is why a quote for scanning alone rarely reflects the real project. Fragile, bound, oversized or torn material dictates the equipment, and on damaged files the preparation time regularly exceeds the time spent scanning.
See alsoDocument digitizationDPI (dots per inch)Image cleanupBatch separation
- DPI (dots per inch)
DPI, or dots per inch, is the resolution at which a page is captured, and it governs how much fine detail survives into the digital image: small type, signature strokes, stamp edges, faint carbon copy.
300 DPI is the usual working baseline for ordinary office text and is what most OCR engines are tuned around. Going higher helps with small print, poor originals and detailed drawings, and does nothing for a clean typed letter except multiply the file size. Scanning too low is the more damaging mistake, because detail that was never captured cannot be recovered without pulling the paper back out of storage, so set the resolution against the worst documents in the set rather than the average ones.
See alsoDocument scanningOCR (optical character recognition)Image cleanupDigital archive
- Enterprise content management (ECM)
Enterprise content management is the combined strategy, platform and governance model an organisation uses to manage all of its unstructured information (documents, images, email, media and records) across the whole lifecycle from creation to disposal.
ECM reaches well past storage: it covers the retention schedules, workflow, classification and audit trails that decide what happens to content, not merely where it sits. Buyers usually meet ECM as a platform name, but the platform is the smaller half of the problem, and the taxonomy and retention rules are what make it work or quietly fail. A migration into ECM is a good moment to fix an inherited filing structure and a poor moment to carry it across unchanged.
See alsoDocument management system (DMS)TaxonomyRecords retentionDigital archive
- Full-text search
Full-text search is retrieval that queries the entire recognised body text of a document rather than only its filename or index fields, so any word on any page becomes a way into the collection.
It rests entirely on OCR quality, which is why it performs well on clean printed material and unevenly on handwriting, faded fax copy or dense Arabic script. Because it returns everything containing a term, it complements indexed search instead of replacing it: index fields narrow the set, full text finds the outlier nobody thought to index. A collection with reliable metadata and mediocre OCR is generally more usable than the reverse.
See alsoOCR (optical character recognition)Searchable PDFIndexingMetadata
- ICR (intelligent character recognition)
ICR is the recognition of hand-printed and handwritten characters, using trained models that interpret letterforms varying from one writer to the next instead of matching against the fixed shapes a machine produces.
OCR assumes a machine wrote the text; ICR assumes a person did, and that single assumption changes how the engine works and how far you can trust the result. Hand-printed characters in boxed form fields read reasonably well, while free cursive scrawled in a margin often does not. This is why ICR output normally passes through validation or a double-key check on any field that will drive a payment, a claim decision or a licence.
See alsoOCR (optical character recognition)Data extractionIntelligent document processing (IDP)Document classification
- Image cleanup
Image cleanup is the set of automatic corrections applied to a scanned page before it is stored or read by OCR: deskewing a crooked capture, despeckling scanner noise, removing black borders and punch-hole marks, and cropping to the page edge.
It is the cheapest accuracy gain available in a digitization project, because recognition engines lose a surprising amount to a page sitting three degrees off square. Aggressive cleanup carries its own risk, since over-filtering can erase a faint stamp or a thin signature stroke that matters legally. Keep the settings documented and conservative on anything with evidential value, and retain the raw capture where the record is contentious.
See alsoDocument scanningOCR (optical character recognition)DPI (dots per inch)Microfiche
- Indexing
Indexing is the assignment of agreed search values to each digitized document, such as an account number, a date, a name or a contract reference, so files can be retrieved by what they are about rather than by browsing folders.
Indexing and metadata get used interchangeably, but they are not the same thing: metadata is the data attached to the document, and indexing is the work and the rules that produce it. Cost scales directly with depth, so the useful question is how few fields you can capture and still retrieve reliably. Three well-chosen fields backed by full-text search usually beats twelve fields captured inconsistently by tired operators.
See alsoMetadataFull-text searchTaxonomyDocument digitization
- Intelligent document processing (IDP)
Intelligent document processing is the automated pipeline that carries a document from capture through classification, data extraction, validation and release into a business system, combining OCR, ICR and machine learning with business rules.
The point of IDP is not recognition; it is the decision that follows recognition, which is why exception handling and validation rules matter more than which recognition engine sits underneath. Vendors quote accuracy per character or per field, but the operational number is the straight-through rate, meaning the share of documents that finish without a human. It earns its cost on steady repeating volumes, and a one-off backfile of mixed unstructured paper is a harder and less rewarding target than a daily stream of the same form.
See alsoOCR (optical character recognition)ICR (intelligent character recognition)Data extractionDocument classification
- Legal hold
A legal hold is an instruction that suspends the normal destruction of specified records because they are relevant to actual or anticipated litigation, investigation or audit, and it overrides the retention schedule until it is formally lifted.
Automated retention turns into a liability the moment a hold exists that the system cannot enforce, because routinely disposing of held material is far more damaging than keeping too much. Any records platform under evaluation should be able to freeze a defined set, log who applied the hold and why, and block deletion by administrators as well as ordinary users. Holds also need a release process, or the exception silently becomes permanent and the retention policy stops meaning anything.
See alsoRecords retentionChain of custodyEnterprise content management (ECM)Redaction
- Metadata
Metadata is the structured descriptive information attached to a digital record (title, date, author, document type, department, retention class) that makes it findable, governable and meaningful outside its own contents.
Some of it is captured automatically at scan time, including resolution, capture date, operator and batch reference, and some has to be decided by people who understand the records. The governance metadata is the part organisations skip and later regret, because a document carrying no retention class cannot be disposed of correctly when its time comes. Settle the field list before capture starts; adding one field to two million already-scanned images is a project in itself.
- Microfiche
Microfiche is a flat sheet of film carrying a grid of miniaturised document images, used widely from the mid-twentieth century for archives, land records, engineering drawings and newspapers before digital storage became practical.
Fiche and roll microfilm need dedicated film scanners rather than document scanners, and the recoverable quality is capped by the original filming, not by the scanner you point at it. Many archives across the UAE and the wider GCC still hold fiche created decades ago, and film degrades with heat and humidity, so the material tends to get worse while it waits for a decision. OCR results from fiche are usually weaker than from paper, which makes indexing at the point of conversion more important rather than less.
See alsoDocument scanningDigital archiveImage cleanupIndexing
- OCR (optical character recognition)
OCR is the conversion of printed or typed text inside a page image into machine-readable characters, so the words on the page become text a computer can search, copy and process.
An engine is only ever as good as what it is handed: clean capture of laser-printed text reads close to perfectly, while a fifth-generation fax copy covered in stamps and handwritten annotations does not. Accuracy is also language-dependent, and Arabic, being cursive, context-shaped and frequently mixed with English on the same page, is harder than Latin script and deserves a sample test before anyone commits. Treat quoted accuracy figures as conditional until the vendor has run your worst documents rather than a tidy sample set.
See alsoICR (intelligent character recognition)Searchable PDFFull-text searchDPI (dots per inch)
- PDF/A
PDF/A is the archival subset of the PDF standard, which requires everything needed to display a file correctly (fonts, colour profiles, image data) to be embedded inside the file itself and prohibits features that would break future rendering, such as external links, scripting and encryption.
The aim is a file that opens and looks the same decades from now, on software nobody has written yet. It is commonly specified for records under long statutory retention, and it costs nothing extra to produce when it is set at the outset instead of run as a later conversion pass. Note what it does and does not promise: PDF/A guarantees the rendering, not the content, so a searchable PDF/A is still only as good as the OCR that built its text layer.
See alsoSearchable PDFDigital archiveRecords retentionDocument digitization
- Records retention
Records retention is the policy that sets how long each class of record must be kept before it is destroyed or transferred to permanent archive, driven by statutory, regulatory, tax and contractual obligation rather than by how much storage happens to be available.
Digitizing does not reset the clock, since a scanned record inherits the retention period of the original, and whether the paper itself can then be destroyed depends on the record type and the rules that apply to it. Retention only works when it is attached to the record as metadata at the point of capture, because applying a schedule afterwards means reviewing a repository document by document. The end of a period should trigger a logged, defensible disposal step rather than a quiet deletion nobody can account for.
See alsoMetadataLegal holdDigital archiveEnterprise content management (ECM)
- Redaction
Redaction is the permanent removal of specific information from a document, such as identity numbers, bank details, personal data or commercially sensitive terms, so that a copy can be released without exposing what was taken out.
A black rectangle drawn over a PDF is not redaction. The text sits underneath it and can be selected, copied or recovered, and that is precisely how organisations leak data while believing they have protected it. Genuine redaction deletes the underlying content along with the corresponding portion of the text layer. In a digitization project it is normally applied to a derived release copy, with the unredacted original retained under tighter access control.
- Searchable PDF
A searchable PDF is a scanned document carrying an invisible OCR text layer aligned behind the page image, so the file still looks exactly like the original paper while its words can be searched, selected and copied.
It is the default delivery format for most digitization work because it satisfies two audiences at once: people who need to see the document as it was signed and stamped, and systems that need the text. What it does not give you is structured data, since a searchable PDF of an invoice knows the characters but not which of them is the total. Where the downstream requirement is posting into a finance or claims system, ask for extraction output alongside the PDF rather than instead of it.
See alsoOCR (optical character recognition)Full-text searchPDF/AData extraction
- Taxonomy
A taxonomy is the agreed classification structure for an organisation's records: the document types, the categories they sit within, and the controlled vocabulary used to label them, which keeps naming and filing consistent across departments.
It is the piece most often deferred to the end of a project and the piece that decides whether anything can be found once the project is over. A useful taxonomy is shallower than people expect and is built from how records are actually searched for, not from an org chart that will be redrawn within two years. Where a legacy filing structure is being migrated, settle the taxonomy before the migration runs, because moving several million documents twice is the expensive route to the same place.
See alsoDocument classificationIndexingMetadataEnterprise content management (ECM)
Tell us what is in your archive.
Send us your page or box estimate and we will come back with a scoped approach, a security plan and a written quotation.
Or call +971 55 430 1681