Technologies
The equipment and software behind a digitization project
Projects use production sheet-fed, flatbed, book cradle, wide-format and microform capture, followed by image processing such as deskew, despeckle, adaptive binarisation and colour dropout. Recognition runs as OCR, ICR or full intelligent document processing. Output is delivered as searchable PDF, PDF/A, TIFF or structured data, shaped for the system it is going into.
Last reviewed
Buyers are often shown a technology list before anyone has looked at their documents. It reads well and it decides nothing. The useful conversation runs the other way: what is the material, what condition is it in, what has to be findable afterwards, and what system is it going into. Equipment and software follow from those answers. This page describes the categories we work with and what each one is actually for.
Capture hardware
One scanner class cannot cover a mixed archive. Sending bound volumes through a sheet-fed transport destroys them, and running loose A4 across a book cradle wastes most of a working day. Real projects sort material by physical characteristics during preparation and route each stream to the right device.
| Category | What it is for | Typical material |
|---|---|---|
| Production sheet-fed | High-throughput capture of loose, single-sheet material with automatic feeding, duplex capture and ultrasonic double-feed detection | Correspondence, invoices, application forms, personnel files after preparation |
| Flatbed | Single-sheet capture where the item cannot be fed: fragile, torn, stapled-in-place, thick or unusually sized | Damaged originals, photographs, passports and ID pages, small receipts |
| Book and overhead cradle | Face-up capture with a supported spine, so a bound item is never pressed flat or disbound | Ledgers, minute books, registers, bound archives, library volumes |
| Wide-format | Roll-fed or flatbed capture of oversized sheets that no office device can accept | Drawings, site plans, maps, engineering and architectural sets |
| Microform | Conversion of film-based archives back into digital images, frame by frame | Microfilm reels, microfiche, aperture cards |
Capture settings are decided per record class rather than set once for a whole project. Resolution, bit depth and colour mode are chosen against what has to remain legible and what the recognition stage needs downstream, then validated on a sample batch before volume production begins. Faint carbon copies, coloured forms and pencil annotations all behave differently, and the sample is where that shows up.
Image processing
Raw output from a scanner is rarely the image you want to keep. Processing sits between capture and recognition, and it is where most of the difference between a usable archive and a frustrating one is made. Applied too lightly, text stays broken and OCR guesses. Applied too heavily, thin strokes and light pencil disappear permanently, which is why processing settings are proved on a sample and then locked.
- Deskew, which straightens a page that entered the transport at an angle. Recognition accuracy falls quickly on skewed text
- Despeckle, which removes the scattered noise that photocopying and ageing paper introduce, without eating punctuation
- Adaptive binarisation, which sets the black-and-white threshold region by region rather than once for the whole page. This is what rescues a document with a coffee stain, a shadowed gutter or uneven ageing, where a single global threshold would wash out one half of the sheet
- Colour dropout, which removes a form's pre-printed grid or coloured background so only the entered data survives into the image. It makes field extraction dramatically more reliable on structured forms
- Automatic cropping and border removal, so pages are not delivered inside a black scanner frame
- Blank page detection and page orientation correction, both of which reduce manual handling more than they sound like they should
Recognition: OCR, ICR and IDP
These three terms get used interchangeably by suppliers and they should not be. They describe different things, and confusing them is how expectations get set wrongly at the start of a project.
- 1
OCR: optical character recognition
Converts printed and typed characters in an image into machine-readable text. It is what makes a scanned page full-text searchable. It behaves well on clean printed originals and gets progressively less reliable on faint, skewed, stamped or heavily photocopied material.
- 2
ICR: intelligent character recognition
Extends recognition to hand-printed characters, most usefully in constrained boxes on structured forms. Free-flowing cursive on unstructured paper remains hard for any engine, and any project relying on it should be tested against a real sample before it is priced.
- 3
IDP: intelligent document processing
The layer above recognition. It classifies a document by type, locates the fields that matter, extracts their values, and validates them against business rules or a reference system. The output is structured data, not a page of text, and that is the difference that lets a document post into a business system.
Arabic and English archives sit side by side across the UAE and both are handled, including bilingual index fields. Arabic script is genuinely harder for recognition engines than Latin script, particularly on handwriting and degraded originals, so we run a representative sample first and set expectations from the result rather than from a brochure figure.
Output formats
The format decision looks technical and is really a governance decision. It determines whether the file opens in fifteen years, whether the text is searchable, and whether anything downstream can consume it.
| Format | What it is | When it is the right choice |
|---|---|---|
| Searchable PDF | The page image with a hidden text layer produced by recognition sitting behind it | The default for general business archives. Looks like the original, searches like a document |
| PDF/A | A constrained PDF profile built for long-term preservation, with fonts and colour information embedded and external dependencies removed | Records with long or permanent retention, and anything that has to survive a system migration |
| TIFF | Uncompressed or losslessly compressed page images, single-page or multi-page | Preservation masters, and legacy repositories that will only ingest TIFF |
| JPEG | Lossy compressed continuous-tone images | Photographs and colour material where file size matters more than pixel-exact fidelity |
| Structured data | Extracted field values as XML, JSON or CSV, with or without the source image alongside | Anything that has to post into an ERP, a finance system or a line-of-business application |
Many projects take two outputs: a preservation master and a lighter access copy for daily use. It costs little at capture time and it is expensive to recreate later, because the paper has usually gone by then.
Integration targets
Delivery into a folder is the least useful outcome available. The point of indexing is that a system can consume it, and that system is normally already chosen before we arrive. We build the output to fit what you run.
- SharePoint and Microsoft 365 environments, with index values mapped to site columns and content types rather than dumped as filenames
- ECM and EDMS platforms, using their own import format so that metadata, version history and permissions land correctly
- Cloud object storage, with a manifest and a folder convention agreed in advance so the structure is predictable and scriptable
- ERP, finance and line-of-business systems, which want validated structured data rather than images
- Case, patient and student record systems, where the match key back to an existing record matters more than the image itself
- A plain, well-structured file and folder delivery with an index file, for organisations that have not yet chosen a repository
We do not resell any of these platforms and we do not claim a partnership with their vendors. Integration means we produce output in a shape the target system accepts, and we test that ingestion on a sample before the bulk of the archive is processed.
How the stack gets chosen
In the order that keeps the cost down. A physical survey of the material comes first, then the retention and classification position, then the retrieval requirement, then the target system, and only then the equipment and recognition approach. A sample batch runs end to end through the proposed configuration, you look at the result, and the settings are fixed before volume work starts.
Frequently asked questions
Which scanner brands and software do you use?
We describe technology by category on this site rather than naming products, because naming a manufacturer implies an accreditation or reseller relationship we are not going to assert publicly. For a specific project we will tell you exactly which device classes and recognition tooling would be deployed, and you are welcome to put that in the contract.
What is the difference between OCR and IDP?
OCR converts characters in an image into searchable text, so you get a page you can search. IDP works a level above it: classifying the document type, locating the fields that matter, extracting their values and validating them against rules or a reference system. OCR gives you a searchable document. IDP gives you structured data a system can post.
Should we ask for PDF/A rather than ordinary PDF?
For anything with long or permanent retention, yes. PDF/A embeds fonts and colour information and strips external dependencies, so the file still renders correctly after the software that made it has gone. For short-retention operational records, a standard searchable PDF is normally enough and produces smaller files.
Can bound books and ledgers be scanned without damaging them?
Yes, on an overhead cradle scanner. The volume sits face up in a supported cradle and is captured from above, so the spine is never pressed flat and nothing is disbound. It is slower per page than sheet-fed capture, which is why bound material is usually costed as its own stream rather than blended into the average.
Do you handle microfilm and microfiche?
Yes. Film-based archives are converted frame by frame into digital images and can then be processed and indexed like any other capture stream. Condition matters: film that has been stored warm or damp degrades, and legibility on the resulting images is limited by what survives on the original rather than by the conversion.
Not sure which capture approach your archive needs?
Describe the material: formats, condition, bound or loose, Arabic or English, and where the output has to end up. We will tell you what we would run and why.
Or call +971 55 430 1681