Most archives contain two different problems wearing the same cover. There is printed matter, which a recognition engine handles well, and there is everything a person wrote by hand into a box on a form. Intelligent data capture in the UAE is usually bought because of the second category: the values somebody handprinted into a claim form, an application or a site sheet, which currently reach your systems only when a member of staff retypes them.
ICR, OCR and IDP are not the same thing
These three get used interchangeably by vendors, which makes buying harder than it should be. They sit at different levels. One reads characters, one reads handwriting inside known field positions, and one makes decisions about the document as a whole.
| OCR | ICR | IDP | |
|---|---|---|---|
| Reads | Printed and typed characters | Handprinted characters, plus printed text | Whatever the page contains, using OCR and ICR underneath |
| Unit of work | The page | The field | The document, and the batch it arrived in |
| Needs to know the layout? | No, though zoning helps | Yes, either a template or a located field | No. It works out the document type itself |
| Output | A text layer or text file | Named field values with a confidence score each | Validated, structured data plus a routing decision |
| Typical failure | Misread characters on poor originals | Joined-up or overflowing handwriting | Misclassified document, or a rule that silently passes bad data |
| Human step | Sampled quality checking | Verification of every low-confidence field | An exception queue for whatever the rules could not settle |
What handprint recognition can and cannot read
ICR is trained on separated, handprinted characters: block capitals in individual boxes, digits in a combed field, a date written one character per cell. Given that, it performs. Given continuous cursive, it does not, and no amount of tuning changes the fact that joined-up handwriting has no reliable character boundaries to find.
- Reads well: block capitals in printed boxes, digits in combed date and amount fields, tick boxes and shaded response marks.
- Reads with verification: unboxed handprint on ruled lines, mixed upper and lower case, characters that overflow their box.
- Does not read reliably: cursive script, signatures, handwritten Arabic, overwritten corrections, anything crossed out and rewritten.
- Reads better than expected: numeric-only fields, because the character set is small and check digits give the validation step something to test against.
The form itself decides much of this. Comb boxes, drop-out colours that vanish under the scanner lamp, generous field spacing and a clear instruction to use capitals will each do more for capture quality than any software setting. Where a client is redesigning a form anyway, an hour spent on capture-friendly layout pays for itself for years.
Confidence scores are the point, not a footnote
Every field ICR returns carries a score describing how sure the engine is. That number is the mechanism that makes automated capture safe to use, because it separates the values you can accept untouched from the values a person needs to look at.
You set the threshold, and you set it per field rather than per form. An optional comments box and a national identifier do not deserve the same treatment. In practice we set identifiers, amounts, dates and anything with a downstream financial or legal consequence to a strict threshold, and let descriptive fields run looser. A system tuned so that nothing is ever flagged is not accurate, it is untested.
Human-in-the-loop verification
- 1
Capture and score
The engine reads each defined field and attaches a confidence value to the result.
- 2
Route by threshold
Fields above the threshold pass through. Fields below it queue for verification, together with anything a validation rule rejected.
- 3
Verify against the image
The operator sees a cropped image of that field beside the proposed value, not the whole page. Correcting a single field takes seconds and keeps the reviewer looking at exactly the thing in doubt.
- 4
Double-key the critical fields
Nominated fields are keyed independently twice and compared. Where two operators disagree, the record escalates. This is how identifiers and amounts are protected.
- 5
Validate as a record
Cross-field checks run last: does the date make sense, does the total match the lines, does the reference exist in your master data.
- 6
Release and report
Verified data is exported, and the batch report shows how many fields needed intervention and which ones. Those figures are how the next batch gets configured better.
When ICR is worth it, and when it is not
ICR earns its place where the same structured form arrives repeatedly and in volume. Ten thousand identical claim forms is exactly the shape of problem it was built for. A few hundred one-off handwritten letters is not: setting up field templates for material that never repeats costs more than typing it. We will say so rather than sell the licence.
Legacy archives sit somewhere in the middle. Where a historical form was used unchanged for years, a single template can unlock decades of boxes. Where the form was revised every eighteen months, each version needs its own template, and the sensible answer is often to capture only the handful of fields you will actually search on and leave the rest as a searchable image.
Frequently asked questions
What is the difference between OCR and ICR?
OCR reads printed and typed characters across a whole page. ICR is built for handprinted characters inside known field positions on a form, and returns named values rather than a block of text. Most real projects use both: OCR for the printed body, ICR for the completed fields.
Can ICR read joined-up handwriting?
Not reliably. ICR depends on finding boundaries between individual characters, and cursive script does not provide them. Block capitals in boxes read well, unboxed handprint reads with verification, and continuous cursive is normally captured by an operator instead. We assess a sample before committing either way.
What is a confidence score used for?
It tells you how certain the engine is about a specific field, which is what decides whether that value is accepted automatically or sent to a person. Thresholds are set per field, so an identifier or an amount is treated far more strictly than a free-text comments box.
How do you stop wrong data reaching our system?
Three layers. Confidence thresholds route doubtful fields to a verification screen showing the field image beside the read value. Critical fields are keyed twice independently and compared. Validation rules then test the record as a whole, checking dates, check digits and references against your master data.
Does the design of our form affect capture quality?
Considerably. Comb boxes, drop-out colours that disappear under the scanner lamp, wide field spacing and an instruction to use capitals all raise the proportion of fields read cleanly. If a form is being revised anyway, capture-friendly layout is the cheapest quality improvement available.
Can captured data go straight into our existing system?
Yes. Output can be delivered as CSV, XML or JSON for import, or written directly into an EDMS, ERP or line-of-business application through its API or a scheduled load. The field names in the export are mapped to your system fields during setup, not afterwards.
Do you handle handwritten Arabic forms?
We handle them, but with human verification rather than an automated pass. Handwritten Arabic combines cursive joining with positional letter forms, and on aged or carbon-copy originals it is the hardest material in document capture. We price those batches as operator-verified work from the outset.
Related reading
- OCR and text conversionThe printed-text layer that ICR sits alongside on most forms.
- Intelligent document processingWhere captured fields get classified, validated and routed automatically.
- How a digitisation project runsPreparation, capture and quality control before any field is read.
- Request a sample batch assessmentThe only honest way to price handwritten material.
