Skip to main content

Data & Intelligence

Intelligent data capture: reading handwriting and form fields

Intelligent character recognition reads handprinted characters and pulls values out of defined form fields, where standard OCR reads printed text only. Every value carries a confidence score, and anything below the agreed threshold is routed to an operator for verification rather than passed through as fact.

Last reviewed

Close-up of a document on scanner glass with translucent highlight boxes marking individual data fields being detected

Most archives contain two different problems wearing the same cover. There is printed matter, which a recognition engine handles well, and there is everything a person wrote by hand into a box on a form. Intelligent data capture in the UAE is usually bought because of the second category: the values somebody handprinted into a claim form, an application or a site sheet, which currently reach your systems only when a member of staff retypes them.

ICR, OCR and IDP are not the same thing

These three get used interchangeably by vendors, which makes buying harder than it should be. They sit at different levels. One reads characters, one reads handwriting inside known field positions, and one makes decisions about the document as a whole.

OCR, ICR and IDP compared
OCRICRIDP
ReadsPrinted and typed charactersHandprinted characters, plus printed textWhatever the page contains, using OCR and ICR underneath
Unit of workThe pageThe fieldThe document, and the batch it arrived in
Needs to know the layout?No, though zoning helpsYes, either a template or a located fieldNo. It works out the document type itself
OutputA text layer or text fileNamed field values with a confidence score eachValidated, structured data plus a routing decision
Typical failureMisread characters on poor originalsJoined-up or overflowing handwritingMisclassified document, or a rule that silently passes bad data
Human stepSampled quality checkingVerification of every low-confidence fieldAn exception queue for whatever the rules could not settle

What handprint recognition can and cannot read

ICR is trained on separated, handprinted characters: block capitals in individual boxes, digits in a combed field, a date written one character per cell. Given that, it performs. Given continuous cursive, it does not, and no amount of tuning changes the fact that joined-up handwriting has no reliable character boundaries to find.

  • Reads well: block capitals in printed boxes, digits in combed date and amount fields, tick boxes and shaded response marks.
  • Reads with verification: unboxed handprint on ruled lines, mixed upper and lower case, characters that overflow their box.
  • Does not read reliably: cursive script, signatures, handwritten Arabic, overwritten corrections, anything crossed out and rewritten.
  • Reads better than expected: numeric-only fields, because the character set is small and check digits give the validation step something to test against.

The form itself decides much of this. Comb boxes, drop-out colours that vanish under the scanner lamp, generous field spacing and a clear instruction to use capitals will each do more for capture quality than any software setting. Where a client is redesigning a form anyway, an hour spent on capture-friendly layout pays for itself for years.

Confidence scores are the point, not a footnote

Every field ICR returns carries a score describing how sure the engine is. That number is the mechanism that makes automated capture safe to use, because it separates the values you can accept untouched from the values a person needs to look at.

You set the threshold, and you set it per field rather than per form. An optional comments box and a national identifier do not deserve the same treatment. In practice we set identifiers, amounts, dates and anything with a downstream financial or legal consequence to a strict threshold, and let descriptive fields run looser. A system tuned so that nothing is ever flagged is not accurate, it is untested.

Human-in-the-loop verification

  1. 1

    Capture and score

    The engine reads each defined field and attaches a confidence value to the result.

  2. 2

    Route by threshold

    Fields above the threshold pass through. Fields below it queue for verification, together with anything a validation rule rejected.

  3. 3

    Verify against the image

    The operator sees a cropped image of that field beside the proposed value, not the whole page. Correcting a single field takes seconds and keeps the reviewer looking at exactly the thing in doubt.

  4. 4

    Double-key the critical fields

    Nominated fields are keyed independently twice and compared. Where two operators disagree, the record escalates. This is how identifiers and amounts are protected.

  5. 5

    Validate as a record

    Cross-field checks run last: does the date make sense, does the total match the lines, does the reference exist in your master data.

  6. 6

    Release and report

    Verified data is exported, and the batch report shows how many fields needed intervention and which ones. Those figures are how the next batch gets configured better.

When ICR is worth it, and when it is not

ICR earns its place where the same structured form arrives repeatedly and in volume. Ten thousand identical claim forms is exactly the shape of problem it was built for. A few hundred one-off handwritten letters is not: setting up field templates for material that never repeats costs more than typing it. We will say so rather than sell the licence.

Legacy archives sit somewhere in the middle. Where a historical form was used unchanged for years, a single template can unlock decades of boxes. Where the form was revised every eighteen months, each version needs its own template, and the sensible answer is often to capture only the handful of fields you will actually search on and leave the rest as a searchable image.

Frequently asked questions

What is the difference between OCR and ICR?

OCR reads printed and typed characters across a whole page. ICR is built for handprinted characters inside known field positions on a form, and returns named values rather than a block of text. Most real projects use both: OCR for the printed body, ICR for the completed fields.

Can ICR read joined-up handwriting?

Not reliably. ICR depends on finding boundaries between individual characters, and cursive script does not provide them. Block capitals in boxes read well, unboxed handprint reads with verification, and continuous cursive is normally captured by an operator instead. We assess a sample before committing either way.

What is a confidence score used for?

It tells you how certain the engine is about a specific field, which is what decides whether that value is accepted automatically or sent to a person. Thresholds are set per field, so an identifier or an amount is treated far more strictly than a free-text comments box.

How do you stop wrong data reaching our system?

Three layers. Confidence thresholds route doubtful fields to a verification screen showing the field image beside the read value. Critical fields are keyed twice independently and compared. Validation rules then test the record as a whole, checking dates, check digits and references against your master data.

Does the design of our form affect capture quality?

Considerably. Comb boxes, drop-out colours that disappear under the scanner lamp, wide field spacing and an instruction to use capitals all raise the proportion of fields read cleanly. If a form is being revised anyway, capture-friendly layout is the cheapest quality improvement available.

Can captured data go straight into our existing system?

Yes. Output can be delivered as CSV, XML or JSON for import, or written directly into an EDMS, ERP or line-of-business application through its API or a scheduled load. The field names in the export are mapped to your system fields during setup, not afterwards.

Do you handle handwritten Arabic forms?

We handle them, but with human verification rather than an automated pass. Handwritten Arabic combines cursive joining with positional letter forms, and on aged or carbon-copy originals it is the hardest material in document capture. We price those batches as operator-verified work from the outset.

Ready to talk about ICR & Data Capture?

Send us your page or box estimate and we will come back with a scoped approach, a security plan and a written quotation.

Or call +971 55 430 1681

CallWhatsAppGet Quote