Skip to main content

Data & Intelligence

Document indexing: designing how a record gets found again

Document indexing attaches structured field values to each scanned file, such as account number, document type, date and department, so it can be retrieved by search rather than by browsing folders. Indexing depth is the main cost driver in most digitisation projects, and the main determinant of whether the archive is usable afterwards.

Last reviewed

Close-up of a document on scanner glass with translucent highlight boxes marking individual data fields being detected

Every other page in this section is about reading a document. This one is about finding it three years later, when the person who scanned it has left and the only thing the requester knows is a name and roughly which year. Document indexing services in the UAE are, in practice, the difference between a digital archive and a very tidy pile of PDFs.

Index fields are chosen from the question, not the page

The common mistake is to index whatever the document conveniently displays. The useful method runs the other way. Ask who requests these files, and what they know at the moment they ask. If HR always arrives with an employee number, that is a mandatory index field. If they arrive with a name and a rough joining year, then name and year have to be indexed, and indexing the employee number alone will not help anybody.

Three or four well-chosen fields usually outperform a dozen indifferent ones. Extra fields are not free, and each one adds capture time, verification time and another opportunity for an inconsistent value to enter the archive.

How deep to index

Depth is the single biggest lever on both the cost of a scanning project and the usefulness of what it produces. It is worth deciding deliberately rather than defaulting, and it is entirely reasonable to apply different depths to different parts of the same archive.

Indexing depth, what it costs and what it buys
DepthWhat is capturedSuitsEffect on cost and retrieval
Box and file levelContainer, file title, date rangeDormant archive kept only for retention obligationsLowest capture cost. Retrieval means opening a file and reading through it
Document levelDocument type, date, one primary reference such as an account or employee numberThe common default for operational recordsModerate cost. A specific document can be pulled directly by its reference
Document level plus full textThe above, plus a searchable OCR text layer across every pageContracts, correspondence, anything read rather than filedSmall additional cost where OCR is already running. Finds phrases nobody thought to index
Field levelNamed business values extracted from the page: amounts, parties, expiry dates, statusRecords that feed reporting, alerts or a downstream systemHighest capture and verification cost. Enables filtering, sorting and automation, not just retrieval

Metadata schemas and why consistency beats richness

A schema states, for each field, what it is called, what type of value it holds, what format that value takes, and whether it is mandatory. It sounds bureaucratic until the first time somebody searches for a date recorded four different ways across the same archive.

  • Dates in one format throughout, stored as dates rather than as text, so a range query works.
  • Controlled vocabularies for anything categorical. Department, document type and status should be picked from a list, never typed. Free text guarantees drift.
  • Reference numbers validated at capture against a check digit or your master data, so a transposed digit is caught while the document is still on the desk.
  • Names normalised to one convention. Arabic transliteration in particular produces several plausible spellings of the same name, and the archive needs to settle on one while recording the variants.
  • Mandatory fields kept genuinely minimal. Any field an operator can only complete by guessing will be filled with a guess.

Taxonomy: the shape above the fields

Taxonomy is the agreed set of document types and how they relate. It is where most internal disagreement surfaces, because two departments frequently call the same thing by different names, or the same name by two meanings. Resolving that before capture is far cheaper than reindexing afterwards.

Keep the hierarchy shallow. Deep trees push the person filing into a judgement call at every level, and judgement calls at capture speed produce inconsistency. A short list of document types combined with good index fields outperforms an elaborate structure that nobody applies the same way twice.

This is also the natural point to attach retention metadata. If each document type carries its retention class from the moment it is indexed, disposal schedules can be run later as a query rather than as a manual review of the whole archive.

How index quality is controlled

An index value that is wrong is worse than one that is missing, because a missing value prompts a search and a wrong one ends it. Key fields are keyed independently twice and compared, batches are sampled and checked against the images, and references are tested against your master data where a master list exists. Where automated extraction supplies the index, low-confidence values are verified before release rather than accepted on the assumption that the engine was probably right.

Delivery into your system

Index data is delivered as a load file mapped to your destination fields, whether that is SharePoint columns, EDMS metadata or a database table. Mapping is agreed before capture begins, because renaming fields after a million values have been captured is not a rename, it is a migration. Where you have no destination system yet, the index is delivered in an open, structured form that any future platform can ingest.

Frequently asked questions

What is document indexing?

Indexing attaches structured values to each digital document, such as document type, date, department and a reference number, so records can be retrieved by searching those fields. It is what turns a folder of scanned files into an archive people can actually use without knowing where anything was filed.

How many index fields do we need?

Usually fewer than expected. Three or four fields that match how your teams actually search will outperform a dozen chosen because they appear on the page. Every extra field adds capture cost, verification cost and another chance for inconsistent values to enter the archive.

Does indexing affect the price of a scanning project?

It is normally the largest single variable. Scanning cost per page is fairly stable; indexing cost scales with how many fields are captured and verified per document. Two projects with identical page counts can differ substantially in price purely on indexing depth.

Is full-text search enough on its own?

Rarely. Full-text search finds every document containing a word, which is useful for phrases nobody thought to index and unhelpful for common names or numbers. Index fields return the specific record. Most archives need both, with full text as the safety net behind structured retrieval.

Who decides the taxonomy, you or us?

You own it; we facilitate it. We bring the structures that work in comparable archives and the awkward questions that surface disagreements early, such as two departments using one name for different things. The final vocabulary has to be yours, because your staff are the people who will apply it.

Can you index an archive we have already scanned?

Yes. Existing images can be indexed retrospectively, either by operators reading each document or by running recognition and extraction across them first. The practical limit is image quality: files captured at low resolution or heavily compressed may not support automated extraction and will need manual indexing.

Ready to talk about Indexing & Metadata?

Send us your page or box estimate and we will come back with a scoped approach, a security plan and a written quotation.

Or call +971 55 430 1681

CallWhatsAppGet Quote