Skip to main content

Scanning & Capture

Archive digitization for historical and long-term records

Archive digitization converts historical and long-term records into digital form while preserving the arrangement that gives them meaning. Fragile, brittle and oversized material is captured on glass rather than through a feeder, original order is kept, and output is delivered as a preservation master plus a compressed access copy. Retention review before capture usually reduces the collection first.

Last reviewed

Rows of steel shelving in a records store room packed with archive boxes and lever-arch files

An archive is not a large pile of documents. It is an arrangement, and the arrangement carries meaning. Which series a file sits in, what came before it in the box, which department accessioned it and when: that context is part of the record. Archive digitization services have to preserve the arrangement as carefully as they capture the pages, because a beautifully scanned collection with its original order shuffled is a collection you can read but can no longer rely on.

What an archive usually holds

  • Bound registers and ledgers with entries running straight across the spine.
  • Brittle and foxed paper that will tear if it goes anywhere near a feeder.
  • Carbon copies and thermal paper where the text has already begun to disappear on its own.
  • Oversized plans folded down to fit a standard file, with the creases now split.
  • Microfilm reels and microfiche cards holding material that was itself the preservation copy of something older.
  • Photographs, negatives and tape media filed alongside the paper because that is where they were accessioned.
  • Arabic and English documents in the same file, occasionally on the same page.

Preservation master and access copy

Serious archive work produces two outputs from one capture. The preservation master is made to be the best available surrogate for the original: uncompressed or losslessly compressed, at the resolution and bit depth the source justifies, and processed no further than is needed to represent the page honestly. The access copy is derived from it, compressed, run through OCR and sized for daily use in a document system.

The distinction matters because image processing is a one-way door. Once contrast has been pushed to make faint text readable on a screen, the detail thrown away to get there is gone. If the original is fragile, there may not be a second attempt to recapture it. Keeping an unprocessed master means every future decision about presentation can be revisited. Where your policy calls for a format with a long readability horizon, the access layer is delivered as PDF/A rather than a general-purpose PDF.

Handling material that cannot take a feeder

  • Nothing brittle goes through a roller path. Fragile sheets are captured face-up on a flatbed with the page supported, at the cost of a much slower page rate.
  • Repairs are minimal and reversible: a tear is supported so the sheet can be handled safely, not restored to look undamaged.
  • Fasteners that have rusted into the paper are cut around rather than pulled through, because pulling takes the paper with them.
  • Sheets folded for decades are opened carefully and captured with the fold visible, rather than forced flat and split along the crease.
  • Anything too damaged to capture safely is flagged, photographed in place and referred back to you for a decision instead of being pushed through and hoped for.

Keeping the arrangement intact

  1. 1

    Accession survey

    The collection is walked as it stands: series, boxes, files, condition, and whatever finding aid, register or index already exists. That existing register is often the single most valuable document in the room and should never be treated as scrap.

  2. 2

    Numbering that mirrors the shelf

    The digital hierarchy repeats the physical one, so a reference in a decades-old index still resolves to the right item. Renumbering to something tidier is how you strand every citation that points at the collection.

  3. 3

    Capture in original order

    Batches follow the order on the shelf. Items are never resequenced for scanning convenience, and anything found out of place is recorded as found rather than quietly corrected.

  4. 4

    Description and metadata

    Each level gets the fields your researchers or auditors actually search: date range, originating department, series, reference code and subject. Description is budgeted as work, because it is the retrieval route for everything OCR cannot read.

  5. 5

    Reconciliation and refile

    Items go back in the order they left, and the digital collection is checked against the survey before the project closes.

Retention: settle it before capture

Not everything in an archive should be digitized. Some of it has passed its retention period and can be disposed of under policy. Some is duplicated in a system you already run. Some is permanent and needs the full treatment. Putting the collection against your retention schedule before the project starts routinely removes a meaningful share of it, and that is the cheapest cost reduction available on any archive project. It is also the moment to decide which series need description at item level and which are adequately served at file level.

What the finished archive looks like

For printed twentieth-century material, expect full-text search that works well. For a good deal of older content, OCR is a bonus rather than the retrieval method. Handwriting from the mid-century and earlier, faded dyeline, and Arabic script written by hand will not produce text you can trust in a search index. For those series the index you can rely on is the one a person typed, which is exactly why description work is planned and priced as work rather than assumed to fall out of the scanning.

Frequently asked questions

How is archive digitization different from normal document scanning?

Normal scanning optimises for throughput on uniform, healthy paper. Archive digitization optimises for fidelity and context: fragile material captured on glass, original order preserved, an unprocessed preservation master kept alongside the access copy, and descriptive metadata written at collection, series and file level. It is slower per page and produces a record you can still cite in twenty years.

Can documents too fragile to handle still be digitized?

Usually, on a flatbed with the sheet supported face-up so nothing is pulled through rollers. Brittle paper, split folds and rusted fasteners all have handling routines. Where an item is genuinely too damaged to capture without further loss, it is flagged and referred back to you with the condition recorded, rather than being pushed through on a judgement call.

Do you keep the original file order?

Yes. Capture follows the order on the shelf, the digital hierarchy mirrors the physical arrangement, and items found out of sequence are recorded as found rather than silently rearranged. Original order is evidence in its own right, and a collection that has been tidied during digitization has lost information that cannot be recovered afterwards.

What file formats should an archive be delivered in?

Typically two layers. A preservation master in an uncompressed or losslessly compressed image format, kept unprocessed as the reference copy, and a compressed access copy for daily use. Where a long readability horizon is a policy requirement, the access layer is delivered as PDF/A. The exact formats are agreed against your own preservation policy at the survey.

Can microfilm and microfiche be digitized in the same project?

Yes, and it is usually better to do so. Many collections were partly filmed decades ago, leaving one series split between paper and film. Converting both together lets the reels, fiche and paper files be indexed into a single hierarchy, so researchers search one collection rather than discovering the split for themselves.

Should we digitize the whole archive?

Almost never. Run the collection against your retention schedule first: material past its retention period can often be disposed of under policy, and duplicates of records already held in a system can be excluded. Deciding depth also helps, since some series need item-level description while others are adequately served at file level.

Will old handwritten records be searchable after digitization?

Not reliably through OCR. Historical handwriting, faded dyeline and hand-written Arabic produce a good image but text no search index should be trusted with. For those series, retrieval comes from descriptive metadata captured by a person: date range, department, reference code and subject. That work is planned and priced explicitly rather than assumed.

Ready to talk about Archive Digitization?

Send us your page or box estimate and we will come back with a scoped approach, a security plan and a written quotation.

Or call +971 55 430 1681

CallWhatsAppGet Quote