Skip to main content

Guide · 8 min read

What document digitization actually costs, and what decides it

Document digitization is priced per page, per box or per project, and the rate depends on far more than page count: document condition, preparation effort, how many index fields you need, OCR requirements, on-site or off-site working, output format and turnaround. Two archives of identical size can differ several times over in cost for those reasons alone.

Last reviewed

We are not going to print a rate on this page, and it is worth being straight about why. A rate quoted without knowing your document condition and indexing requirement is either padded to cover the worst case or set low enough to require renegotiation once the first box is opened. Neither helps you budget. What does help is understanding the variables well enough to predict roughly where your project sits, and to recognise a quote that has not been thought through.

By the end of this page you should be able to size your own archive, describe it in the terms a supplier prices on, and compare two proposals that use different units.

The variables that move the price

Cost drivers, and the direction each one pushes. Ranked roughly by how much they typically move a quote.
DriverPushes cost up whenPushes cost down when
Preparation effortHeavy stapling, bulldog clips, treasury tags, bindings, mixed sizes, sticky notes, taped repairsLoose sheets, uniform A4, already unfastened, one consistent format
Indexing depthSix or more captured fields, fields not printed on the page, values requiring judgement or lookupOne or two fields, printed clearly, ideally barcoded or already on a file cover
Document conditionTorn, brittle, damp-damaged, faded thermal paper, carbon copies, tightly folded, oversized plansClean, flat, printed on standard office paper, stored in a controlled environment
VolumeSmall or fragmented batches; each mobilisation carries fixed setup costLarge continuous volumes that let one configured line run without reconfiguration
OCR requirementArabic and English mixed, handwriting, degraded originals, structured field extraction with validationClean machine print in one language, text layer only with no field extraction
Location of workOn-site, in your premises, under supervision, with your access constraintsOff-site at a production facility with full equipment and shift working
Output and integrationPDF/A with embedded metadata, per-document splitting, direct load into a DMS with field mappingFlat PDF per batch delivered on encrypted media or a portal
TurnaroundCompressed deadlines requiring additional shifts or parallel linesA schedule that lets a steady line run at its natural rate
Colour and resolutionColour capture, higher resolutions, large-format equipmentBitonal or greyscale at a standard text resolution
Security requirementsVetted staff only, no-network rooms, escorted transport, per-file chain of custodyStandard confidentiality controls under a normal NDA

Three of those deserve expanding, because they are the ones buyers consistently underestimate.

Preparation is the hidden majority of the labour

A scanner running a clean stack of loose sheets is fast and cheap. Everything before that point is a person at a desk removing fasteners, flattening folds, repairing tears with archival tape and inserting separator sheets. A box of tightly stapled files with taped-in receipts can take many times longer to prepare than a box of loose printed sheets, and no scanner speed compensates for it. When you compare quotes, compare their preparation assumptions first. That is where a low number is usually hiding.

Indexing depth is a decision, not a given

Each index field is a value a person reads and types, or that software extracts and a person verifies. Going from two fields to six does not add a little cost. It roughly triples the keying work on every single document. The discipline is to ask, for each proposed field, what search it enables and who will run that search. Fields nobody searches on are pure cost.

Condition is what wrecks estimates

Faded thermal receipts, brittle carbon copies, oversized drawings folded for years, pages with stamps printed over text. These need slower handling, sometimes flatbed capture rather than a feeder, and they generate exceptions that a person has to resolve. This is precisely why a serious supplier wants to open a representative box before quoting, and why a supplier who quotes from an email is guessing.

Per page, per box, per project: how the models differ

Pricing models compared on the risk each one transfers
ModelHow it worksWho carries the volume riskBest whenWatch for
Per pageA rate per scanned image, usually banded by volume, with preparation and indexing priced separately or as add-onsYou do. The final invoice moves with the real page countYou have a reliable page estimate and mixed document typesWhether the rate covers preparation and indexing or only capture, and whether a double-sided sheet counts as one page or two
Per boxA rate per standard archive box, assuming an average page count per boxThe supplier does, within their assumed averageUniform, well-filled boxes of similar materialThe assumed pages per box, and what happens when your boxes are denser than assumed. Ask for the assumption in writing
Per projectOne fixed price for a defined scope after a physical assessmentThe supplier does, priced with a contingencyA one-off archive with a hard budget ceilingThe change-control clause. A fixed price is only fixed while the scope is
Per hour or resourceCharged by operator day, common for on-site work and unusual materialYou do, entirelySensitive records that cannot leave, or archives too varied to unitiseProductivity assumptions. Ask what output per operator day is expected, and what happens if it is not met
Scan on demandPhysical storage plus a retrieval fee per file scanned when requestedSharedDormant archives with occasional, unpredictable retrievalThe retrieval turnaround commitment and the cost if usage rises

Comparing across models is where buyers get caught. A per-box price and a per-page price cannot be compared until you know the pages-per-box assumption behind the first. Ask for it, then check it against a box you have counted yourself.

Estimate your own scope before you ask anyone

You can do this in an afternoon and it changes the conversation entirely, because you arrive with numbers of your own rather than accepting someone else's.

  1. 1

    Count the containers

    Boxes, lever-arch files, filing-cabinet drawers, shelf metres. Record them by category, not as one total. Different categories will be priced differently.

  2. 2

    Sample a representative box, not a convenient one

    Pick one that reflects the messy reality. Count the sheets in it by hand. That is your real pages-per-box figure. Do not use anyone's default, including a supplier's.

  3. 3

    Record what is in the sample

    How many sheets are double-sided, how many staples and clips there are, how many pages are torn or oversized, how many documents the box contains. The document count matters as much as the page count because indexing is priced per document.

  4. 4

    Multiply, then sample again

    Extrapolate to the full archive, then repeat the count on a second box from a different category. If the two disagree sharply, your archive is not homogeneous and it should be scoped as separate lots.

  5. 5

    Write down your index fields

    List every field you want searchable, and next to each one write who will search by it and how often. Cut anything you cannot answer. This list will move your price more than anything else on it.

  6. 6

    Decide the boring things

    Colour or bitonal, output format, where files land, what happens to the paper afterwards, whether records must stay on your premises, and what your real deadline is as opposed to your preferred one.

What a proper quote contains

  • The unit of pricing, stated explicitly, and whether a double-sided sheet is one image or two.
  • Preparation treated as its own line, with the assumed condition it is based on.
  • The number of index fields included, and the rate for each additional field.
  • Whether OCR is included on all pages, on some, or priced separately.
  • The quality standard, how it is measured, on what sample size, and the remedy when a batch fails.
  • Handling of exceptions: oversized, damaged, illegible and unidentifiable documents, and who decides what happens to them.
  • Transport, insurance and chain of custody, if records leave your premises.
  • The re-fastening and return, storage or destruction arrangement for the originals.
  • The change-control mechanism, in writing, before you need it.

Questions worth asking every supplier

  1. Will you open a box before quoting, and will the quote change if the sample was unrepresentative?
  2. What is your assumed pages per box, and how did you arrive at it?
  3. How many index fields does the quoted rate include?
  4. Who does the preparation work, where, and is it included in the rate?
  5. How do you evidence that every document that arrived was returned and captured?
  6. What is your exception rate assumption, and what do you do when a document cannot be read?
  7. If the deadline moves in, what changes: the price, the scope, or the quality tolerance?

Frequently asked questions

Why will no one publish document scanning prices?

Because page count alone does not describe the work. Condition, fastener density, indexing depth, language, output format and turnaround can move the required labour per page substantially. A published rate is either a worst-case figure that overprices simple archives, or a headline figure that changes after assessment. Ask instead for the rate card structure and the assumptions behind it.

Is per-page or per-box pricing better for a buyer?

Per page is more transparent and moves with reality. Per box transfers volume risk to the supplier but only within their assumed pages per box, so it rewards you if your boxes are dense and penalises the supplier if they are not. Count one box by hand before choosing, then check the assumption in the quote against your count.

Does double-sided scanning cost twice as much?

Not twice, but more than single-sided. Production scanners capture both sides in one pass, so the extra handling is small, while the image count doubles and so does the quality-check and OCR work. What matters commercially is whether the quote prices per sheet or per image. Confirm which, in writing, before comparing suppliers.

How much does indexing add to the cost?

Enough that it is usually the second-largest component after preparation. Each field is a value someone reads and types, or that software extracts and someone verifies, on every document. Adding fields multiplies that work. The most effective way to control a digitization budget is to cut index fields nobody will search by.

Can I reduce cost by preparing documents myself?

Sometimes, and it is worth discussing. Removing staples and clips in-house can lower the quoted preparation effort, but only if done to the supplier's standard, since badly prepared batches cause double feeds and rework. Ask for their preparation specification first, then trial one box before committing your team to the whole archive.

What is the cheapest way to digitize a large archive?

Split it. Index the categories people actually search and audit, and capture the dormant remainder as images with a minimal index recording box, category and date range. This concentrates the expensive work where it earns its keep, and leaves the rest re-indexable later as a targeted job rather than a second full project.

About this article

Written and reviewed by the digitization delivery team at Document Digitization Services, the specialist division of Athena Global Technologies LLC. Content is reviewed against how projects are actually run, and updated when that changes.

Tell us what is in your archive.

Send us your page or box estimate and we will come back with a scoped approach, a security plan and a written quotation.

Or call +971 55 430 1681

CallWhatsAppGet Quote