Programmes at this scale rarely fail because scanners are slow. They fail because the volume was estimated rather than counted, because the document mix in box four hundred bore no resemblance to the sample, because the index schema turned out to be unusable during acceptance testing, or because nobody agreed in writing what a defect was until there was an argument about one.
Everything below is about producing numbers you can defend. This article contains none of its own, deliberately: any rate quoted without knowing your documents would be fiction, and the whole method here is designed to replace fiction with measurement.
Count, do not estimate
The first number everyone wants is total pages, and it is the number most often wrong. Converting shelving metres or box counts into pages using a rule of thumb produces an error that then propagates through every downstream calculation: price, duration, staffing, storage.
- Establish the physical unit you can actually count: boxes, files, drawers, linear metres. Count all of it. This part is not sampled.
- Stratify by record series and by era. A 1990s personnel file and a 2020s contract file behave differently in preparation, capture and indexing.
- Physically count the pages in several containers from each stratum, chosen at random rather than pulled from the front of the shelf. One box tells you nothing about variance.
- Record dispersion, not just the average. If containers in one stratum range from a few hundred to a few thousand sheets, that spread is the planning risk and the average conceals it.
- Count what is not paper as well: bound volumes, drawings, photographs, media, and anything in a format nobody has mentioned yet. There is always something.
Then record the mix, because a page is not a unit of work. The attributes that move effort are staples and bindings, sheet size, single or double sided, condition and repairs, whether colour is required, how many index fields the document type needs, and how document boundaries are identified. Two archives with identical page counts can differ by a large multiple in cost, and every bit of that difference lives in the mix.
The pilot exists to produce rates
A pilot that proves scanning is possible has proved nothing anybody doubted. A pilot that produces defensible per-stage rates for each document class is the foundation of the entire plan.
- 1
Sample the real mix, including the ugly strata
Stratified sampling across every document class, deliberately including the damaged, the oversized, the bound and the bilingual. A pilot run on the tidiest boxes produces rates that will be missed for the rest of the programme, and the shortfall will be blamed on the production team.
- 2
Time each stage separately, with the clock genuinely running
Preparation, capture, quality control, indexing, export. Separately, because they have different bottlenecks and different staffing. Include setup, changeover between batches, and interruptions. A rate measured on an uninterrupted hour is not a rate you can plan a shift with.
- 3
Measure rework, not just first-pass output
Count how many images failed QC and why, how long the correction took, and how often a batch had to be re-run. Rework is the difference between a plan that holds and a plan that slips, and it is the number most often left out of a pilot report.
- 4
Time indexing per document, per field set
Indexing scales with fields per document and with how each field is captured: typed, chosen from a list, or extracted and verified. Time a document with the full field set, not an average. If extraction is in scope, measure the straight-through rate and the time to resolve a queued exception.
- 5
Record the exception rate and its causes
Every archive contains documents that cannot be processed as specified. Establish what proportion they are and what they are, because the exception handling process is a resourcing line and an agreement in the contract, not an afterthought.
- 6
Test delivery end to end on a real batch
Push a completed batch all the way into the destination system with the real folder structure, naming convention and load file. Ingestion problems, path length limits, unsupported filename characters and volume constraints are far cheaper to discover on a pilot batch than at month five.
Building the throughput model
Once rates exist, capacity is arithmetic. Build it as a spreadsheet with the pilot rates as visible inputs, so it can be re-run when the mix turns out differently, which it will.
| Stage | What the pilot measures | What that rate sizes |
|---|---|---|
| Preparation | Minutes per container by document class, including repairs and re-sequencing | Prep bench count and prep staffing. Usually the first bottleneck |
| Capture | Images per scanner-hour by class and sheet size, including loading and changeover | Number and type of scanners, and shift pattern |
| Image QC | Images checked per hour, defect rate found, rework minutes per defect | QC headcount and the realistic rework allowance |
| Indexing | Seconds per document at the agreed field set, plus exception resolution time | Indexer headcount. Frequently the true constraint on the whole line |
| Export and delivery | Batch assembly, conversion and transfer time per batch | Batch size, delivery cadence and network or media requirements |
| Exceptions | Proportion of documents that cannot be processed as specified, and why | The exception team, and the contractual process for handling them |
Two adjustments separate a working model from an optimistic one. First, effective hours are well short of paid hours once breaks, handovers, equipment downtime and batch changeover are counted, and the gap should come from the pilot rather than from a guess. Second, new operators are slower and make more errors during their learning period, so a ramp curve belongs in the model. A plan that assumes full productivity from day one on a team that has to be recruited will miss its early milestones and never recover the deficit.
Model the line as a line. Capacity is set by the slowest stage, and adding scanners to a programme constrained by preparation buys nothing but idle scanners.
Quality control that means something
At volume, quality control has to be a defined sampling regime rather than a vigilance policy. Three components, and they layer rather than substitute.
- Operator checks at capture, covering the faults that are obvious at the scanner: blank pages from a double feed, folded corners, clipped edges, skew, wrong colour mode. Cheap, immediate, and no substitute for an independent check.
- Independent sampled audit on defined lots, with a stated lot size, sample size and acceptance number, agreed before production. Defects must be classified as critical, major or minor, with written definitions and examples, because otherwise the classification is argued after the fact.
- Index verification by blind re-keying of a random sample, compared field by field. Document-level accuracy is a misleading measure. One wrong field in ten can make the document unfindable, which is a total loss for that record rather than a ten per cent one.
The rule that matters most is what happens when a lot fails. The whole lot is reworked and re-sampled. Correcting only the defects the sample happened to catch leaves the same underlying error rate in everything that was not sampled, while producing a report that says the problem was fixed.
Phasing, acceptance and change control
Sequence phases by business value and by risk, not by shelf order. The series people request most often should be delivered first, so the archive starts earning its cost early and so real users find schema problems while they are still cheap to fix. Put a technically awkward stratum early too, for the same reason.
Deliver in batches with a defined acceptance step per batch: a stated review window, stated criteria, and a signature. Holding all delivery to the end of the programme means discovering a systematic fault after it has been repeated across the whole archive.
Change control needs to exist before it is needed. Define what counts as a change: a new document class, an added or altered index field, a different output format, a revised naming convention, a retrospective re-index. Every one of those carries an impact assessment, and the assessment must state the cost of applying the change to everything already processed. A mid-programme index-schema change is not a small request. It can mean revisiting every document delivered so far, and both parties are better off knowing that before the change is agreed rather than after.
What actually goes wrong at scale
- The mix is not what the pilot saw. Almost always a sampling failure, and it surfaces as a throughput shortfall that gets misdiagnosed as a productivity problem.
- Document boundary detection breaks down. Separator sheets are missed or misread, and documents merge or split. This corrupts every index field on both affected documents and is the most expensive defect class to remediate.
- The index schema fails on contact with users. Discovered at acceptance testing, when the fields nobody searches by have been captured on four hundred thousand documents and the field everyone needs has not.
- The destination system cannot take the delivery. Volume limits, folder depth limits, path length limits, unsupported characters in file names, or an ingestion rate slower than the production rate.
- Originals are needed back mid-programme and there is no retrieval process. Emergency retrieval from a live production line is disruptive and error-prone, so agree a service level for it at the start.
- A retention or legal hold issue emerges mid-flight, and nobody can tell which boxes are affected because classification was not captured at indexing.
- Acceptance criteria were never written down, so quality is negotiated retrospectively under time pressure.
- Storage and transfer were sized on an assumption of bitonal office paper, then colour and large-format content arrived.
- Attrition. Trained operators leave and are replaced by people on the learning curve, and the plan assumed a steady rate.
- The client-side subject matter expert who understands the filing structure moves on, taking the only explanation of the numbering scheme with them. Document that knowledge early, while somebody still holds it.
A programme planned this way still meets surprises. The difference is that the model shows what the surprise costs, and the phased delivery means it is found in the first hundred thousand pages rather than the last.
Frequently asked questions
How long does it take to scan a million pages?
There is no honest general answer, because duration depends on document mix far more than on page count. Preparation effort, sheet sizes, indexing depth and exception rate can change the total by a large multiple. The way to get a real figure is a stratified pilot that times each stage on your own documents, then a capacity model built from those rates.
What is the bottleneck in a large scanning project?
Rarely the scanners. It is usually document preparation, because removing staples and repairing sheets is manual and does not scale by buying equipment, or indexing, because effort scales with fields per document. Model the operation as a line and size each stage independently. Adding capture capacity to a prep-constrained line produces idle scanners.
How large should a digitization pilot be?
Large enough to cover every document class with a meaningful sample of each, including the damaged, oversized and bilingual strata, rather than a fixed page count. The purpose is to produce per-stage rates and an exception rate per class with enough observations to be credible. A pilot that only sampled the easy material will understate the programme.
What sampling rate should quality control use?
Set it from an acceptance sampling plan with a defined lot size, sample size and acceptance number, agreed before production and tightened or relaxed based on observed performance. The rate itself matters less than two rules: defect classes defined in writing with examples, and failed lots reworked in full rather than patched where the sample happened to look.
What happens if we need to change the index fields mid-project?
Everything already indexed must be revisited unless the new value can be derived from existing data or the OCR text. On a large programme that is a re-indexing exercise priced per document, and it can rival the original indexing cost. Test the schema against real retrieval requests during the pilot, when changing it is still nearly free.
About this article
Written and reviewed by the digitization delivery team at Document Digitization Services, the specialist division of Athena Global Technologies LLC. Content is reviewed against how projects are actually run, and updated when that changes.
Read next
- Choosing a document scanning companyThe evaluation and contract terms that a large programme depends on
- Document indexing and metadata explainedIndex schema design, which is the change most expensive to make later
- Security questions to ask a digitization providerChain of custody and reconciliation, which matter more as volume rises
- How a digitization project runsThe stage-by-stage process this model is built on
- Request a page-count assessmentLarge programmes are quoted after a physical assessment, not before