Legacy document archive migration protects revision history when the project team treats every revision, relationship and approval record as controlled data. Moving the latest file alone does not preserve the record. A reliable migration connects each document to its previous revisions, metadata, status and supporting audit evidence.
In engineering and document-intensive organisations, an old drawing may explain how an asset reached its current configuration. A poor import can separate files from their registers or promote an obsolete copy as the current record.
The safest legacy document archive migration starts with discovery and rules. The team inventories the archive, identifies the authoritative records, maps the metadata and tests a representative sample. The legacy system stays available until the organisation confirms that the new electronic document management system, or EDMS, contains a complete and usable record.
Why does revision history disappear during archive migration?
Revision history rarely exists in one place. The file name may contain a revision number, while a document register holds the status, issue date and approver. An old database may link a drawing to previous versions. Email or workflow logs may contain the reason for a change.
A simple file copy preserves none of those relationships unless the migration team captures and rebuilds them. Common causes of lost history include:
- Importing only the latest approved file
- Treating every revision as an unrelated document
- Replacing source dates with the migration date
- Dropping comments, transmittals or approval records
- Mapping free-text revision fields without checking their meaning
- Deleting duplicates before confirming whether they hold useful context
- Renaming files without preserving the original identifier
Revision labels can conflict. One source might use letters, another numbers and a third a mix of preliminary and contractual codes. The team must understand each convention before translating it.

Start legacy document archive migration with discovery
Create an inventory before changing, renaming or deduplicating anything. The inventory provides the baseline for reconciliation and exception handling.
Capture fields such as:
- Source system, drive, cabinet or storage medium
- Original path, box or folder reference
- File name and document identifier
- File type, size and last-modified date
- Revision, version, status and issue date
- Author, reviewer and approver where available
- Parent document or revision-series relationship
- Access restriction, retention rule or sensitivity marker
- Checksum for digital files
- Condition and scanning requirement for paper records
The National Archives recommends identifying file formats, creating content lists and recording descriptive metadata. It also recommends checksums to test file integrity during transfer. Read The National Archives’ digital preservation workflow guidance.
Did You Know? A checksum creates a digital value from a file. Comparing that value before and after transfer can reveal whether the file changed. The National Archives uses SHA-256 in its preservation workflow, although organisations should select methods that fit their own technical and governance requirements.
Decide which record holds authority
Archives often contain several candidates for the same document. Legacy document archive migration needs an authority rule for each document class.
One archive may treat an approved native file and controlled PDF as one record. Another may treat the signed scan as the formal record. Legal, contractual and operational requirements should guide the decision.
Do not let the migration script make this judgement from a file date alone. A newer timestamp may reflect a server move, file restoration or format conversion rather than a genuine revision.
Record the source, reason, owner and treatment of conflicting copies. Place unresolved cases in an exception queue.

Build a metadata crosswalk before importing files
A metadata crosswalk maps each source field to its destination. It also defines how legacy document archive migration handles blanks, conflicting terms, invalid dates and retired codes.
For revision history, the crosswalk should cover:
- Source document identifier to destination identifier
- Source revision to destination revision
- Revision sequence and parent-child relationship
- Document status and suitability
- Original issue and approval dates
- Author, checker, reviewer and approver
- Workflow outcome and comments
- Superseded or withdrawn indicators
- Links to transmittals, mark-ups and supporting records
Avoid collapsing several source fields into one description box. Structured metadata allows the EDMS to sort, filter and report on revision history. A free-text note may preserve the words without preserving the relationship.
Keep the original value when the new system needs a normalised version. Store the source revision code beside the mapped code so reviewers can investigate discrepancies.
Organisations planning a controlled migration can explore NanoServe’s electronic data solutions and discuss suitable metadata, indexing and workflow requirements. NanoServe should confirm all DataViewer capabilities against the current version, licence, infrastructure and proposed configuration.
How should you handle duplicates and missing revisions?
Run duplicate checks after creating the inventory. Matching content does not always mean matching business context. Two copies may carry different access restrictions or transmittal relationships.
Classify duplicates into clear groups:
- Exact duplicates with no additional context
- Matching content with different metadata
- Renditions of the same source document
- Near-duplicates that require visual or technical comparison
- Conflicting candidates for the authoritative record
Never invent a missing revision to complete a sequence. Record the gap, note the sources checked and assign an owner to investigate it. If the team cannot resolve the issue, retain the exception in the migration report. A visible gap gives future users more reliable information than a false sense of completeness.
Scan paper records to the agreed specification, link each image to the correct identifier and capture stamps, signatures or annotations that affect interpretation. NanoServe’s document-scanning services can support paper and large-format digitisation.

Migrate in controlled batches and validate the results
Test legacy document archive migration with a representative pilot. Include long revision chains, multiple formats, duplicates, missing metadata and restricted content. A pilot filled with easy examples proves little.
For each batch:
- Freeze or control changes in the source records covered by the batch.
- Export the files, metadata and relationship data.
- Record file counts, revision counts and checksums.
- Import the batch into a non-production or controlled staging area.
- Validate identifiers, revisions, dates, permissions and relationships.
- Compare source and destination checksums where appropriate.
- Review a risk-based sample manually.
- Record failures, corrections and approval to proceed.
The reconciliation report should count documents, revisions, renditions, rejected records and exceptions. Correct file totals do not prove correct document relationships.
Protect the old archive until sign-off
Do not remove the source archive when the first import finishes. Keep it read-only and available for verification until the authorised stakeholders approve the migration outcome. The sign-off criteria should cover record counts, revision chains, metadata accuracy, search results, permissions and agreed exceptions.
Personal data also needs a defined retention approach. The UK Information Commissioner’s Office states that organisations should not keep personal data for longer than necessary and should justify retention periods. Migration does not create a reason to retain every record indefinitely. Review the ICO’s storage-limitation guidance.
After sign-off, follow the approved retention and disposal process for the old system. Keep the inventory, mapping rules, exception log, reconciliation results and approval record as evidence.

Plan legacy document archive migration around evidence
Legacy document archive migration succeeds when the new EDMS preserves context, not just files. The destination should show each document’s identity, revision sequence, status, relationships and supporting decisions. Users should understand which version holds authority and why.
Begin with an inventory. Define authority rules and build the metadata crosswalk. Test difficult records in a pilot, migrate controlled batches and reconcile every result before cutover. Keep the source available until formal sign-off.
NanoServe can help organisations assess legacy archives, plan conversion and configure a controlled import into an electronic document environment. The final approach should reflect the archive, required workflows, retention obligations and current system configuration.
FAQ
How do you preserve revision history during an EDMS migration?
Export each revision with its document identifier, sequence, status, dates and relationship to earlier or later versions. Map those fields to the destination schema before importing files. Test complete revision chains during the pilot. Reconcile the number of documents and revisions after each batch, then inspect a risk-based sample manually.
Should you migrate every duplicate document?
Not automatically. Inventory the archive before removing duplicates. Check whether matching files carry different metadata, locations, permissions or transmittal relationships. Preserve useful context even when the binary content matches. Record every deduplication rule and retain an exception queue for records that need human review.
What should happen when a revision is missing?
Do not create a replacement revision or silently close the gap. Record the missing revision, list the sources checked and assign an owner to investigate it. If the team cannot locate it, keep the gap visible in the destination record and include it in the migration exception report.
How do you prove that files did not change during migration?
Create checksums before transfer and compare them with checksums calculated after transfer. Matching values provide evidence that the file content remained unchanged during the move. Keep the results in the migration log. Checksums support integrity checking, but they do not confirm that metadata mappings or revision relationships are correct.
When can the organisation retire the legacy archive?
Retire it only after authorised stakeholders approve the reconciliation results, revision chains, metadata, permissions, search tests and unresolved exceptions. Keep the source read-only during validation. Apply the organisation’s approved retention and disposal process after sign-off, including any legal, contractual and data-protection requirements.
Can DataViewer preserve existing revision history?
The answer depends on the source data, available metadata, DataViewer version, licence and chosen configuration. NanoServe should assess the archive and confirm how the target structure can represent revision relationships before migration. A pilot import should prove the proposed approach before the organisation commits to full-scale conversion.
Plan a controlled archive migration with NanoServe: Contact NanoServe




