Why Digital Documents Are Secretly Exciting
Digital documents might seem ordinary at first glance, but beneath the surface, they’re layered with intrigue! As someone who specialises in document analysis, this is my favorite type of investigation …. it’s underrated but absolutely fascinating. These files tell a story …. a journey of where they’ve been, how users have interacted with them, and the changes they’ve undergone. Documents remind me of that classic show This Is Your Life, where the document is the star, and we create our own red book to chronicle its existence. From hidden metadata to revision histories, every layer adds to the narrative. It’s really not just about the content you can see …. it’s about understanding the life of the document itself, and that’s where the excitement lies!
I’ve had the privilege of working with a wide range of documents, each one a new adventure in uncovering hidden artefacts. I absolutely love teaching others about document analysis, because once you see the depth of information these files hold, you’ll never look at a PDF the same way again!

Where Documents Steal The Show!
Digital documents play a key role in a broad range of investigations. Below are examples of different case types where my team and I have analysed digital documents, illustrating their importance across many sectors!
Table showing examples of where Documents can feature in investigations
| Investigation | Type | Sector | Document Type | Notes |
|---|---|---|---|---|
Contract Disputes | Civil | Legal | Legal Agreements, Contracts | Analyse the history of revisions. |
| Corporate Fraud | Criminal | Finance | Spreadsheets, PDFs | Trace financial activity. |
| Cybersecurity Investigations | Criminal | Cybersecurity | PDFs | Analyse phishing attacks or malware delivery |
| Data Breaches | Criminal | Cybersecurity | Leaked Documents | Track the dissemination of sensitive information. |
| Employee Misconduct | Corporate | HR | Memos, Emails | Investigate improper behaviour or policy violations. |
| Financial Audits | Corporate | Finance | Spreadsheets, Financial Statements | Detect irregularities in financial statements and audit trails. |
| Industrial Espionage | Corporate/Criminal | Corporate | Spreadsheets, Reports, Notes | Uncover stolen trade secrets. |
| Insider Trading | Criminal | Finance | Emails, Presentations, Reports, Spreadsheets | Identify suspicious trading patterns. |
| Intellectual Property Theft | Corporate | Corporate | Presentations, Documents, Notes | Reveal unauthorised use of proprietary content. |
| Litigation Support | Civil | Legal | Emails, Reports, Presentations | Provide evidence in civil cases. |
| Regulatory Compliance | Corporate | Various | Reports, Documents | Ensure adherence to industry regulations and standards. |
| Tax Evasion | Criminal | Finance | Financial Documents, Spreadsheets | Investigate illegal practices to avoid taxes. |
| Validating Journalists’ Sources | Corporate/Legal | Media | All Document Types | Verify the authenticity of documents provided by sources. |
| Whistleblower Investigations | Corporate/Legal | Various | Emails, Internal Documents | Investigate and verify claims |
| Workplace Harassment Cases | Corporate/Legal | HR/Legal | Emails, Memos, Notes | Gather evidence for investigation. |

Handle with Care … The Fragility of Digital Evidence
When it comes to digital documents, handling them with care is crucial. Just like physical evidence at a crime scene, digital evidence can be easily altered or destroyed if not properly managed. But there’s another layer to be aware of … your operating system might be trying to “help” you. Even when you do nothing, the OS can trigger changes to document metadata or index the document, potentially leading to data leaking into your lab machine. This is why preserving the integrity of digital documents is key in any investigation.
From the moment a document is identified as potential evidence, it must be treated with the same level of caution as any other piece of evidence. Chain of custody, proper documentation, and using forensically sound tools are essential to ensuring that the evidence remains admissible in court. It’s a delicate balance between uncovering the truth and maintaining the original state of the document, but getting it right is important to solve the case.

Understanding Document Types and Construction
Digital documents come in many forms, and each type has its own unique structure and so many quirks. Popular document formats include PDF, OLE2 (DOC, XLS, or PPT), and OOXML (DOCX, XLSX, or PPTX). Understanding how these files are constructed is crucial for effective forensic analysis. Each file type contains different layers of data, from the visible content to the hidden metadata and embedded objects like images, files, or links.
For example, a PDF might seem like a simple static document, but it can contain embedded fonts, images, and even scripts that execute when the document is opened. An XLSX file isn’t just a spreadsheet …. it’s a collection of XML files zipped together, each storing different aspects of the document, like cell values, formulas, and formatting.
This is what I really love about documents, often we are talking about a specification rather than a strict rigid structure. Often this leads to interesting interpretations of how to put the specification into practice (PDF drivers are a big example of this!) and that leads to facinating and unique forensic traces!!
Knowing the construction of these documents helps forensic analysts determine where to look for hidden information, how to recover data from partially corrupted files, and what artefacts might be left behind by the document’s interactions with different software.

Hidden Layers … Internal Components!
Within every digital document, internal embedded objects …. like images, macros, links. Each object is like mini crime scenes that demands their own separate analysis. Each of these components holds unique clues and history that must be thoroughly examined. But the job doesn’t end there …. once these mini scenes are analysed, they need to be placed back into the wider context of the document. This process is like piecing together evidence from different rooms of a crime scene, ensuring that the history of the document is revealed and contextually understood.

The Art of Timelining
Imagine solving a mystery where the clues are scattered across different times and places, but some pieces are missing or have unexpected twists. In document forensics, timelining and reconstruction help us piece together these clues to form a contextual narrative. After analysing each part of the document we arrange them in chronological order to understand the document’s journey.
When analysing a document, you’re not just looking at the visible content … you’re examining the entire environment it’s been through. The metadata helps show when and where actions were taken. Embedded objects, like images or hidden links, are the clues scattered around, waiting to be discovered. Even deleted or altered content can leave traces, much like fingerprints long forgotten on a surface but still detectable.
This process involves reconstructing the document’s timeline …. who interacted with it, what changes were made, and how it has evolved. It’s about understanding the document’s journey …. reconstructing the events leading up to it’s present state. Every detail, no matter how small, can be crucial in uncovering the truth and solving the mystery hidden within the digital layers.
However, this is where it gets challenging …. we might not be able to identify every single revision or change, which can leave gaps in our timeline. But, the story doesn’t end there. Some document formats have a few surprises up their sleeves. They might store unexpected information, like clipboard data or snippets from other open applications, that weren’t intended to be part of the document. These unexpected details can add vital context, filling in the gaps and giving us a more complete picture of what happened.

Recovering Deleted and Incomplete Documents
One of the most exciting parts of document forensics, for me, is working with deleted or incomplete documents. It’s a bit like being a digital archaeologist, digging through layers of data to uncover hidden or lost artefacts. When a document is deleted, it’s often not entirely gone … traces of it can linger in the system. Similarly, the same can be true for deleted information within a document itself! Recovering these fragments requires skill, patience, and a bit of luck, but when you piece them together, it’s such a thrill!
Incomplete documents present a different challenge … they’re like a jigsaw puzzle with missing pieces. Sometimes, the document was saved in a corrupted state, or perhaps only parts of it were backed up. In these cases, using forensic tools (ok mostly a very VERY lovely hex editor) can help retrieve and reconstruct as much of the original content as possible …. filling in the blanks where data might be missing.
What makes this process even more exciting is the unexpected discoveries. You might uncover draft versions that were thought to be erased, hidden metadata that reveals more than intended, or even remnants of other documents that got mixed in.
This part of document forensics is a blend of science and detective work, where each recovered fragment brings you one step closer in your investigation.

The Final Reveal
After focusing on the individual components, reconstructing timelines, and recovering hidden or deleted data, the final step in document forensics is to bring everything together with contextual analysis. It’s like the grand end reveal in a mystery novel, where all the clues come together to form a complete picture. Getting to our reveal often requires experimentation … testing hypotheses, performing simulated user activity, and exploring different angles to understand the full context of the document.
This stage gathers all the evidence to build a contextual narrative that answers the core questions of the investigation … Who? What? When? How? It’s not just about connecting the dots ….. it’s about understanding the broader context, identifying patterns, and drawing informed conclusions.
The thrill of seeing everything come together is why document forensics is so fascinating to me! It’s the moment where all the pieces fit, and the full story of the document is finally told. This process not only uncovers the truth but also transforms what might seem like a boring file into a crucial piece of evidence that can make or break a case!

Become a Document Detective …. Take Our Short Course on Analysing Digital Documents
If you find the world of document forensics as fascinating as I do, you’re in luck! I’m excited to introduce our new short course, “Document Detective: Analyzing Digital Documents.” This course covers the core aspects of document forensics, including:
- History of Documents
- Document Formats (Binary Specifications and Structures)
- Working with Documents
- Embedded Object Analysis
- Timelining
- Deleted Data Recovery
- Reassembling Corrupt/Incomplete Documents
- Contextualising and Presenting Document Data
These are just the core topics …. there’s much more to explore! The course is taught using immersive, case-driven teaching, making it hands-on with each component featuring an in-depth, low-level practical exercise. All material, data, and walkthroughs are provided on a Virtual Learning Environment (VLE), along with a variety of self-study support exercises.
The course is available online as independent self-driven learning or in a traditional taught module format in person at the University of Southampton. Currently, entry is for practitioners only, but we’ll be opening it to all soon! If you’re interested or want more information on our wide suite of technical digital forensics courses, please get in touch!

For more information on our range of Digital Forensics Short Courses please visit https://sarahmorris.prof/shortcourses.html
