What happened
On 17 August 2026, 404 Media published an investigation into the purchase and scanning of books for AI. The outlet said a tracking device placed in a shipment led it to an Amazon facility in Las Vegas.
Its report described employees cutting book bindings to speed up scanning, destroying the printed copies in the process. These are the publication’s reported findings. They do not, by themselves, settle the rights associated with every title or establish the terms under which any resulting material was used.
The story turns an abstract phrase, “training data”, into something tangible.
Why it matters
When a business buys access to AI, it is easy to treat data as a technical ingredient somewhere upstream. When the same business starts building its own knowledge system, that ingredient becomes an immediate management responsibility.
A folder of documents can contain several different things: material the company created, content licensed for a particular purpose, personal information and records supplied in confidence. Having access to the folder does not answer every question about using it in another system.
I would begin with a simple inventory. For each collection, identify its owner, origin, permitted uses, sensitivity and retention requirements. Where the position is unclear, ask the relevant rights holder or specialist rather than letting the uploading process make the decision.
The business value of a data collection includes its reliability and legitimate usability. A large collection with unclear provenance can create expensive uncertainty precisely when a project is ready to expand.
The bigger shift
AI projects can bring overlooked company knowledge back into view. Historical manuals, product records and resolved support cases may help staff answer questions. But digitising or indexing material should preserve enough context to interpret it correctly.
An old service instruction might apply only to a discontinued product. A successful proposal might contain terms that were exceptional. A customer complaint might record an allegation, followed elsewhere by a correction. Removing that context can make a confident answer less trustworthy.
I would assign a domain owner to approve the material used in an initial knowledge pilot. Start with a small, well-understood collection. Retain references to the original records, distinguish superseded versions and test whether users can check the evidence behind an answer.
These are practical parts of AI governance, alongside the decision about which model to use.
My take
The most useful lesson here is to make the origin of information visible before celebrating what a system can generate from it.
For a company knowledge project, run a provenance review on a sample of documents before importing the rest. Ask whether someone can explain why each document belongs, whether it is current and who can resolve an objection. Include the process for removing material and checking what depends on it.
Also keep the original business records. A convenient AI interface should not become the only route to institutional knowledge or replace the evidence needed to resolve a dispute.
Data quality is more than clean formatting. It includes context, permissions and responsibility. Those qualities are harder to demonstrate in a product launch, but they are easier to appreciate when an important answer turns out to depend on the wrong document.
Sources
Read our editorial policy for our approach to sourcing, analysis and corrections.
Let’s put these ideas to work.
Planning a leadership event, developing your team or rethinking your strategy? Let’s discuss how I could support your organisation through a keynote, executive workshop or advisory engagement.
Book a Call with Prof.Christian