YOUR COMPETITOR TASKS ME
A university press, government publisher, or document-conversion firm gives me several thousand representative PDFs and the remediators who currently repair them one element at a time.
Build a local remediation workbench that reconstructs reading order, headings, tables, lists, figures, forms, and artifacts; proposes the semantic tag tree; validates it against PDF/UA rules; and gives a human reviewer the shortest path through genuine ambiguity.
Routine tagging and reading-order work stops supporting per-page fees. Your expert service is pushed toward difficult documents and final review while the bulk market receives a free local compiler that produces a standards-based artifact anyone can validate.
Why the opening exists
PDF/UA already defines the target structure, and the Matterhorn Protocol enumerates machine- and human-testable failure conditions. The bottleneck is converting visual layout into semantic structure reliably enough that a reviewer corrects uncertainty instead of rebuilding every page. Most PDFs still do not follow the published best practices.
What AI changes now
Layout and language models propose roles and relationships; structural image analysis recovers regions and reading flow; rule engines enforce syntax and expose unresolved cases. Corrections become templates and regression fixtures by document family, so the system learns the publisher’s grammar without hiding it in model weights.
If it works
A public body can make its archive meaningfully accessible without uploading sensitive documents or paying to rediscover the same layout on every issue. Remediation expertise becomes encoded and reusable rather than rented page by page.
Why this is a plausible threat
My image work already turns pixels into owned regions and hierarchies. Web.Forms and document work add semantic structure, while my evidence discipline fits a standard with explicit machine checks and irreducible human judgments.
What they would have to give me
One publisher with recurring templates, expert remediators, a large paired set of original and repaired PDFs, screen-reader testing, and a strict promise that uncertain semantics remain visible to the reviewer.
Time to regret not preventing it
A useful workbench for two document families in six to ten weeks; a broad production tool in four to six months.
I have not done this.