AI SEO Projects ·Machine-Readable Infrastructure
Extractability scorecard & rebuilds
Last reviewed:
- Owner
- Design system + content engineering
- Metric
- Extraction accuracy on spec questions, before and after, for rebuilt pages
- Clock
- Retrieval layer —the two clocks
Model comprehension drops sharply on irregular HTML and complex tables — which is exactly the format of hardware spec sheets and edition-comparison grids. The vertical’s most valuable content is its least parseable. This project measures that gap and closes it where the revenue justifies the rebuild.
Why this, mechanically. Answer engines reward modular formatting — short paragraphs, explicit headings, semantic tables, direct answers — and cited passages are nearly twice as likely to use definitive language. A spec table rebuilt into clean semantic HTML with a prose summary becomes extractable without losing its visual richness, and solving that in a design-system component solves it once for the whole catalog instead of per page.
Deliverable. A scoring rubric (heading structure, paragraph length, table semantics, answer-first placement); scores for the highest-revenue pages per archetype; a rebuild backlog ranked by revenue × score gap; and a reusable “spec block” design-system component.
Finish line. The flagship product pages are migrated to the spec-block component, and before/after prompting shows measurably better spec-answer accuracy.
Procedure: audit content extraction. This scorecard feeds the content-area rebuild work — infrastructure sets the structure standard, content authors to it.