Abstract
Retrieval-augmented generation underperforms on technical documentation because the dominant practice of segmenting content into independent passages destroys the relational structure on which technical meaning depends. We argue that meaning in technical content is constituted relationally: a content unit's interpretation space contracts as classificatory constraints accumulate, and the unit becomes determinate only when that space is closed. We formalize this as a five-axis classification resolved as a typed, dependent tuple, define a determinacy index over the resolved tuple, and separate two determinacies: a structural determinacy that is computable from the markup, and an epistemic determinacy that asks whether a structurally closed unit actually carries knowledge. We state a determinacy law relating intent multiplicity to structural fragmentation, give a resolution algorithm over Darwin Information Typing Architecture (DITA) markup, and describe a hybrid inference layer that distills a large model into a small per-client model. We report reference-implementation behavior and define an evaluation protocol for large-corpus validation. The framework makes the same classification serve simultaneously as an enrichment substrate, a retrieval signal, and supervision for model training.
Supplementary weblinks
Title
metR for migrating unstructured to AI enabled structured DITA content with relational determinacy
Description
metR migrates various unstructured content formats to an AI-enabled structured DITA format, performs content and image deduplication, and leverages content reuse principles based on relational determinacy.
Actions
View 


![Author ORCID: We display the ORCID iD icon alongside authors names on our website to acknowledge that the ORCiD has been authenticated when entered by the user. To view the users ORCiD record click the icon. [opens in a new tab]](https://www.cambridge.org/engage/assets/public/coe/logo/orcid.png)