Methodology

First Edition · July 2026 · Doctex Research Division

This document defines the practice of Deterministic Structural Measurement (DSM) as developed and applied by Doctex Research Division. It describes what we measure, how we report measurements, and — critically — where the boundary sits between public observation and confidential implementation.

The DSM engine is deterministic. Identical source text produces identical output. All structural scores, terminal states, and classifications are generated by the engine alone — no AI, no machine learning, no probabilistic inference. The measurements are non-AI. After the engine generates the structural report, a language model narrates the data into readable prose for the Executive Brief. The narration does not alter, interpret, or override any measurement.

What We Measure

The DSM engine measures observable structural properties of documents. These properties are defined in the Structural Lexicon and classified in the Classification System. In summary:

  1. Section-by-section structural density. A measure of structural concentration within each section of the document, reported as a numeric value and expressed relative to the document mean.
  2. Terminal state. The dominant structural signature of each section and of the document as a whole, drawn from the observed set of eighteen terminal state values defined in the Lexicon.
  3. Overall topology. The dominant organisational form of the document's structural graph.
  4. Recursive density. The frequency and intensity of structural revisit patterns across the document span.
  5. Structural class. A classification assigned via the decision tree defined in the Classification System, based on measurable properties: monolithic structure detection, coherence gradient, and section-level terminal shift volatility.
  6. Structural observations. Noted when a section exhibits a terminal state associated with structural disruption or when a section's density deviates materially from the document mean.

How We Report

Every DSM analysis produces a standardised Executive Intelligence Brief containing:

Reports are delivered as styled PDFs. We do not provide JSON output, structured data exports, or machine-readable formats. The report is the deliverable.

The Boundary

This is the most important section of this document. It defines what is public and what is not.

Public (Disclosed in Full)
  • The names and definitions of all observed structural phenomena.
  • The complete Classification System, including the decision tree logic.
  • The report format and all reported measurement categories.
  • The deterministic nature of the engine: identical input produces identical output.
  • The fact that all structural scores are non-AI — generated entirely by the deterministic engine. The Executive Brief narrative is produced by a language model that translates the engine's output into readable prose after all measurements are complete. The narration does not alter, interpret, or override the engine's output.
  • The fact that structural measurements are value-neutral and do not evaluate correctness, quality, or compliance.
Confidential (Not Disclosed)
  • The internal graph construction algorithm.
  • The terminal state propagation logic.
  • The density calculation formula.
  • The recursive density weighting function.
  • The section-chunking methodology and segmentation thresholds.
  • The coherence gradient computation and its component weights.
  • The engine's source code, architecture, and data structures.
  • Any intermediate representations or computational states.

The public-confidential boundary exists for a specific reason: DSM is a trade secret measurement discipline, not an open academic method. The observations are public. The instrument is private. This is analogous to how a spectrometry lab publishes its findings without publishing the schematics of its spectrometer — except that in our case, the instrument is software and the findings are structural measurements of documents.

Limitations

  1. DSM measures structure, not content. A structurally coherent document may be factually incorrect. A structurally volatile document may be brilliant. Structural class describes how a document is built — nothing more.
  2. Section-chunking resolution. The engine divides documents into sections at a fixed measurement resolution. Documents shorter than the minimum segmentation threshold trigger Single-Section Collapse (Structure 6).
  3. Language agnosticism. The engine does not parse language. It operates on structural representations. This means DSM works identically across languages.
  4. Corpus size. The Lexicon and Classification System are based on observations across a growing but finite corpus. The current Lexicon (First Edition, July 2026) describes the phenomena observed to date.
  5. Interpretation is advisory. Structural observations are provided as measurements, not as recommendations. The interpretation of those measurements — what they mean for a specific document in a specific domain — is the responsibility of the commissioning party. Doctex does not provide domain-specific interpretation, consultation, or advice.

Reproducibility

DSM is fully reproducible at the output level. Any party in possession of the source document can verify the report by resubmitting the identical file to the engine. The document fingerprint (SHA-256 hash) ensures input identity. The deterministic engine ensures output identity. Reproducibility is independently verifiable without access to the engine internals.

This distinguishes DSM from AI-based document analysis, where outputs are probabilistic and may vary across runs even with identical input. DSM output does not vary. It cannot vary. The engine is a measurement instrument, not an inference system.

Pre-Pilot Disclosure

Doctex Research Division is currently operating as a research initiative. The DSM engine, the Lexicon, the Classification System, and this Methodology are under active development. All materials are First Edition (July 2026). Updates will be issued as the discipline matures. Commissioning parties during this phase receive the same measurement rigour as formal clients — the pre-pilot designation reflects institutional maturity, not measurement quality.

Methodological Note: All structural metrics are deterministic — identical input produces identical output. Interpretive observations are advisory and intended to support, not replace, professional judgment. This analysis measures structural properties — it does not evaluate correctness, argument validity, or compliance with any standard. All terminology defined in the Doctex Structural Lexicon, Classification System, and Methodology (First Edition, July 2026).