← Back to Research
NORTHLINE RESEARCH SERIES · GOVERNED AI SYSTEMS · NR-03

Evidence Coverage and Traceability

Knowing What Supports a Conclusion

Steven BoyleNorthline Advisory
|October 2026 · 25 min read
Abstract

Generative AI can usually explain why an answer appears reasonable. That does not mean the system can establish which governed evidence actually supported the answer. This distinction becomes important when AI participates in consequential work. A model may select evidence, interpret its significance, synthesize several sources and generate a persuasive conclusion. If the relationships between evidence, proof requirements and output are not preserved during that process, later traceability becomes reconstruction. The system can ask the model what probably supported the conclusion, but it can no longer distinguish remembered structure from a plausible after-the-fact explanation. This paper examines an implemented evidence-coverage and traceability architecture developed within the Executive Positioning System (EPS). EPS separates semantic judgment from deterministic relationship control. Models determine meaning: relevance, significance, selection and synthesis. The surrounding framework preserves governed identities, declared support relationships, selection history, downstream survival, explicit output references and diagnostic state. A deterministic diagnostics layer then evaluates those governed relationships without making a model call. It can identify whether declared proof obligations have surviving support, whether the same evidence is carrying a disproportionate share of the argumentative burden, whether unused governed evidence remains available for a proposed repair, and where exact output traceability is or is not available. The central proposition is that Traceability is not the ability to produce a plausible explanation of where a conclusion came from; it is the ability to preserve and inspect the governed relationships that actually support it. The architecture also exposes a second requirement: a trustworthy system must represent what it cannot establish. Where exact support has not been preserved, EPS reports the relationship as unavailable rather than reconstructing provenance through semantic similarity.

Get this research paper

Receive the complete paper as a PDF by email.

1. A good answer is not necessarily a supported answer

Large language models are exceptionally good at explanation.

Give a model a conclusion and a body of evidence and ask why the conclusion is justified. In many cases it will produce a coherent account connecting the evidence to the result.

That ability is useful.

It is not the same thing as traceability.

This distinction has established precedent in both provenance and language-model research. W3C PROV provides a formal model for representing and interchanging provenance information, including relationships among entities, activities and agents [1]. In natural-language generation, Rashkin et al. separately define attribution by asking whether externally grounded generated statements can be verified against identified source material [2]. Both approaches treat the relationship between an output and its source history as something that must be represented or evaluated rather than assumed from the plausibility of the output itself.

Consider a system that has generated an executive recommendation, risk assessment, policy argument or application document. The output may contain several conclusions assembled from many records. After the fact, a reviewer asks:

Which evidence supports this statement?

A sufficiently capable model can search the available material and produce a convincing answer.

But what has actually been established?

Perhaps the identified evidence genuinely informed the original reasoning. Perhaps the model has found different evidence that also happens to support the finished statement. Perhaps several sources could plausibly justify it. Perhaps the relationship never existed at all and the model has reconstructed one because it was asked to provide an answer.

From the reader's perspective these explanations may look almost identical.

Architecturally, they are not.

The first describes provenance. The others describe inference.

That distinction became increasingly important while developing EPS. The system had already begun separating authoritative evidence from generated artifacts, and authoritative versions from merely newer ones. The next problem was more granular:

How can the system establish what supports the content inside an authoritative artifact?

The answer could not simply be, "ask the model."

A model's ability to explain its output does not establish the historical relationship between the evidence and the decision that produced that output.

Research on model-generated reasoning provides a related caution. Turpin et al. showed that chain-of-thought explanations can systematically misrepresent factors that influenced a model's prediction and can rationalize answers without revealing those influences [3]. Their study does not examine EPS provenance, but it reinforces the narrower point that a generated explanation should not be treated as a reliable historical record of how an output was produced.

The first principle of NR-03 is therefore:

Explanation is not traceability.

And the stronger form:

A system should not confuse its ability to explain a conclusion with its ability to prove what supported it.

2. Meaning and relationship integrity have different owners

The problem is not solved by removing semantic judgment from the model.

The model is often the component best suited to decide that one piece of evidence is relevant to a particular requirement, that another is weak, or that two experiences together support a broader argument.

Those are meaning-making tasks.

What EPS changes is who owns the resulting relationship.

An accepted architectural decision in the system states [6]:

Deterministic graph state is framework owned.

Earlier EPS contracts asked models to return not only semantic decisions but also reverse graph relationships describing the consequences of those decisions. That created redundant state.

For example, a model might declare that Evidence Record A supports Proof Obligation 3 and also return a separate structure saying that Proof Obligation 3 is supported by Evidence Record A.

Both represented the same relationship.

If one changed and the other did not, the system could contain two conflicting versions of its own evidence graph.

The architecture was changed.

The model now provides the semantic judgment. The framework constructs deterministic reverse relationships, coverage structures and derived diagnostic state.

The distinction is important:

The model may determine:

  • whether a record is relevant;
  • whether it should be selected;
  • what proof requirement it supports;
  • whether its relationship is primary or secondary;
  • how several records should be synthesized;
  • what argument should be written.

The framework determines:

  • whether the referenced record actually exists;
  • whether the proof identifier is valid;
  • how the declared relationship is stored;
  • how reverse mappings are constructed;
  • whether identifiers resolve consistently;
  • which governed version created the relationship;
  • what deterministic diagnostics follow from the relationship.

The framework does not independently invent the semantic assertion: Record X supports Proof Obligation Y.

But once that relationship has been declared within a governed contract, the model should not also be responsible for maintaining its structural integrity.

Formal provenance models distinguish represented provenance relationships from the content of the entities involved [1]. NR-03 applies that separation operationally to model-mediated evidence selection: the model may judge that a support relationship exists, but the framework owns the integrity of the persisted relationship after declaration.

This produces a useful architectural boundary:

The model determines meaning. The framework governs the integrity of the relationships created by that meaning.

That is a narrower and more useful separation than "AI versus rules." Probabilistic reasoning remains responsible for semantic work. Deterministic infrastructure is responsible for making the declared result inspectable.

Diagram separating model-owned semantic judgment from framework-owned deterministic traceability, showing a declared evidence-to-proof relationship flowing into identity validation, persistence, reverse mappings and diagnostics.
Figure 1. Semantic Judgment and Deterministic Traceability
The model performs semantic judgment about relevance, meaning and support. Once a governed support relationship is declared, deterministic infrastructure validates identities, persists the relationship, constructs reverse mappings and computes diagnostics.

3. Coverage begins with what the output must establish

Many traceability approaches begin with the source.

EPS begins one step earlier.

Before asking whether an output contains supporting evidence, the system represents what the output is supposed to establish.

In the Stage 07 evidence-selection process, those requirements are represented as proof obligations.

A proof obligation describes something the application needs to demonstrate within a particular role or decision context. Examples might include:

  • evidence of enterprise governance;
  • evidence of operational improvement;
  • evidence of large-scale transformation;
  • evidence of executive influence;
  • evidence of a particular capability required by the target mandate.

The important architectural point is that proof obligations have governed identities.

They are role-qualified.

A proof identifier such as PO1 cannot be treated as globally meaningful because two roles may each contain a local PO1 representing completely different requirements.

EPS therefore constructs identities such as:

Role A::PO1

and:

Role B::PO1

as distinct proof references.

Evidence-based generation research makes a related distinction between producing citations and establishing whether the claims in an output are actually supported. Gao et al.'s ALCE benchmark evaluates citation quality separately rather than treating the presence of citations as sufficient evidence of support [4]. NR-03 extends that concern upstream by representing what an output is expected to establish before evaluating whether governed evidence covers those obligations.

Evidence coverage has meaning only in relation to a defined requirement.

Counting citations does not establish coverage. A document could contain ten evidence references and still fail to support its most important proposition. Conversely, a critical proof obligation might be strongly supported by two high-quality records even if the document contains few total references.

The system therefore asks: What needs to be proven? and then: What governed evidence has been declared in support of it?

That creates the basis for deterministic coverage analysis.

4. Declared support can be evaluated deterministically

Once the semantic relationships have been established, EPS's Phase 2 Graph Diagnostics layer can evaluate their structural condition without asking a model to reassess their meaning.

For every declared Stage 07 proof obligation, the diagnostic layer resolves its role-qualified proof identity, its obligation text, its declared importance, primary and secondary supporting evidence records, each record's selection disposition, and whether those supporting records survived into promoted output.

The implementation then applies an explicit coverage policy. Current EPS policy defines five states.

Coverage state Deterministic meaning
Strong At least two distinct primary support records survive into promoted output
Covered One primary support record survives into promoted output
Partial Declared support resolves, but no primary support survives into promoted output
None No support has been declared
Unavailable No authoritative coverage-map entry exists

The important point is not that these five categories represent a universal theory of evidence quality. They do not. They represent an explicit and inspectable policy.

That distinction matters.

A model could be asked: Is this proof obligation strongly supported? It might reasonably consider evidence quality, persuasive strength, context, redundancy and semantic fit.

The deterministic diagnostic asks a narrower question: Given the support relationships already declared by the governed process, how many primary records survive into authoritative output?

That answer is reproducible.

The architecture therefore separates two different questions.

Semantic question: Is this evidence meaningful and persuasive support for the requirement?

Deterministic question: Does the governed relationship structure satisfy the explicit coverage rule?

Both may be useful. They should not be represented as the same kind of knowledge.

Attribution research similarly treats source support as a distinct property of generated content [2,4]. NR-03 goes further architecturally by preserving the declared support graph and computing policy-defined diagnostics without asking a model to reinterpret the relationship.

Semantic judgment determines what evidence means. Deterministic diagnostics determine what follows from the relationships the system has actually recorded.

5. Evidence availability is not the same as evidence survival

Coverage also exposed another distinction.

An organization may possess strong evidence for a proposition without that evidence appearing in the final authoritative output.

Consider a simplified workflow. A proof obligation is defined. Several records are identified as potentially supporting it. One is selected. A model writes a section using that record. The section is revised. A whole-document review changes the composition. A later version is promoted.

At the beginning of the process, evidence existed. At the end, the final document may or may not still contain it.

Those are different states.

EPS therefore preserves evidence history across multiple stages. The case overlay can record whether evidence was: selected; held in reserve; rejected; used in Stage 08; incorporated into a Top Section; retained through Stage 09; used in a Stage 10 argument; associated with human review; or preserved through promotion.

The diagnostic question then becomes more useful:

Did the evidence required to support this proof obligation survive into the authoritative output?

This changes how coverage is interpreted. Suppose a proof obligation has excellent primary evidence in the approved evidence package, but none of that evidence survives into the promoted output. It would be misleading to describe the final output as fully covered merely because the evidence exists somewhere in the system.

EPS therefore distinguishes the evidence universe from output survival.

Attribution research typically evaluates whether finished generated content is supported by identified sources [2,4]. NR-03 adds a workflow dimension: evidence may be valid and available earlier in the process yet fail to survive into the particular artifact that ultimately acquires authority.

Having evidence available is not the same as using evidence in the result that actually acquires authority.

Traceability needs to follow evidence through the decision process, not stop when retrieval succeeds.

Black-and-white journal diagram showing evidence moving from a proof obligation through declared support, selection, generated output and promotion, with branches indicating that evidence may be selected, reserved, omitted, survive or fail to survive.
Figure 2. Evidence Coverage Through the Governed Workflow
Evidence may exist in the governed evidence universe without surviving into authoritative output. Coverage therefore depends not only on evidence availability and selection, but on whether declared support remains present through generation, review and promotion.

6. Traceability exists at more than one level

The term traceability is often used as though it describes one capability.

In practice, several different relationships are involved.

Provenance can be represented at different levels of relationship granularity. W3C PROV models provenance relationships among entities, activities and agents [1], while attribution research evaluates support relationships between generated statements and identified sources [2,4]. NR-03 uses those broader precedents but distinguishes the traceability levels required by the EPS workflow.

EPS currently distinguishes at least four useful levels.

Artifact lineage

NR-02 examined this level [6]. Artifact lineage answers: Which governed artifact or version contributed to another artifact? This is structural system provenance.

For example: Stage 08 Section Version 3 → Stage 09 Resume Version 1. It tells the system which governed objects depend upon which others.

Selection traceability

Selection traceability asks: What happened to this evidence record within this case?

The EPS case overlay can preserve information such as: active, reserve or rejected disposition; model rank; human rank; final order; selection group; source version; Stage 08 use; Stage 09 survival; and promotion relationships.

The evidence record remains the same authoritative record. The selection event is opportunity-specific.

Output-reference traceability

Where an output contract explicitly preserves a support relationship, the system can answer: Which records support this particular output reference or block? This is more granular than artifact lineage.

Statement-level traceability

The strongest possible question is: Which exact evidence supports this exact statement?

That capability is attractive. It is also where systems are most likely to overstate what they know.

EPS deliberately does not infer this relationship when it has not been explicitly captured. That design choice is central to NR-03.

7. Unavailable is a valid traceability result

The Phase 2 implementation creates an exact output trace only where promoted artifacts already contain an explicit support relationship.

Several current contracts provide those relationships.

Stage 08 employment sections: An achievement_traceability record can connect an evidence record to a supporting_output_reference.

Top Section: A block can carry explicit supporting_record_ids.

Stage 10 cover letter: argument_support can connect evidence records to a governed internal argument.

These relationships allow deterministic tracing. But their granularity differs. A Top Section relationship may operate at block level. A cover-letter relationship currently operates at argument level. A Stage 08 output reference can be more precise.

The diagnostic layer does not quietly normalize these differences into a fictional sentence-level certainty.

It reports the trace at the granularity actually supported by the contract.

And where exact support is absent, it reports the capability as unavailable.

This is deliberately stricter than asking a model to reconstruct likely support after generation. Existing cited-generation benchmarks show that even systems designed to produce sourced answers can leave claims without complete citation support [4]. The architectural response in NR-03 is therefore not to infer the missing edge more intelligently, but to preserve the distinction between an explicit governed relationship and a plausible reconstruction.

It does not: align free text to evidence through embedding similarity; ask a language model which sentence probably came from which record; invent sentence boundaries; or infer unsupported edges to make the trace look complete.

This produces one of the most important principles in the research programme:

Unavailable is a legitimate result. Inferred provenance presented as fact is not.

This is not merely conservative engineering. It defines an epistemic boundary.

The system distinguishes: what it knows because a governed relationship exists; what a model might reasonably infer; and what has not been established.

A system that cannot represent that distinction will tend to manufacture completeness.

8. A governed evidence chain follows the decision

Traditional provenance often records document relationships.

For example: Document A → Document B. That is useful. It becomes less sufficient when AI participates in complex synthesis.

A more meaningful chain may be:

Evidence Record → Proof Obligation → Selection Decision → Output Reference → Review → Promotion

This is the sequence through which evidence becomes part of authoritative work.

I describe this as a governed evidence chain:

A governed evidence chain is the inspectable sequence connecting authoritative evidence to a declared proof requirement, its selection and use, and the promoted output in which it survives.

The concept matters because the same evidence can participate differently in different cases. An achievement may be highly relevant to one opportunity, reserve evidence in another, and rejected for a third.

Those contextual decisions should not rewrite the underlying evidence record.

The EPS knowledge architecture therefore separates the shared evidence backbone from opportunity-specific case overlays. The authoritative evidence remains stable. The case overlay records what happened to that evidence within a particular decision.

That structure becomes especially important if the system later learns from historical decisions. History can describe what happened. It should not silently change what the evidence is.

Six-stage governed evidence chain connecting an authoritative evidence record to a proof obligation, selection decision, output reference, human review and promoted authoritative artifact.
Figure 3. The Governed Evidence Chain
Traceability follows the decision pathway from an authoritative evidence record through proof obligation, selection, explicit output use, review and promotion. Each transition preserves a different aspect of the evidence relationship.

9. Evidence availability should be checked before AI attempts a repair

Coverage diagnostics become particularly useful when the system begins optimizing its own work.

Suppose an evaluator identifies a weakness: The document does not provide enough evidence of operational transformation.

The obvious response is to ask another model to improve the section.

But there is a prior question: Does additional governed evidence exist that could support the proposed improvement?

If the answer is no, rewriting may only change language. It cannot create proof.

EPS therefore implements an evidence-availability check for proposed repairs.

The request does not contain arbitrary reviewer prose. It identifies one or more explicit role-qualified proof references. The deterministic service then examines the evidence relationships already declared for those obligations.

A candidate record is considered eligible when it: resolves to an authoritative evidence record; is Active or Reserve rather than Rejected; is not already used in a promoted role section, Top Section block or promoted cover-letter argument; and is not explicitly excluded by the request.

The result is available or not_available.

An unknown proof reference fails closed. An unresolved supporting record fails closed.

The system does not ask a model to search the evidence corpus for something that "sounds like" it might address the problem.

This produces a critical distinction:

A writing deficiency and an evidence deficiency are not the same problem.

If relevant unused evidence exists, another optimization cycle may be useful. If no governed evidence remains, additional rewriting may simply produce increasingly polished language around the same evidentiary ceiling.

This becomes important in NR-04 and NR-05, where optimization itself becomes governed.

10. Reuse exposes hidden concentration of proof

Evidence coverage does not only concern gaps.

A system can also rely too heavily on a small amount of evidence.

A document may contain different paragraphs, different sections and different arguments while repeatedly depending on the same underlying record. From the surface, the output looks varied. From the evidence graph, it is concentrated.

EPS therefore calculates evidence reuse across distinct governed output references.

The current Phase 2 implementation counts uses across: promoted Stage 08 role-section references; promoted Top Section blocks; and promoted Stage 10 cover-letter arguments.

Duplicate relationship rows do not increase the count. Only distinct governed uses do.

The current policy classifies reuse as: two uses — low; three uses — medium; four or more uses — high.

Again, these thresholds are policy. They do not prove that four uses are inherently problematic.

What the diagnostic can truthfully say is:

This evidence record is carrying a large amount of the argumentative burden.

That observation can inform review. It may reveal: repetitive evidence; a narrow proof base; over-dependence on one role; a lack of supporting diversity; or a legitimate anchor achievement that deserves repeated use.

The semantic judgment about whether reuse is appropriate remains separate. The deterministic diagnostic identifies the concentration.

This creates another useful principle:

Content diversity is not necessarily evidence diversity.

A model can paraphrase the same proof many times. Stable record identities reveal whether the underlying evidence has actually changed.

11. Complementary proof is not the same as multiple records

EPS also computes complementary-proof diagnostics where two or more distinct records support the same proof obligation.

The diagnostic preserves information including: record identity; source role; canonical domain; and number of contributing records.

This can help expose whether the system has multiple sources of support.

But the diagnostic deliberately stops short of claiming that those records are semantically complementary.

Consider two records supporting a governance proof obligation. The deterministic system can say:

Two distinct records from two governed identities support this obligation.

It should not automatically say:

Together they demonstrate both strategic governance and operational implementation.

That second statement requires interpretation. A model may make that interpretation. A reviewer may agree or disagree. The graph itself should not quietly transform record multiplicity into semantic meaning.

This is an example of the broader architectural discipline in EPS:

Deterministic facts should remain deterministic facts. Semantic conclusions should remain semantic conclusions.

The point is not to minimize what the model can do. It is to avoid representing inference as structure.

12. Selection history makes evidence decisions auditable

The implemented Selection History Service creates a stable case-level projection of evidence decisions.

The service records opportunity-specific events without changing the authoritative evidence graph [6].

For each selection event it can preserve information such as: evidence record identity; opportunity identity; role identity; active, reserve or rejected disposition; model rank; human rank; final ordering; canonical primary domain; secondary domains; selection group; Stage 08 use; rendered heading; Stage 09 survival; and provenance.

The service also captures bounded argument support from Stage 10.

This makes an important question answerable:

What happened to this evidence in this specific decision context?

That is different from asking: What does this evidence mean in general?

The former is history. The latter is knowledge.

The distinction becomes important if future systems use historical patterns to inform new decisions. A frequently selected achievement may be useful historical information, but that history should not automatically become a new fact about what the achievement proves [8].

For NR-03, the important point is simpler:

Traceability should preserve the decisions made about evidence without allowing those decisions to rewrite the underlying evidence.

13. Missing identity should fail closed

The Phase 2 diagnostic service has explicit integrity behaviour. A diagnostic build fails when required governed relationships cannot be resolved.

Examples include: no approved Stage 07 package exists; the Phase 1 ledger cannot be constructed or validated; required graph, repository or ontology pins are unavailable; a coverage map references a proof obligation that was never declared; a support record cannot be resolved to the appropriate role's evidence history; duplicate role-qualified proof identities exist; an evidence-availability request references an unknown proof obligation; or a governed diagnostic artifact cannot be persisted correctly.

This may appear strict. That is intentional.

The alternative is to let the system continue by reconstructing what it thinks the missing relationship probably was. That would improve completion rates. It would weaken integrity.

The concern is consistent with the broader distinction between generated explanation and recorded provenance [1,3]. A plausible reconstruction may be useful as inference, but it is not equivalent to a preserved historical relationship.

The current test suite explicitly verifies this failure behaviour. Unknown repair proof references raise an integrity error. Unresolved support IDs also fail rather than being ignored or semantically substituted.

The architectural principle is:

If the system cannot resolve the governed evidence relationship, it should not manufacture one simply to complete the workflow.

That becomes increasingly important as systems become more autonomous. A model that is rewarded for producing an answer will usually attempt to produce one. The deterministic layer must be capable of saying that the required relationship does not exist.

14. Diagnostics should not become another reasoning engine

Knowledge graphs and diagnostic services can themselves become sources of hidden inference.

EPS deliberately avoids that in Phase 2 [6].

The diagnostic service performs deterministic operations over governed identifiers and relationships: joins; set operations; counts; classifications; projections.

It makes no model call. It does not infer new semantic graph facts. It does not select evidence for a future application. It does not alter evidence authority. It does not promote or reject artifacts.

A later model may consume a bounded diagnostic result and use it to explain a problem or propose an action. But the model does not become the author of the underlying diagnostic fact.

The Phase 2 contract explicitly prevents downstream consumers from using the diagnostic layer to: create new support edges; change coverage classifications; resolve missing record identities; convert arbitrary reviewer prose into proof references without another governed semantic contract; or alter evidence authority.

This is an important architectural constraint:

Diagnostics should expose governed relationships, not quietly become a second source of semantic authority.

That principle keeps the system inspectable. The graph describes what the governed workflow recorded. The model interprets what that information might mean. Those are different responsibilities.

15. The deterministic layer can be reproducible even when the AI is not

One of the Phase 2 acceptance requirements states that repeated diagnostic builds from unchanged inputs must produce identical diagnostic content apart from the generation timestamp.

The tests verify this.

That property is important because the underlying AI operations may themselves be probabilistic. The model may not produce identical wording or identical semantic judgments if a stage is rerun.

But once a governed set of decisions and relationships has been persisted, the diagnostic interpretation of that state can be deterministic.

Formal provenance representation is designed to preserve and interchange provenance relationships independently of the process that originally generated the associated content [1]. NR-03 applies that principle to persisted AI-system state: probabilistic reasoning may vary, while diagnostics over an unchanged governed relationship graph remain reproducible.

The pattern is:

Probabilistic reasoning → Governed declarations and artifacts → Deterministic diagnostics

This illustrates a broader point about governed AI architecture. The objective is not to make the entire system deterministic. That would eliminate many of the capabilities for which AI was introduced.

The objective is to identify which parts of the system need reproducibility. In this case: semantic interpretation can remain probabilistic; evidence identity cannot; declared relationships cannot; diagnostic classification should not vary because a model happened to interpret the same graph differently on Tuesday than it did on Monday.

This produces a particularly useful architectural result:

A probabilistic system can still have a reproducible account of its governed state.

16. What the system can know

The implementation examined here makes the system's epistemic boundary unusually visible.

Given its current governed contracts, EPS can deterministically establish: which proof obligations were declared; which evidence records were declared in support; whether those record identities resolve; whether support was primary or secondary; whether evidence was selected, reserved or rejected; whether it was used in promoted Stage 08 output; whether it survived Stage 09; whether it supports a governed Stage 10 argument; how many distinct outputs use the record; whether additional declared unused evidence remains; which exact output references contain explicit support relationships; and which repository, graph, ontology, artifact and policy versions underpin the diagnostic.

There are also things the deterministic layer cannot establish from those relationships alone.

Attribution frameworks distinguish whether generated content is supported by identified source material from broader questions about the ultimate truth of the conclusion [2]. NR-03 makes an analogous architectural distinction: the graph can establish that a governed support relationship exists without proving that the original semantic judgment was correct.

It cannot prove: that the evidence is persuasive; that the model's original support judgment was semantically correct; that two records are genuinely complementary in meaning; that a particular sentence was supported by a record if that relationship was never captured; that missing evidence exists outside the governed evidence universe; that a conclusion is true merely because several records support it; or that the argument presented is the best possible argument.

Those require other forms of judgment.

This boundary should not be treated as a weakness. It is part of the governance model:

The system becomes more trustworthy when it explicitly distinguishes what it knows, what it can infer and what remains unavailable.

Two-column journal figure contrasting relationships the system can deterministically verify with claims it must not invent, ending with the principle that unavailable is a legitimate result.
Figure 4. The Boundary of System Knowledge
Deterministic diagnostics can verify governed identities, declared relationships, selection state, output survival, reuse, availability and version integrity. They must not manufacture missing provenance, semantic meaning or factual certainty that the governed record does not establish.

17. What organizations need to decide

Broader AI risk-management guidance emphasizes governance, documentation, transparency, evaluation and accountability across the AI lifecycle [5]. The questions below are NR-03's more specific architectural interpretation of what evidence traceability requires in a consequential AI workflow.

Organizations do not need EPS's exact graph architecture to apply the underlying principles. They do need to make several decisions explicitly.

What must the output establish? If the requirements are not represented, evidence coverage cannot be meaningfully assessed. A system may have many sources without proving the things that actually matter.

When are evidence relationships captured? If the organization waits until after publication to reconstruct support, traceability will depend on inference. Relationships should be preserved as part of the governed workflow where practical.

Do support relationships resolve to stable evidence identities? A citation string or copied paragraph is weaker than an authoritative evidence record and version. Stable identity makes downstream validation possible.

Can the system distinguish availability from use? Evidence existing somewhere in a repository does not mean it appears in the authoritative output. The workflow should be able to distinguish the two.

Can the system return "unavailable"? A traceability system that must always produce an answer will eventually manufacture certainty. Unknown and unavailable states need first-class representation.

Are structural facts kept separate from semantic judgments? Counts, identifiers, survival and explicit relationships can be deterministic. Persuasiveness, significance, complementarity and interpretation remain semantic. Blurring those categories makes the graph appear smarter while making the system less honest about what it actually knows.

18. Limitations

The implementation supports strong claims about how EPS records and diagnoses evidence relationships. It does not support every possible claim about evidence quality or traceability.

A structurally valid relationship may still reflect poor semantic judgment. The framework can establish that Record A was declared as support for Proof Obligation B. It cannot prove from the graph alone that the original semantic decision was correct. A model or reviewer may have misunderstood the evidence. Structural integrity does not eliminate semantic error.

Coverage states are policy definitions. The current categories — strong, covered, partial, none and unavailable — are explicit EPS policies. They are useful because they are transparent and reproducible. They should not be interpreted as universal scientific categories of evidentiary sufficiency.

Statement-level traceability remains intentionally incomplete. EPS currently emits detailed traces only where promoted contracts contain explicit support relationships. Some outputs therefore have block-level or argument-level traceability rather than exact sentence-level support. The architecture treats missing granularity as unavailable rather than reconstructing it.

Coverage is not truth. Two or more support records can strengthen an argument. Their existence does not prove the conclusion. Evidence quality, interpretation and reasoning remain important.

The architecture has not been experimentally compared with alternatives. NR-03 is based primarily on implemented architecture, deterministic services and tests. It demonstrates that EPS can preserve and reproduce these evidence relationships under its current contracts. It does not establish through controlled comparative experimentation that this architecture is superior to every alternative traceability design.

19. From traceability to governed optimization

Once a system can determine which proof obligations are covered, which are only partially supported, where evidence did not survive, which records are overused, whether additional governed evidence remains, and where support relationships are unavailable — it can make a more informed decision about what to do next.

That changes optimization.

A model should not necessarily rewrite a weak section merely because an evaluator can describe a weakness. The system can first ask: Is this a writing problem? Is this an evidence-selection problem? Is additional unused evidence available? Has the evidence ceiling already been reached? Would another optimization cycle likely improve anything material?

Those questions lead directly into the next paper in this series.

NR-04 examines how an AI system can use governed evaluation and evidence state to improve its own work without allowing iteration to override evidence integrity, authority or the best valid state [7].

Conclusion

Generative AI can explain almost anything it produces. That capability should not be confused with provenance.

In consequential systems, traceability needs to survive the workflow as governed structure: evidence identities, proof obligations, declared support relationships, selection decisions, output references, human review and promotion.

The implementation examined in EPS suggests that semantic judgment and traceability should have different owners.

Models can determine meaning. They can assess relevance, compare evidence, synthesize sources and propose arguments.

Deterministic infrastructure can preserve the relationships produced by those decisions, verify their integrity and compute reproducible diagnostics from them.

Just as importantly, the system must be capable of representing the absence of knowledge. If exact support was not preserved, the correct answer may be unavailable rather than probably this.

Existing work establishes that provenance can be formally represented [1], that generated statements can be evaluated for attribution to identified sources [2], that generated explanations need not faithfully reveal what influenced a model's answer [3], and that cited generation still faces problems of citation correctness and completeness [4]. NR-03 contributes the architectural layer around those concerns: semantic support judgments may remain probabilistic, while the identities, persistence, integrity, coverage and downstream consequences of declared evidence relationships are governed deterministically—and missing relationships remain explicitly unavailable rather than reconstructed as fact.

That boundary matters because trustworthiness is not created by always having an explanation. It is created by knowing which explanations are grounded in governed evidence and which are inference.

A trustworthy AI system should not only know what supports a conclusion. It should know when it does not know.

Research Status and Evidence Basis

This paper is based primarily on implemented EPS architecture, deterministic diagnostic services and associated tests. The repository evidence reviewed includes: EPS Phase 2 Graph Diagnostics Implementation Specification; ADR-003 Deterministic Graph State is Framework Owned; EPS Knowledge Graph Learning Flow; shared graph diagnostics, selection history, and graph model implementations; associated test suites; and Stage 07, Stage 08 and Stage 10 governed support contracts. The implementation and tests support claims concerning EPS's handling of proof coverage, selection history, evidence reuse, evidence availability, explicit output support relationships, deterministic diagnostics, failure behaviour and reproducibility. They do not establish that EPS's coverage-policy thresholds are universally optimal, that structurally valid evidence relationships are necessarily semantically correct, or that this architecture produces superior decision outcomes compared with alternative traceability approaches. Broader architectural propositions should be understood as conclusions derived from the implemented case rather than experimentally established universal results.

References
[1] Lebo, T., Sahoo, S., and McGuinness, D., eds. "PROV-O: The PROV Ontology." W3C Recommendation, 30 April 2013.
[2] Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G. S., Turc, I., and Reitter, D. "Measuring Attribution in Natural Language Generation Models." Computational Linguistics, 49(4), pp. 777–840, 2023. DOI: 10.1162/coli_a_00486.
[3] Turpin, M., Michael, J., Perez, E., and Bowman, S. R. "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." Advances in Neural Information Processing Systems 36, 2023.
[4] Gao, T., Yen, H., Yu, J., and Chen, D. "Enabling Large Language Models to Generate Text with Citations." Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 6465–6488, 2023. DOI: 10.18653/v1/2023.emnlp-main.398.
[5] Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. National Institute of Standards and Technology, 2023. DOI: 10.6028/NIST.AI.100-1.
[6] Boyle, S. Artifact-Oriented AI Architecture: Governing State, Authority and Lineage. Northline Research Series: Governed AI Systems, NR-02. Northline Advisory, 2026.
[7] Boyle, S. Governed Self-Optimization: How an AI System Can Improve Its Work Without Losing Control. Northline Research Series: Governed AI Systems, NR-04. Northline Advisory, 2026.
[8] Boyle, S. Governed Learning: How AI Systems Can Learn Without Turning Experience Into Truth. Northline Research Series: Governed AI Systems, NR-06. Northline Advisory, 2026.
Next in the series
NR-04
Governed Self-Optimization
How an AI System Can Improve Its Work Without Losing Control
Want a copy of this paper?

Receive a link to the complete paper by email.

Steven Boyle
Northline Advisory

Steven Boyle is the founder of Northline Advisory, a technology advisory and research practice focused on technology leadership, enterprise transformation, governance, data and governed AI. His work draws on more than two decades of executive and operational experience across higher education and public-interest organizations.

Related Practitioner Insights
Practitioner Insight · Governance

Governed AI with Gated Learning: Lessons from Building EPS

A governed architecture pattern for introducing learning into consequential AI workflows, illustrated through the development of EPS.

Steven BoyleSeptember 11, 2026 · 8 min readRead →
Related Research
Research Paper · Data & AI

Beyond the Model

A Governed Architecture for Consequential AI Systems

Consequential AI requires an architecture that governs evidence, artifacts, evaluation, authority and learning around the model.

Steven BoyleNR-01October 2026 · 18 min readRead →
Research Paper · Data & AI

Artifact-Oriented AI Architecture

Governing State, Authority and Lineage

An implemented architecture for treating durable, versioned and governed artifacts—not transient model responses—as the unit of enterprise AI control.

Steven BoyleNR-02October 2026 · 22 min readRead →