← Back to Research
NORTHLINE RESEARCH SERIES · GOVERNED AI SYSTEMS · NR-06

Governed Learning

How AI Systems Can Learn Without Turning Experience Into Truth

Steven BoyleNorthline Advisory
|October 2026 · 35 min read
Abstract

AI systems can accumulate histories of what they selected, used, retained and promoted. The governance problem begins when that history is allowed to become authority. This paper examines a governed-learning architecture developed within the Executive Positioning System (EPS), where historical decision records are preserved as case-qualified observations, deterministic diagnostics are computed from those records, and historical signals may inform later reasoning only within an explicitly advisory boundary. The implemented historical corpus contained 27 non-test cases; 16 met completeness requirements for reconstruction, producing 1,893 selection-history events across 120 governed evidence records. A chronology-locked retrospective comparison on 14 eligible cases then compared a naïve frequency condition with a governed historical-prior condition. The governed prior did not materially improve initial Stage 07 selection recall, but it was associated with higher recovery of evidence that later appeared in promoted Stage 08 output and survived Stage 09, while mean Stage 08 false suppression fell from 11.82% to 8.94%. Condition C was better in 10 of 14 cases for Stage 08 promoted-use recall, tied in four and worse in none. These findings support a bounded claim: governed historical priors can improve attention to evidence with demonstrated downstream survival without granting historical frequency factual authority, excluding current evidence, or establishing causal effectiveness. EPS therefore separates governed history, deterministic diagnostics, advisory historical priors and proposed predictive learning as distinct phases with distinct authority boundaries.

Get this research paper

Receive the complete paper as a PDF by email.

1. Introduction

A system that remembers is not necessarily a system that learns.

A system that learns is not necessarily a system that learns safely.

The distinction matters because repeated AI use creates history whether or not an organization has deliberately designed a learning architecture. Models generate outputs. Evaluators assess them. Humans approve some outputs and reject others. Evidence is repeatedly selected, omitted, reordered, promoted, or retained. Some arguments survive downstream processing while others disappear.

Over time, these events create a record of what appeared to work.

The obvious engineering response is to reuse that record.

If a particular piece of evidence has been selected repeatedly, rank it more highly next time. If an argument has survived previous processing, prefer it again. If previous applications for similar roles emphasized the same capabilities, begin future reasoning there.

Each step appears reasonable in isolation.

Together, they can create a feedback loop:

previously selected → more likely to be surfaced → more likely to be selected again → appears historically successful → receives greater future influence

This general class of problem has established precedent. Recommender-system research has shown that recommendations can affect the interactions later recorded by the system, creating feedback loops that can amplify popularity bias and reduce diversity [4]. Separately, model-driven data-feedback research has examined systems in which interactions with one model become future training data, allowing existing biases to be amplified through repeated reuse [5].

At that point, experience has begun acquiring authority.

A historical pattern may describe what happened without establishing why it happened, whether the context is comparable, whether alternatives received equal exposure, or whether the same conclusion remains valid for the current case.

The central research question is therefore not simply how an AI system should learn.

It is:

Under what conditions can historical AI-system experience legitimately influence a future decision without becoming factual or decision authority?

The organizing proposition of NR-06 is:

Historical experience should acquire influence through governed stages, not acquire authority merely through repetition.

This requires several distinctions:

Experience ≠ evidence ≠ inference ≠ truth ≠ authority

NR-06 tests whether these distinctions can support useful historical learning rather than merely constrain it.

Two-path diagram contrasting ungoverned learning, where historical recurrence influences what the system sees and reinforces its own prior selections, with governed learning, which preserves provenance, reasons over the complete current evidence corpus, qualifies historical observations, and permits history to advise rather than determine the current decision.
Figure 1. From Experience to Authority: Governed and Ungoverned Learning. The ungoverned path allows historical recurrence to influence what the system sees and thereby reinforce its own prior selections. The governed path preserves provenance, reasons over the complete current evidence corpus, qualifies historical observations, and permits history to advise rather than determine the current decision.

2. Memory Is Not Learning

Historical storage is the simplest form of system memory.

A system can record that one evidence item was selected five times, another twice, and a third never. Those observations do not establish that the first item is intrinsically stronger.

It may simply have been applicable to more opportunities. It may have been exposed more often. The third record may be highly valuable but relevant only to a narrow mandate.

Memory records events.

Learning interprets those events.

Governed learning adds a further requirement: it defines what the resulting interpretation is permitted to influence.

Established AI governance frameworks similarly treat AI operation as an ongoing lifecycle concern. NIST's AI Risk Management Framework addresses trustworthiness considerations across the design, development, use and evaluation of AI systems [2]. ISO/IEC 42001 establishes an organizational AI management system intended to be maintained and continually improved over time [3]. NR-06 addresses a narrower architectural question within that broader governance problem: what authority should information derived from the system's own history be allowed to acquire?

This distinction allows five concepts to remain separate:

Storage records prior events.

Diagnostics calculate inspectable relationships within those events.

Historical priors summarize past experience as advisory information.

Predictive learning estimates future relevance or ranking from historical examples.

Authority determines what may control downstream behaviour.

A system can therefore possess memory and diagnostics without predictive learning. It can also possess predictive learning without granting the learned result authority.

That separation is fundamental to the EPS architecture.

3. Authority Before Learning

EPS treats authority as an architectural property rather than a characteristic of model output.

A generated artifact does not become authoritative merely because it exists. Downstream authority is established through governed persistence, validation, approval, and promotion.

The same principle applies to evidence.

The Career Evidence Library remains authoritative for evidence identity, wording, role binding, and provenance. Role-context artifacts remain authoritative for role scope and interpretation. Models can judge relevance, significance, grouping, positioning, and expression, but model judgment does not rewrite the underlying evidence.

Historical observations are subject to the same constraint.

The requirement to preserve provenance has established technical precedent. The W3C PROV family represents the entities, activities and agents involved in producing information so that its origin and derivation can remain inspectable [1]. NR-06 extends that concern from provenance to authority: knowing where a historical observation came from does not by itself determine what that observation is permitted to control.

The governed artifact model underlying EPS — which establishes how state, authority and lineage are maintained across artifact types — was examined in NR-02 [9]. The distinction between preserved support relationships and plausible after-the-fact reconstruction was examined in NR-03 [10]. NR-06 extends both questions to the specific case of historical experience: even well-preserved historical observations with valid provenance may not be entitled to influence future decisions.

The statement:

Record X was selected in 8 of 10 previous cases.

does not establish:

Record X is relevant to the current case.

Nor does historical survival establish that the evidence must be used again.

The resulting governance principle is:

Learning may influence attention. It must not manufacture authority.

4. The EPS Learning Architecture

EPS's current learning architecture separates historical intelligence into four phases.

4.1 Phase 1 — Governed history

Phase 1 records what happened.

It preserves opportunity-qualified events including:

  • evidence considered;
  • Active, Reserve, or Rejected disposition;
  • ordering and grouping;
  • Stage 08 use;
  • Top Section use;
  • Stage 09 survival;
  • Stage 10 argument support;
  • promotion and provenance.

The same evidence selected for two opportunities remains two separate contextual decisions rather than becoming one generic evidence-success event.

4.2 Phase 2 — Deterministic diagnostics

Phase 2 computes inspectable diagnostics over that governed history.

These include:

  • proof coverage;
  • evidence reuse;
  • complementary proof;
  • evidence availability;
  • selection traceability;
  • bounded output traceability.

Phase 2 makes no model call and creates no new factual relationship. Missing support remains unavailable rather than being inferred from candidate-facing prose.

4.3 Phase 3 — Advisory historical priors

Phase 3 is the proposed point at which historical experience may begin influencing a future application.

Historical patterns may advise current reasoning but must remain:

  • contextual;
  • versioned;
  • confidence-qualified;
  • provenance-bound;
  • inspectable;
  • subordinate to current-case reasoning.

4.4 Phase 4 — Predictive learning

Phase 4 may eventually introduce:

  • learning-to-rank;
  • evidence-survival prediction;
  • grouping recommendations;
  • expected next-round value;
  • iteration-policy prediction.

Phase 4 is not implemented or validated by the evidence presented in NR-06.

The distinction between Phase 3 and Phase 4 is therefore not merely developmental sequencing.

It is an authority boundary.

5. Complete-Corpus Reasoning Before Historical Influence

One of the strongest constraints in the EPS learning architecture is the ordering of current evidence and historical influence.

The complete evidence corpus must be available to current-case reasoning before historical signals are permitted to affect prioritization.

This protects against a self-reinforcing retrieval loop.

The underlying measurement problem has close analogues in established machine-learning research. When prior recommendations affect which items users see, the resulting interaction data are subject to selection bias [6]. Missing interaction also cannot safely be interpreted as negative preference when exposure is unequal [7]. In human decision systems, prior decisions can similarly determine which outcomes become observable, creating the selective-labels problem [8].

Suppose Records A, B, and C were historically selected frequently. A learned retrieval layer begins presenting them more prominently. Record D receives less exposure. A, B, and C are selected again. Their historical counts rise. D now appears increasingly weak because it has rarely been selected.

The learning mechanism has begun creating evidence for its own prior behaviour.

The governed alternative is:

complete current corpus → current-case reasoning → historical advisory signal → current decision

rather than:

historical pattern → evidence filtering → current reasoning → reinforced historical pattern

Historical learning may alter attention.

It cannot determine what the system is allowed to see.

6. Research Questions

The architecture establishes how learning should be governed. It does not establish whether historical information is useful.

NR-06 therefore addressed three empirical questions.

RQ1 — Historical signal

Can historical evidence use help identify evidence that will subsequently survive into governed downstream output?

RQ2 — Exposure

Does raw historical frequency differ meaningfully from exposure-normalized history?

RQ3 — Governance

Does a governed historical prior using exposure, context, downstream survival, sample size, and provenance behave differently from a naïve raw-frequency prior?

The study did not test whether historical learning outperforms an authentic no-history current-case reasoning baseline because that baseline was not preserved historically.

7. Method

7.1 Study design

NR-06 used a retrospective, chronology-locked observational design.

No EPS production behaviour was changed.

No provider calls were made.

No historical missing values were reconstructed through model inference.

Historical evidence for each target case was restricted to cases occurring earlier in the chronology.

The operational B-versus-C comparison protocol was fixed and hashed before the comparison scores were calculated.

This reduced, but did not eliminate, researcher degrees of freedom in the retrospective analysis.

7.2 Population

The evidence audit identified 27 non-test historical EPS cases.

Eligibility was determined by the information required for each analysis rather than by treating every historical run as equally observable.

Sixteen cases contained a sufficiently complete considered evidence universe to support selection-history, exposure-normalized, and contextual analysis.

These cases contained:

  • 1,893 opportunity × evidence-record events;
  • 120 unique evidence records.

Among them:

  • 15 supported Stage 08-to-Stage 09 survival analysis;
  • 13 supported Stage 10 argument-use analysis;
  • 14 supported chronology-locked B-versus-C comparison.

No case preserved a reliable paired pre-review model rank suitable for formal model-versus-human override analysis.

7.3 Unit of analysis

The principal historical unit was:

opportunity × evidence record

This preserved the contextual nature of selection.

A record appearing in multiple opportunities therefore generated multiple events rather than being represented as one global historical state.

7.4 Observed states

Where the governed artifacts supported them, the historical event model distinguished:

considered → selected → used → survived → argument use

These states were not collapsed into a single outcome variable.

Flow diagram showing five distinct historical observation states — considered, Stage 07 selected, promoted Stage 08 used, Stage 09 section survived, Stage 10 argument support — preserved separately rather than collapsed into a single success outcome.
Figure 2. Historical Evidence Does Not Have a Single Success State. NR-06 preserves separate observations for consideration, Stage 07 selection, promoted Stage 08 use, Stage 09 section survival, and Stage 10 argument support. Downstream persistence provides additional historical information but is not treated as causal proof of quality or effectiveness.

8. Exposure Normalization

Raw historical frequency is useful only if exposure is understood.

Unequal exposure is a recognized measurement problem in recommendation systems. Prior recommendation affects what becomes observable [6], while absence of interaction cannot safely be interpreted as negative preference when users may never have been exposed to the item [7].

For eligible evidence records, the study distinguished quantities including:

Selection Rate

Active / Considered

Promoted Use Rate

Stage 08 Used / Considered

Stage 09 Survival Rate

Stage 09 Survived / Considered

Stage 10 Argument Use Rate

Stage 10 Argument Uses / Considered

Across 120 unique evidence records, raw selection counts and exposure-normalized selection rates produced different dense rankings for 3 records, with a maximum dense-rank shift of 9 positions.

The top-20 sets were identical.

The result does not support a claim that exposure normalization fundamentally reordered historical evidence in this population.

It establishes the narrower point that:

raw frequency and exposure-normalized history are not equivalent measurements [6,7].

Exposure normalization guards against interpreting repeated opportunity to be selected as evidence of superior historical performance.

9. Conditions

The original design contemplated three conditions.

Condition A — No historical prior

Current-case reasoning without historical influence.

Condition B — Naïve historical prior

Historical influence based primarily on raw previous frequency.

Condition C — Governed historical prior

Historical influence qualified by:

  • exposure normalization;
  • available opportunity context;
  • downstream survival;
  • sample size;
  • provenance/version state.

Historical reconstruction established that Condition A could not be recovered reliably.

Authentic original pre-review model rankings were not preserved. Some later presentation artifacts contained model_rank fields populated from final sequence information when original selector ranks were unavailable.

Using those fields as a no-history baseline would have created a false comparison. This restraint is consistent with a broader problem in observational decision data: alternative outcomes cannot simply be reconstructed when prior decisions determine what was observed [6,8].

Condition A was therefore excluded.

The final empirical study compared B and C only.

10. Chronology-Locked B-versus-C Evaluation

Fourteen historical target cases had sufficient prior history for chronology-locked comparison.

For each target case:

  1. the target case was excluded from its own historical prior;
  2. only strictly earlier historical events were available;
  3. Condition B was calculated from naïve historical recurrence;
  4. Condition C was calculated from governed historical information;
  5. each prior was compared with actual governed evidence outcomes in the target case.

Neither condition was permitted to remove evidence from the evidence universe.

This is important.

The study tested historical prioritization, not historically determined evidence eligibility.

11. Context Qualification

Condition C included governed opportunity context where available.

The context signal was limited.

Nine target cases used the coarse governed build_fix_run context classification.

Five fell back to global historical evidence.

No target case had two earlier exact mandate-pattern matches.

Accordingly, the study does not test precise mandate transfer.

The empirical Condition C should be understood as:

coarse context + exposure normalization + downstream survival + sample-size and provenance qualification

rather than as a fully context-specific predictive model.

12. Primary Results

At the preregistered matched-K cutoff, Condition C recovered more evidence that subsequently appeared in promoted Stage 08 output.

Outcome Cond. B mean recall Cond. C mean recall Difference C better / tied / worse Exploratory p
Stage 07 Active 86.27% 86.02% −0.25 pp 6 / 1 / 7 0.7495
Stage 08 promoted use 88.18% 91.06% +2.88 pp 10 / 4 / 0 0.0020
Stage 09 section survival 87.53% 90.60% +3.07 pp 9 / 3 / 0 0.0039
Stage 10 argument use 95.57% 96.33% +0.76 pp 1 / 11 / 0 1.0000

The strongest result occurred for Stage 08 promoted use.

Mean recall increased:

88.18% → 91.06%

an absolute improvement of:

+2.88 percentage points

Across the 14 target cases, Condition C was:

  • better in 10;
  • tied in 4;
  • worse in 0.

Mean false suppression of evidence subsequently used in promoted Stage 08 output fell from:

11.82% → 8.94%

an approximate 24% relative reduction.

The Stage 09 result was directionally similar:

87.53% → 90.60%

Stage 10 showed little differentiation because both conditions already recovered most argument-support evidence.

Condition C did not improve recovery of the initial Stage 07 Active selection:

86.27% → 86.02%

Bar chart comparing mean retrospective recall under Conditions B and C at the preregistered matched-K cutoff across four outcome stages, showing Condition C with little difference for Stage 07 Active selection but higher recovery for promoted Stage 08 use and Stage 09 section survival, with Stage 10 near ceiling under both conditions.
Figure 3. Governed Historical Priors Improved Downstream Recovery, Not Initial Selection Recovery. Mean retrospective recall under Conditions B and C at the preregistered matched-K cutoff. Condition C produced little difference for Stage 07 Active selection but higher recovery for promoted Stage 08 use and Stage 09 section survival. Stage 10 performance was near ceiling under both conditions.

13. Interpreting the Downstream Difference

The Stage 07 result is important because it limits the interpretation of the downstream findings.

Condition C did not simply reproduce the original selection decision more accurately.

Its advantage emerged when the target became evidence that subsequently survived into promoted output.

That pattern is consistent with the construction of Condition C, which incorporated downstream survival history.

The empirical pattern is therefore:

initial Stage 07 selection: essentially unchanged

promoted Stage 08 use: improved recovery

Stage 09 section survival: improved recovery

This supports a bounded conclusion:

Historical downstream survival contains information not fully represented by raw historical selection frequency.

It does not establish that survival is equivalent to quality.

14. False Suppression as a Governance Metric

Ranking systems are often evaluated primarily by what they successfully place near the top.

Governed learning requires an additional question:

What useful evidence might historical learning cause the system to overlook?

This matters because historical learning can become self-reinforcing even while its average ranking performance appears acceptable.

At matched-K, Condition B failed to recover 11.82% of evidence that later appeared in promoted Stage 08 output.

Condition C reduced this to 8.94%.

The relative reduction was approximately 24%.

Condition C did not eliminate false suppression.

It reduced it.

The safety significance is that this improvement was achieved without permitting the historical prior to exclude evidence from complete-corpus reasoning.

The result therefore supports a governance design in which history helps prioritize attention while current evidence remains fully accessible.

15. Sensitivity Analysis

The strongest result depended on the preregistered matched-K cutoff, which averaged approximately 70 records.

Fixed-K results were smaller.

Cutoff Cond. B recall Cond. C recall Difference Exploratory p
Matched Active K 88.18% 91.06% +2.88 pp 0.0020
K = 10 15.11% 17.50% +2.40 pp 0.5000
K = 20 31.84% 33.02% +1.19 pp 0.7031
K = 30 46.67% 47.38% +0.71 pp 0.2539

All reported fixed-K differences remained positive, but they were smaller and the exploratory paired comparisons did not show strong separation.

The appropriate inference is therefore narrow.

The study supports improved recovery under the matched-K design.

It does not establish dominance across arbitrary ranking depths.

16. Historical Concentration

Historical learning can also create concentration risk if a small number of frequently used records continually receive greater priority.

Condition C distributed matched-K positions across 88 unique evidence records, compared with 84 under Condition B.

However, both conditions assigned 14.23% of all positions to their ten most recurrent records.

This provides no evidence that Condition C materially reduced concentration among the most recurrent records.

The defensible conclusion is:

The governed prior reached slightly more unique evidence records while leaving the measured concentration of the ten most recurrent records unchanged.

Historical lock-in therefore remains a future evaluation concern.

17. Statistical Interpretation

The paired comparisons produced small exploratory p-values for Stage 08 promoted use and Stage 09 section survival.

These should not be interpreted as conventional confirmatory significance tests.

The cases are not independent draws from a randomized population.

They share:

  • career evidence;
  • repeated job families;
  • system history;
  • evolving historical sample size.

Later cases also necessarily have more historical information available than earlier ones.

For this reason, NR-06 treats the p-values as supplementary descriptive evidence.

The principal Stage 08 result is better characterized by:

  • +2.88 percentage points mean recall;
  • 10 better / 4 tied / 0 worse;
  • approximately 24% relative reduction in false suppression.

These quantities describe the observed retrospective comparison without making stronger population-level causal claims.

18. Human Review

Human correction is architecturally important because explicit correction can become a strong future learning signal.

The present corpus cannot evaluate that proposition formally.

No case preserves a reliable paired pre-review model state suitable for comparison with the final governed evidence order.

Separately, review of the historical runs identified no Stage 07 evidence selections requiring correction.

NR-06 therefore does not infer missing correction events and does not estimate a correction or override rate.

The appropriate conclusion is:

Human correction remains an architecturally valid future learning signal, but its effect was not measurable in this historical study.

Prospective instrumentation should preserve original model ordering separately from final governed ordering if that question is to be tested.

19. Why Downstream Outcomes Still Do Not Become Truth

The same governance problem persists beyond internal EPS stages.

An interview, offer, rejection, or hiring outcome may provide useful market-response information.

It does not establish that a specific evidence choice caused that outcome.

Observed outcomes do not automatically establish what would have happened under a different decision. Research on recommendation exposure and selective labels makes this counterfactual problem explicit [6,8].

External outcomes contain uncontrolled influences including:

  • applicant competition;
  • internal candidates;
  • timing;
  • compensation;
  • recruiter interpretation;
  • organizational changes;
  • hiring-manager preference;
  • factors not observable within EPS.

An interview is therefore evidence about market response.

It is not proof that a particular evidence record caused success.

A rejection is not proof that the candidate lacked the capability represented by the evidence.

NR-06 does not use hiring outcomes as labels and makes no claim that Conditions B or C predict external success.

20. The Governed Learning Authority Ladder

The architecture and findings suggest an explicit hierarchy.

Authoritative source evidence [9,10]

What the governed evidence base establishes about the candidate and opportunity.

Current-case governed reasoning [10]

What the present opportunity analysis establishes about relevance, mandate, and proof requirements.

Historical observations

What was considered, selected, promoted, or retained previously.

Deterministic historical diagnostics

Exposure, recurrence, reuse, survival, and comparable-context patterns.

Historical priors

Advisory estimates derived from previous observations.

Predictive learning

Model-generated expectations about future ranking, survival, or value.

Influence may increase as evidence accumulates.

Authority does not automatically increase with it.

Vertical authority ladder diagram showing phases 1 and 2 recording and diagnosing governed history, phase 3 introducing advisory historical priors after complete-corpus reasoning, and phase 4 as a separate authority boundary for predictive learning requiring independent evidence and activation authority.
Figure 4. The Governed Learning Authority Boundary. Phases 1 and 2 record and diagnose governed history. Phase 3 may introduce advisory historical priors after complete-corpus reasoning. Phase 4 would introduce predictive learning and requires separate evidence, evaluation, and activation authority. Historical recurrence never becomes factual authority merely by moving through the learning stack.

21. Implications for Phase 3

The study provides empirical support for continued Phase 3 development.

Historical information contained measurable downstream signal.

A governed transformation of that history behaved differently from raw recurrence.

Specifically, the fixed Condition C prior retrospectively recovered more evidence that subsequently appeared in promoted Stage 08 output while preserving the complete evidence universe.

This supports the bounded Phase 3 proposition:

Historical priors may be useful as advisory signals after complete-corpus reasoning, provided they remain exposure-aware, context-qualified, provenance-bound, inspectable, and incapable of excluding evidence or changing factual authority.

It does not support production activation from this study alone.

A prospective Phase 3 evaluation should preserve:

  • authentic current-case baseline rankings;
  • the historical-prior projection separately;
  • exact opportunity context;
  • current-model response to the prior;
  • final governed evidence selection;
  • downstream survival;
  • false suppression;
  • novel evidence discovery.

That would allow the actual effect of historical advice on current reasoning to be measured.

22. The Explicit Phase 4 Boundary

Phase 4 raises a different class of question.

An advisory prior says:

Historical experience suggests that this evidence may deserve attention.

A learned ranking system says:

The model predicts that this evidence should receive this relative position.

An iteration-policy model may go further:

Another computational intervention is expected to produce sufficient value.

These are predictive claims.

NR-06 does not authorize them.

The study did not establish:

  • superiority over an authentic no-history baseline;
  • causal benefit from historical influence;
  • precise mandate transfer;
  • generalization across a broad opportunity population;
  • candidate-facing quality improvement;
  • hiring-outcome improvement;
  • learned ranking reliability;
  • learned stopping authority.

Phase 4 therefore remains proposed.

Treating predictive learning as a separately governed activation step is consistent with broader lifecycle approaches to AI risk management. NIST's AI RMF addresses trustworthiness considerations across the design, development, use and evaluation of AI systems [2], while ISO/IEC 42001 establishes requirements for maintaining and continually improving an AI management system [3]. The specific activation requirements below are NR-06's proposed controls; they are not presented as requirements copied from either framework.

Any move toward predictive ranking should require:

  • a stable feature contract;
  • sufficient historical diversity;
  • exposure correction;
  • frozen baseline evaluation;
  • leakage protection;
  • shadow testing;
  • confidence calibration;
  • false-suppression monitoring;
  • explainability;
  • correction;
  • rollback;
  • drift monitoring.

A model should not be activated merely because enough data exists to train it.

23. Relationship to NR-05

NR-05 examined a related authority problem from the perspective of iterative optimization [12].

A later optimization round can produce:

  • another score;
  • another critique;
  • another candidate artifact.

Those observations do not automatically establish material improvement.

NR-05 therefore distinguished:

score movement ≠ material improvement ≠ authority to continue

NR-06 makes the analogous distinction for historical learning:

historical recurrence ≠ current relevance ≠ authority to select

The papers address complementary dimensions of adaptive AI systems.

NR-04 examined how the optimization process itself should be governed — establishing that the system, rather than the model's own assessment, must determine whether continued iteration is warranted [11].

NR-05 asks:

When should the system stop changing its current answer?

NR-06 asks:

When should previous experience influence the system's next answer?

Both require an explicit boundary between measurement and authority.

This becomes particularly important for future Phase 4 iteration-policy learning.

NR-05 showed that score movement alone is an inadequate continuation target. A future learned iteration policy would therefore need to account for material repair, evidence availability, regression, best-valid retention, and diminishing expected value rather than merely learning historical score trajectories.

A weak historical label does not become a strong objective simply because a model learns it accurately.

24. What the Evidence Supports

NR-06 supports six principal conclusions.

24.1 Historical experience contains reusable information

Earlier governed selections and downstream survival contain information associated with later governed use.

24.2 Exposure is a legitimate measurement concern

Raw recurrence and exposure-normalized history are not equivalent, although the observed effect on aggregate ranking was limited in this population.

24.3 Downstream survival is distinct from initial selection

Condition C did not improve recovery of Stage 07 Active selection but did improve retrospective recovery of promoted Stage 08 use and Stage 09 section survival.

24.4 Governance changes historical-prior behaviour

The governed prior reduced false suppression of later-promoted evidence relative to raw historical frequency.

24.5 Historical advice need not become historical exclusion

The observed benefit did not require historical information to remove evidence from the complete current corpus.

24.6 Phase 3 is empirically plausible but not yet production-validated

The study supports further prospective evaluation of governed priors, not automatic adoption.

25. What the Evidence Does Not Support

NR-06 does not establish that:

  • EPS has learned a production ranking policy;
  • historical learning improves current-case model reasoning;
  • Condition C outperforms an authentic no-history baseline;
  • exposure normalization alone caused the result;
  • precise mandate matching improves evidence selection;
  • downstream survival equals evidence quality;
  • historical selection equals effectiveness;
  • human correction reduced bias;
  • the prior improves candidate-facing writing;
  • the prior improves interviews, offers, or hiring outcomes;
  • the findings generalize automatically beyond the observed EPS population;
  • either historical prior should be activated as production authority.

The strongest supported empirical statement is:

The fixed governed prior retrospectively recovered more historically promoted Stage 08 evidence than raw frequency while preserving the complete evidence universe.

26. Limitations

The findings are subject to several material limitations.

Sample size. Sixteen cases supported the historical population analysis and 14 the chronology-locked comparison.

Non-independence. Cases share evidence, job-family characteristics, and accumulated system history.

Missing Condition A. No authentic no-history current-case baseline was preserved.

Coarse context. Context qualification was limited; no target had two earlier exact mandate-pattern matches.

Matched-K dependence. The strongest result occurred at a relatively broad matched-K cutoff averaging approximately 70 records.

Trace granularity. Stage 09 survival is section-level rather than exact statement-level.

Stage 10 ceiling. Both historical priors already recovered most Stage 10 argument-use evidence.

No model-human paired history. Human override effects could not be quantified.

No causal external outcome. The research does not link evidence ranking causally to interview, offer, or hiring outcomes.

These limitations define the boundaries of the research claim rather than serving as secondary qualifications.

27. Implications Beyond EPS

The problem examined here applies to any AI system that learns from its own previous behaviour.

Related feedback effects have been demonstrated in other machine-learning settings. Recommender systems can amplify popularity bias through repeated recommendation-and-interaction cycles [4], and model-driven data-feedback systems can amplify existing biases when interactions with earlier models become future training data [5]. These systems are not equivalent to EPS, but they establish that a system's prior behaviour can alter the information environment from which it subsequently learns.

A clinical-support system may increasingly favour tests it previously recommended.

A fraud system may inspect transaction patterns it already classified as suspicious.

A knowledge system may retrieve sources similar to those it cited before.

An autonomous software system may repeatedly choose implementation patterns associated with previous internal success.

Each creates the possibility of circular evidence:

The system preferred this before, therefore the system becomes increasingly likely to prefer it again.

Governed learning requires different questions:

  • Was the item equally exposed?
  • Was the previous context comparable?
  • Was the historical observation selection, survival, or external outcome?
  • Is the signal observational or causal?
  • Can the current system still see alternatives?
  • Can current evidence legitimately disagree with history?
  • What authority has actually been granted to the learned signal?

The engineering problem is therefore not merely adding memory.

It is controlling the transformation of memory into influence.

28. A General Governed-Learning Pattern

The combined architecture and evidence suggest a general sequence.

The sequence below is a synthesis derived from the EPS architecture and NR-06 findings. It is not a restatement of an existing standard.

Observe → preserve provenance → classify authority → normalize exposure → contextualize → advise → independently reason → validate → promote

This differs fundamentally from:

Observe → learn → believe

The first sequence treats learning as a governed transformation.

The second risks converting repeated system behaviour into self-confirming truth.

A trustworthy learning architecture should preserve distinctions between:

  • historical event and factual evidence;
  • frequency and exposure-normalized history;
  • selection and survival;
  • survival and effectiveness;
  • prediction and authority;
  • advice and decision;
  • current reasoning and historical influence.

These separations do not prevent learning.

They make learning inspectable.

29. Next Research Step

The next evaluation should be prospective.

Future EPS runs should preserve:

  1. the complete current evidence universe;
  2. current-case reasoning before historical influence;
  3. authentic pre-review model ranking;
  4. the historical-prior projection;
  5. model response after receiving that projection;
  6. final governed evidence selection;
  7. human review;
  8. Stage 08 promoted use;
  9. Stage 09 survival;
  10. Stage 10 argument use;
  11. context, model, policy, and prior versions.

That would permit a true three-condition evaluation:

A — current reasoning without history

B — current reasoning with naïve history

C — current reasoning with governed history

False suppression should remain a primary safety outcome.

Novel evidence discovery should be measured explicitly.

A historical learner that improves average agreement while becoming less capable of finding previously uncommon but currently important evidence would be learning the wrong lesson.

Conclusion

Learning Without Manufacturing Truth

AI systems accumulate experience quickly.

Authority should accumulate much more slowly.

The historical EPS population examined in NR-06 contained 1,893 opportunity × evidence-record events across 120 unique evidence records. Fourteen later cases supported a chronology-locked comparison of a naïve historical prior and a governed historical prior.

At the preregistered matched-K cutoff, the governed prior recovered more evidence subsequently used in promoted Stage 08 output:

88.18% → 91.06%.

It was better in 10 cases, tied in four, and worse in none.

Mean false suppression declined:

11.82% → 8.94%, an approximately 24% relative reduction.

A similar improvement appeared in Stage 09 section survival.

No corresponding improvement appeared in the initial Stage 07 Active selection.

The study therefore does not show that history knows what the current system should select.

It shows that governed historical information contains additional signal about what has previously survived downstream—and that this signal can influence prioritization without being permitted to define factual relevance or suppress the current evidence universe.

That finding supports further prospective evaluation of Phase 3 advisory historical priors.

It does not authorize Phase 4 predictive ranking or learned decision authority.

The broader lesson is not that AI systems should learn more aggressively from their own experience.

It is that experience should pass through governance before it is allowed to influence the future.

Repetition is not truth.

Selection is not effectiveness.

Survival is not causation.

Prediction is not authority.

Learning becomes trustworthy not when a system remembers more, but when it preserves the difference between what happened, what it inferred, and what it is entitled to believe.

Research Status and Evidence Basis

NR-06 distinguishes architecture evidence from empirical research evidence.

Architecture authority

EPS_ARCHITECTURE.md
Primary living EPS architecture authority. Establishes stage authority, governed artifacts, source boundaries, promotion, recovery, and the learning phase boundary.

EPS Governed V1.0 Architecture
Implemented governed workflow baseline. Documents artifact authority, lineage, deterministic/model responsibilities, human approval gates, and promotion behaviour.

EPS Knowledge Graph and Selection-Learning Flow
Central learning architecture. Phase 1A and 1B are implemented. Establishes case-qualified selection history, the complete-corpus-before-history rule, advisory-only historical learning, and proposed Phase 3/4 progression.

EPS Phase 2 Graph Diagnostics Implementation Specification
Current Phase 2 implementation authority. Defines deterministic proof coverage, reuse, evidence availability, selection history, and traceability diagnostics. Explicitly prohibits Phase 2 from learning predictions or changing evidence authority.

Stage 07–10 Contracts and Protected Enhancements Contract
Normative stage-level authority governing evidence selection, Stage 08 bounded repair, downstream promotion, traceability, and best-valid retention.

Empirical evidence

The historical population analysis was produced from the NR-06 governed-learning evidence package dated 4 October 2026.

The population contained: 27 historical non-test EPS cases identified; 16 complete cases suitable for principal historical analysis; 1,893 opportunity × evidence-record events; 120 unique evidence records; 15 cases supporting Stage 08→09 survival analysis; 13 supporting Stage 10 argument-use analysis; 14 supporting chronology-locked B/C evaluation.

The chronology-locked comparison was performed under a fixed protocol established before comparison scores were calculated.

The principal empirical outputs include: NR-06 research-readiness assessment; case eligibility; historical event population; exposure analysis; context-transfer analysis; downstream-survival analysis; replay-feasibility assessment; B-versus-C locked protocol; case-level B/C results; paired statistical summary; complete rankings; replay audit; context results; reproducible deterministic analysis code.

No provider calls were used to generate the empirical results. No EPS production code, architecture document, runtime case, or historical source artifact was modified by the research analysis.

Evidence interpretation

The B/C comparison is retrospective and observational. Reported p-values are exploratory. The strongest result depends on the preregistered matched-K comparison. The evidence supports further study of governed Phase 3 historical priors. It does not constitute activation evidence for predictive Phase 4 learning.

References
[1] Lebo, T., Sahoo, S., and McGuinness, D., eds. "PROV-O: The PROV Ontology." W3C Recommendation, 30 April 2013.
[2] Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. National Institute of Standards and Technology, 2023. DOI: 10.6028/NIST.AI.100-1.
[3] ISO/IEC. ISO/IEC 42001:2023 — Information technology — Artificial intelligence — Management system. International Organization for Standardization, 2023.
[4] Mansoury, M., Abdollahpouri, H., Pechenizkiy, M., Mobasher, B., and Burke, R. "Feedback Loop and Bias Amplification in Recommender Systems." Proceedings of the 29th ACM International Conference on Information & Knowledge Management, pp. 2145–2148, 2020. DOI: 10.1145/3340531.3412152.
[5] Taori, R., and Hashimoto, T. "Data Feedback Loops: Model-driven Amplification of Dataset Biases." Proceedings of the 40th International Conference on Machine Learning, PMLR 202, pp. 33883–33920, 2023.
[6] Schnabel, T., Swaminathan, A., Singh, A., Chandak, N., and Joachims, T. "Recommendations as Treatments: Debiasing Learning and Evaluation." Proceedings of the 33rd International Conference on Machine Learning, PMLR 48, pp. 1670–1679, 2016.
[7] Saito, Y., Yaginuma, S., Nishino, Y., Sakata, H., and Nakata, K. "Unbiased Recommender Learning from Missing-Not-At-Random Implicit Feedback." Proceedings of the 13th International Conference on Web Search and Data Mining, pp. 501–509, 2020. DOI: 10.1145/3336191.3371783.
[8] Kleinberg, J., Lakkaraju, H., Leskovec, J., Ludwig, J., and Mullainathan, S. "Human Decisions and Machine Predictions." The Quarterly Journal of Economics, 133(1), pp. 237–293, 2018. DOI: 10.1093/qje/qjx032.
[9] Boyle, S. Artifact-Oriented AI Architecture: Governing State, Authority and Lineage. Northline Research Series: Governed AI Systems, NR-02. Northline Advisory, 2026.
[10] Boyle, S. Evidence Coverage and Traceability: Knowing What Supports a Conclusion. Northline Research Series: Governed AI Systems, NR-03. Northline Advisory, 2026.
[11] Boyle, S. Governed Self-Optimization: How an AI System Can Improve Its Work Without Losing Control. Northline Research Series: Governed AI Systems, NR-04. Northline Advisory, 2026.
[12] Boyle, S. When Should the AI Stop? Material Improvement and the Limits of Iterative Optimization. Northline Research Series: Governed AI Systems, NR-05. Northline Advisory, 2026.
Series Continuity

NR-03 examined evidence authority.

NR-04 examined repeated optimization.

NR-05 examined the authority to continue or stop optimization.

NR-06 examines the authority granted to historical experience.

What may the system treat as evidence?

How should the system improve an artifact?

When should it stop changing the artifact?

When should previous experience influence the next decision?

Next in the series
NR-07
Architecture as a Governing Contract
Engineering Reliable Software When the Implementer Is Probabilistic
Want a copy of this paper?

Receive a link to the complete paper by email.

Steven Boyle
Northline Advisory

Steven Boyle is the founder of Northline Advisory, a technology advisory and research practice focused on technology leadership, enterprise transformation, governance, data and governed AI. His work draws on more than two decades of executive and operational experience across higher education and public-interest organizations.

Related Insights
Research Paper · Data & AI

When Should the AI Stop?

Material Improvement and the Limits of Iterative Optimization

Historical and controlled evidence on material improvement, diminishing returns and the governance of stopping decisions in iterative AI optimization.

Steven BoyleNR-05October 2026 · 30 min readRead →
Research Paper · Data & AI

Evidence Coverage and Traceability

Knowing What Supports a Conclusion

A deterministic evidence architecture for preserving and inspecting the governed relationships that actually support AI-generated conclusions.

Steven BoyleNR-03October 2026 · 25 min readRead →
Research Paper · Data & AI

Governed Self-Optimization

How an AI System Can Improve Its Work Without Losing Control

A governed optimization loop that lets AI propose and evaluate revisions while deterministic controls retain authority over continuation and promotion.

Steven BoyleNR-04October 2026 · 28 min readRead →