← Back to Research
NORTHLINE RESEARCH SERIES · GOVERNED AI SYSTEMS · NR-05

When Should the AI Stop?

Material Improvement and the Limits of Iterative Optimization

Steven BoyleNorthline Advisory
|October 2026 · 30 min read
Abstract

Generative AI systems can increasingly evaluate and revise their own work. That capability creates an apparently straightforward optimization loop: generate an artifact, evaluate it, improve it, evaluate it again, and continue while quality increases. The difficulty is deciding what improvement actually means. In interpretive work, there may be no objectively optimal output. An evaluator can prefer different wording, emphasis, evidence use, structure or compression without establishing that the resulting artifact is materially better for its intended purpose, so repeated evaluation can continue producing optimization signals even after demonstrable material value has weakened. This paper examines that problem using historical operating evidence, controlled experiments and subsequent architectural changes within the Executive Positioning System (EPS). The historical analysis covered 84 runs and 623 Stage 08 optimization trajectories, including 399 adjacent round transitions sufficiently complete and interpretable for transition analysis. The frequency of explicitly demonstrable material gains declined from 34.4% at Round 1 to 2 to 18.0% at Round 2 to 3 and 13.5% at Round 3 to 4, while median evaluation-score movement declined from +6.0 to +2.5 to +2.0, and to +1.0 at Round 4 to 5. Twenty historical transitions showed divergence between score movement and stronger evidence of material value. Across nine controlled post-Round-1 optimization rounds, two produced material gains, four produced minor or sub-material gains, and three regressed. Separate blind candidate-facing assessment also demonstrated that a higher internal evaluation score did not necessarily correspond to the stronger human-facing artifact. The architectural response was not simply to lower the maximum number of rounds or establish a higher numerical stopping threshold. EPS changed the optimization obligation: material requirements identified by the initial evaluation are carried into a bounded repair cycle, the successor is expected to address those requirements, and subsequent evaluation verifies the repair while checking for regression and evidence-integrity failure.

Get this research paper

Receive the complete paper as a PDF by email.

1. The stopping problem is an improvement problem

Iterative AI optimization appears straightforward in its basic form.

Generate → Evaluate → Improve → Repeat

An artifact is produced. An evaluator examines it and identifies weaknesses. A revised version is generated that addresses those weaknesses. The process continues while quality is increasing.

The difficulty is that this description contains an embedded assumption: that the evaluation will eventually converge on a stable judgment about quality, that "better" can be unambiguously established, and that the improvement obligation is self-evidently complete when the evaluator has nothing more to propose.

None of those assumptions reliably holds in interpretive work.

An evaluator assessing an executive-positioning artifact does not measure against a unique correct answer. It assesses against criteria—evidence use, structural clarity, candidate-facing quality, claim integrity—that do not have a single objectively determined solution. Two evaluators can reach different judgments about the same artifact. The same evaluator can make different assessments in different contexts. Revision instructions can be satisfied while the artifact becomes less useful in ways the evaluator did not explicitly measure.

The central difficulty therefore is not merely computational: how many rounds is the right number? It is epistemic: how does the system determine whether a proposed change represents genuine improvement against the purpose the artifact is meant to serve?

Iterative feedback and revision have established precedent in LLM research. Self-Refine demonstrated that language models can use generated feedback to revise their outputs iteratively and improve performance across a range of tasks [1]. The question examined in NR-05 is not whether iterative refinement can work. It is what evidence should authorize another iteration once improvement becomes interpretive, uncertain and potentially non-monotonic.

This paper examines that problem through operating evidence, controlled experiments and the architectural response implemented in the Executive Positioning System (EPS). The stopping problem is not primarily a problem of limiting computation. It is determining what evidence gives the system authority to continue.

Two-panel diagram contrasting open-ended iterative optimization, where each evaluation produces a newly interpreted improvement opportunity, with bounded iterative optimization, where material repair obligations are identified, held stable, and verified against completion.
Figure 1. Open-Ended and Bounded Iterative Optimization. Open-ended optimization repeatedly evaluates each new artifact against a newly interpreted opportunity for improvement. Bounded optimization identifies material repair obligations, holds those obligations stable through the intervention, and evaluates whether the successor satisfactorily addresses them without material regression.

2. What improvement means in interpretive work

Improvement in interpretive AI work cannot be reduced to a scalar increase in a model-generated evaluation score.

A score captures one evaluator's judgment at one moment about one set of criteria. It may reflect genuine improvement in evidence use. It may reflect a stylistic preference. It may reflect a different interpretation of the same underlying material. It may reflect the model's tendency to reward length, variety or novelty even when those properties do not improve candidate-facing quality.

That distinction matters because an iterative optimizer has a strong incentive to maximize whatever metric the continuation policy treats as the quality signal. If the signal is evaluator score, the optimizer will produce candidates that score highly. Whether those candidates are better for their intended purpose is a separate question.

The risk that an optimization target diverges from the underlying objective has established precedent. Goodhart-type failures describe cases in which a metric useful for improving a system becomes less reliable as stronger optimization pressure is applied to it [2]. In machine learning, Gao, Schulman and Hilton showed experimentally that stronger optimization against an imperfect proxy reward can eventually reduce performance against a separate gold-standard objective [3].

The EPS problem is analogous but not identical. Framework score is not a reinforcement-learning reward function. The relevant parallel is that a measurable evaluator signal can be an incomplete proxy for a richer underlying objective—candidate-facing quality, evidence integrity and substantive usefulness—when movement in that signal is treated as sufficient evidence to continue.

The specific mechanism in iterative AI optimization is this: the evaluator is both the source of the improvement signal and a probabilistic system with its own interpretation of quality. When the writer optimizes against the evaluator's feedback, it learns the evaluator's expressed preferences, which may or may not align with what makes the artifact more useful to the person for whom it was prepared.

Model-based evaluation can be useful without being neutral. Zheng et al. found that strong LLM judges could achieve high agreement with human preferences while still exhibiting identifiable position, verbosity and self-enhancement biases, along with limitations in reasoning [4]. These findings do not establish the specific biases present in EPS's evaluator; they reinforce the need to treat model-generated evaluation as evidence rather than ground truth.

This creates the possibility of an optimization loop that produces artifacts that are increasingly well-tailored to the evaluator's implicit model of quality—and decreasingly grounded in the evidence, reasoning and candidate-specific considerations that should be guiding the work.

Score improvement and material improvement are related but distinct. The ability to identify another change in evaluation is not by itself evidence that the change represents genuine material value.

3. Historical evidence of declining material gains

EPS maintains detailed production records that allowed a retrospective analysis of optimization trajectory behaviour across a substantial corpus.

The analysis examined 84 unique historical runs and 623 Stage 08 optimization trajectories. Of these, 379 had sufficiently complete quantitative records for trajectory analysis. The 399 complete and interpretable adjacent round transitions supported the core transition analysis. Trajectory depth varied: 211 trajectories reached at least three rounds, 81 reached at least four, 39 reached at least five, and 16 reached at least six.

Historical evidence Count
Unique runs 84
Stage 08 trajectories 623
Strictly eligible quantitative trajectories 379
Complete, interpretable adjacent transitions 399
Trajectories reaching 3+ rounds 211
Trajectories reaching 4+ rounds 81
Trajectories reaching 5+ rounds 39
Trajectories reaching 6+ rounds 16

The analysis examined whether each transition provided explicit evidence of material improvement from the perspective of candidate-facing quality, evidence use or substantive structural gain—rather than movement in evaluation score alone.

The frequency of explicitly demonstrable material gains declined across successive transitions. At R1→R2, 34.4% of transitions showed explicit evidence of material gain. At R2→R3, the rate fell to 18.0%. At R3→R4 it fell further to 13.5%. R4→R5 showed 20.0% across a smaller sample of 15 transitions. R5→R6+ showed 0% across a very small sample insufficient for stable inference.

Transition Interpretable transitions Explicit material-gain rate Median score movement
R1→R2 215 34.4% +6.0
R2→R3 128 18.0% +2.5
R3→R4 37 13.5% +2.0
R4→R5 15 20.0% +1.0
R5→R6+ very small sample 0% insufficient for stable inference

Median evaluation-score movement also declined: +6.0 at R1→R2, +2.5 at R2→R3, and +2.0 at R3→R4, declining further to +1.0 at R4→R5.

These figures describe historical trajectory behaviour across the analysed corpus. They do not establish a universal stopping threshold, and the later-round samples are small enough that those figures should be treated as descriptive observations rather than stable estimates.

Twenty historical transitions showed a divergence pattern: score movement present without strong evidence of material value. This is the core pattern the evidence basis addresses—continued score improvement without corresponding evidence of continued material gain.

Two-panel chart showing the frequency of explicitly demonstrable material gains across successive optimization rounds falling from 34.4% at R1 to R2 down to 13.5% at R3 to R4, alongside a corresponding decline in median evaluation-score movement from plus 6.0 to plus 2.0 across the same transitions.
Figure 2. Material Gain and Score Movement Across Successive Optimization Rounds. Historical EPS trajectories show the frequency of explicitly demonstrable material gains falling from 34.4% at R1→R2 to 18.0% at R2→R3 and 13.5% at R3→R4, while median evaluation-score movement declines from +6.0 to +2.5 and +2.0. Later-round samples are smaller and are shown as descriptive observations rather than stable estimates.

4. The indeterminacy problem in historical evidence

A significant limitation of the historical analysis is that material-value classification required independent evidence beyond score movement. The classification used in this analysis asked: can the transition be assessed for material candidate-facing value gain, based on available evidence, reasoning and documented outcomes?

Of the 399 interpretable transitions, 273 were classified as indeterminate for candidate-facing materiality. This is a high proportion—more than two-thirds of the corpus—and it reflects a genuine limitation in the historical record. Many trajectories do not have sufficient independent evidence to distinguish whether score improvement corresponded to genuine material value or to evaluator preference, wording variation, or other non-material factors.

This limitation is important for two reasons.

First, it means the explicit material-gain rates reported in Section 3 represent lower bounds on actual improvement. Some transitions classified as indeterminate may have produced genuine material value that the historical record cannot establish.

Second, it identifies a structural problem with relying on historical trajectory data alone as evidence about stopping. If the system cannot distinguish material value from score movement in most historical transitions, then historical trajectory data cannot establish when continuation was genuinely warranted—only when it was or was not clearly valuable.

This motivated the controlled experimental approach described in Section 5.

This limitation is consistent with the broader proxy-evaluation problem: movement in an optimization signal does not by itself establish that the underlying objective improved [2,3]. NR-05 therefore does not infer material improvement from score movement where independent evidence is unavailable.

5. Controlled experimental evidence

Controlled Stage 08 experiments allowed a stronger quality signal where historical instrumentation was insufficient.

The distinction between an optimized proxy and an independent quality signal has precedent in reward-model overoptimization research, where Gao et al. evaluate optimization against a proxy using a separate gold-standard reward model [3]. NR-05 uses a different domain and experimental design but applies the same underlying evidentiary principle: the optimization signal cannot validate itself.

Experiment 1 examined nine post-Round-1 optimization rounds with controlled starting conditions, detailed pre/post documentation, and subsequent blind candidate-facing assessment independent of Framework scoring.

Of the nine rounds:

Outcome Rounds
Material gain 2
Minor or sub-material gain 4
Regression 3

The experiment also demonstrated two findings with direct implications for stopping policy.

First, regression occurred in three of nine rounds. A later round is not guaranteed to improve on its predecessor. Best-valid retention is therefore not an exception-handling feature—it is a necessary protection against normal probabilistic outcome variance.

Second, blind candidate-facing assessment and Framework score diverged in multiple conditions. A higher internal evaluation score did not necessarily correspond to the stronger human-facing artifact. Earlier research consequently rejected Framework score as sufficient quality truth.

These findings reinforced what the historical evidence suggested: score movement is not material-improvement authority, and continuation justified by score movement alone does not have adequate epistemic foundation.

6. Score movement is not material-improvement authority

The combined evidence supports a clear principle.

Evaluation score is a useful signal. It captures an evaluator's assessment of the current artifact against its criteria. It can identify weakness. It can flag regression. It can distinguish strong candidates from weak ones at a coarse level.

What it cannot do is establish that a positive score delta represents genuine material improvement against the artifact's actual purpose.

That distinction matters because the iterative optimizer's incentive structure pushes toward score maximization. If score is the continuation signal, the system will continue until scores plateau or limits are reached—regardless of whether material quality is actually improving.

The historical evidence showed this pattern directly: 20 transitions exhibited score movement without strong evidence of material value. The controlled experiments confirmed it: blind human assessment and Framework score diverged repeatedly.

Material-improvement authority therefore requires more than score movement. It requires:

  • Evidence that the round addressed a specific material defect;
  • Evidence that the defect resolution improved candidate-facing quality against the artifact's purpose;
  • Absence of material regression on dimensions not targeted by the repair;
  • Absence of evidence-integrity failure in the successor.

This converts score from the terminal decision criterion into one input among several.

A score can inform an optimization decision without being allowed to become the decision.

Existing research establishes both the usefulness of iterative refinement [1] and the risk that increasingly strong optimization against an imperfect objective can separate measured improvement from the intended outcome [2,3,5]. The bounded-obligation architecture below is the EPS response to that tension. It is a Northline architectural synthesis rather than a restatement of an established stopping algorithm.

7. The structure of a material obligation

The architectural response in EPS did not lower the maximum round count or establish a higher numerical stopping threshold. It changed the nature of the optimization obligation.

Instead of asking the evaluator to identify opportunities for improvement, the redesigned architecture asks the evaluator to identify material repair obligations—specific, bounded claims about what the artifact fails to do adequately against its governing criteria.

The distinction is this: an improvement opportunity is open-ended. It invites the writer to find the best way to make the artifact better, given the evaluator's interpretation of quality. A repair obligation is bounded. It specifies what must be fixed, what evidence should support the fix, and what constitutes satisfactory resolution.

Carrying obligations forward rather than re-evaluating freely has a significant consequence.

It separates the evaluation function from the creative function. The evaluator is no longer continuously reinterpreting quality on each round—finding new things to change based on fresh examination of each successive candidate. Instead, it assesses whether the successor has addressed the obligations identified in the prior round, and whether it has introduced new material problems.

This also addresses one of the structural vulnerabilities of open-ended optimization: the tendency for evaluations to identify different weaknesses on successive rounds rather than confirming improvement on the same underlying problem. Carrying obligations forward creates a stable basis for comparison.

A successor that scores higher is not necessarily better. A successor that satisfies its bounded material obligations without regression is better in a governed sense.

8. What a bounded repair obligation looks like

A material repair obligation in the EPS architecture is not a vague instruction to improve. It is a typed, bounded specification carrying several properties.

It identifies: the specific deficiency the successor must address; the evidence available to support the repair; the quality dimension being repaired; what constitutes satisfactory resolution; what existing valid content must be preserved; and what regression conditions would override the improvement.

This creates a governed contract for the successor round. The writer is not asked to reinterpret the entire artifact and improve it by its best lights. It is asked to satisfy specific repair obligations against the available evidence while preserving existing valid content.

The evaluator then does not perform an unconstrained fresh assessment of the successor. It verifies whether the specified obligations have been adequately addressed, whether regression has occurred on preserved content, and whether evidence integrity has been maintained.

This is a narrower scope than traditional iterative evaluation, deliberately so.

The narrower scope reduces the surface for new interpretation to enter. It makes the continuation decision less dependent on whether a fresh evaluator finds something to improve and more dependent on whether the bounded work has been completed satisfactorily.

It also creates an observable stopping condition: when material repair obligations have been satisfied and no new material obligations have been introduced, the bounded optimization is complete.

9. The role of evidence in the stopping decision

Material repair obligations must be resolvable against available evidence.

An obligation that cannot be satisfied from the governed evidence base is not a repair opportunity—it is an evidence ceiling. Asking the writer to satisfy an unsatisfiable obligation creates pressure toward unsupported inference, invented specificity, or increasingly polished restatement of insufficient evidence.

The evidence-coverage and diagnostics architecture documented in NR-03 [7] therefore becomes part of the stopping decision.

Before authorizing continuation against a repair obligation, the system should assess whether the governed evidence base contains what would be needed to satisfy it. An obligation supported by available governed evidence is a legitimate continuation reason. An obligation for which no supporting evidence exists is an evidence ceiling: the system should acknowledge the limitation rather than continue optimizing against it.

This prevents the optimization loop from consuming additional computation on work it cannot legitimately complete.

Material obligation + available governed evidence → continuation authorized
Material obligation + evidence ceiling → limitation acknowledged, continuation stopped

The stopping criterion is therefore not just whether obligations remain—it is whether satisfiable obligations remain.

10. Regression as a stopping signal

The controlled experiments showed regression in three of nine rounds.

Regression is not an exceptional failure condition. It is a normal property of probabilistic optimization. A successor round does not automatically improve on its predecessor.

Iterative-refinement research demonstrates that additional revision can improve model outputs [1], but does not establish that every successor will be better in every task. The EPS evidence therefore treats improvement as something to be established at the candidate level rather than assumed from the existence of another refinement round.

A successor may improve some dimensions while weakening others. It may satisfy the targeted obligation while introducing a new material problem. It may produce a higher aggregate score while making the artifact less faithful to the evidence.

The stopping architecture must account for this.

Best-valid retention—preserving the strongest valid prior candidate independently of the latest—is a prerequisite for rational stopping [6,8]. Without it, the system cannot stop on the best result. It can only stop on the latest result, which may be worse.

But regression is also a stopping signal in a direct sense.

If the successor introduces material regression on previously valid content while failing to adequately address its targeted obligations, continuation has produced a worse outcome. That is not a reason to try again immediately. It is evidence about the difficulty of the repair, the adequacy of the available evidence, and the reliability of the optimization process at this point in the trajectory.

A regression on the successor round is evidence that should affect the stopping decision—not just the retention decision.

11. Evidence integrity as a stopping criterion

The experiments also demonstrated that a candidate can score highly while introducing evidence-integrity problems: unsupported claims, overclaimed outcomes, or evidence use that implies stronger proof than the governed evidence base supports.

An artifact with a high Framework score and an evidence-integrity failure is not a successful optimization outcome. It may be better on the evaluator's stylistic criteria while being less trustworthy as a document.

A high optimization score is not sufficient when the underlying objective has been compromised. AI-safety research describes related failures in which optimization of the specified objective produces behaviour inconsistent with the intended outcome, including reward hacking [5]. NR-05 identifies a narrower version of that problem: an artifact can improve against evaluator criteria while weakening factual support or claim integrity.

Evidence integrity therefore cannot be subordinated to aggregate score in the stopping decision. A successor that fails evidence-integrity checks should not be retained as the best-valid candidate, regardless of its score. And if a trajectory produces repeated integrity failures under continuation, that is an indication that the optimization loop is producing superficial improvements at the cost of reliability.

This is consistent with the governing principle developed in NR-03: a trustworthy system must represent what it cannot establish. A system that optimizes against a score metric while degrading its evidence integrity is failing at the underlying objective even when the metric improves.

12. The human decision boundary

Bounded material obligations change where human judgment is required in the optimization process.

In an open-ended optimization loop, there is no natural stopping point at which the evaluator's judgment is exhausted. The evaluator can always find something to improve. Human intervention would be required either to review every proposed continuation or to establish an arbitrary limit.

In a bounded obligation architecture, completion has a natural definition: material obligations have been addressed without regression. This creates a machine-observable stopping condition that does not require the human to assess each candidate's quality directly. The separation of generation, evaluation and authority established in NR-04 [8] is what makes this possible: the human governs the scope and accepts the outcome, without needing to supervise each intermediate round.

Human judgment is still essential, but at different points:

  • Authorizing the optimization scope and parameters;
  • Reviewing the initial material obligations and their evidence basis;
  • Assessing completion when obligations are satisfied—or determining that the outcome is adequate when obligations cannot be fully satisfied;
  • Making promotion decisions when the best-valid candidate is ready for use.

This is a different governance model than reviewing every round, and a more tractable one at scale.

Bounded obligations give humans a governed handover point rather than a requirement to monitor continuous iteration.

13. Hard limits as circuit breakers, not stopping criteria

Maximum round limits remain necessary in any bounded autonomous system. They prevent runaway iteration in failure conditions, protect against edge cases where the obligation architecture produces an unexpected result, and establish a firm computational boundary.

But a hard round limit is not a stopping criterion in the material sense. It is a circuit breaker.

When a trajectory terminates because it has exhausted its permitted rounds, the system has not established that optimization is complete. It has established that the computational allowance is exhausted. Those are different outcomes, and they should be represented differently.

A trajectory that terminates at the maximum round with unresolved material obligations should be reported differently from one that terminates because obligations were satisfied. A trajectory that terminates because an evidence ceiling was reached should be reported differently from one that was stopped by the hard limit because the obligation architecture failed to converge.

These distinctions matter for trust. A user or governing authority who sees only "optimization complete" cannot assess whether the result is adequate. A system that distinguishes "obligations satisfied" from "limit reached with unresolved obligations" provides a more honest basis for promotion decisions.

A safety ceiling can stop computation. It cannot prove that optimization is complete.

Decision diagram showing the governed stopping decision in which continuation is authorized by unresolved material obligations rather than score movement, with best-valid retention, hard round limits as circuit breakers, and explicit terminal states distinguishing obligation-satisfied from limit-reached.
Figure 4. The Governed Stopping Decision. Another optimization round is authorized by unresolved material obligations, not merely by the model's ability to propose another change or produce a higher score. Repair proceeds against bounded requirements; verification tests completion and regression; best-valid retention protects earlier work; hard round limits remain circuit breakers rather than evidence of convergence.

14. Terminal state transparency

A governed stopping architecture requires explicit terminal states that distinguish how a trajectory ended.

The following terminal states are an architectural synthesis derived from the EPS evidence and bounded-optimization implementation. They are not presented as an established taxonomy from prior literature.

Trajectories can end for several reasons, which should not be collapsed into a single completion status:

Obligations satisfied. The successor addressed its material repair obligations without material regression and without evidence-integrity failure. Optimization is complete in a governed sense.

Evidence ceiling reached. A material obligation cannot be satisfied from the available governed evidence base. The system acknowledges the limitation rather than continuing to optimize against it.

Regression without recovery. Continuation produced regression that the optimization process could not recover. The best-valid state from before the regression is retained.

Hard limit reached. The maximum permitted rounds were exhausted before obligations were fully satisfied.

Integrity halt. Evidence-integrity failure conditions were triggered in a manner that makes continuation inappropriate.

These terminal states are important for the promotion decision. A candidate reaching the "obligations satisfied" state has stronger grounds for promotion than one reaching the "hard limit reached" state with residual obligations. A candidate whose terminal state is "evidence ceiling reached" may be the best the governed evidence can support—which is an honest and appropriate result—but it is different from one where obligations were fully resolved.

Representing these states explicitly enables human reviewers to make promotion decisions with an accurate understanding of what the optimization process established and what it did not.

15. Relationship to governed self-optimization

This paper extends the architecture developed in NR-04 rather than revising it [8].

NR-04 established the core principles: generation and evaluation are separate; continuation requires material and supportable reason; best-valid retention protects against regression; the optimizer cannot control what becomes authoritative. These remain valid.

What this paper adds is a more specific architecture for the stopping decision itself: not just that continuation requires material reason, but what constitutes material reason, how obligations are bounded and carried forward, how completion is defined, and how the terminal state should be reported.

The transition from NR-04 to NR-05 is the transition from a governed optimization loop to a governed stopping policy. The loop architecture establishes the framework. The stopping policy establishes when the loop should terminate and on what basis.

Both are necessary. A governed loop without a governed stopping policy will continue until it hits a hard limit. A governed stopping policy without a governed loop will not have reliable candidates to stop on.

The artifact authority model established in NR-02 [6] remains the foundation: only a promoted artifact is authoritative, and the best-valid state is preserved independently of the latest round result. The evidence-coverage architecture established in NR-03 [7] provides the evidence ceiling that makes some obligations unresolvable regardless of further rounds.

Together they constitute the bounded optimization model: generate and evaluate under governance, continue only for material and satisfiable reasons, retain the best valid state, and stop when bounded obligations have been addressed—or honestly represent why they could not be.

16. What this study establishes—and what it does not

The historical evidence covers a substantial corpus of Stage 08 optimization trajectories across real EPS runs. The declining material-gain rates and score-movement patterns are genuine observations from that corpus. They are not invented, extrapolated or modelled.

The controlled experiments provide a stronger causal signal where the historical record could not distinguish material value from score movement. The experimental findings—two material gains, four minor or sub-material gains, three regressions across nine rounds—are real outcomes from executed experiments, not projected estimates.

What this study does not establish:

The 34.4%, 18.0% and 13.5% material-gain rates are not universal rates applicable to all AI optimization systems. They reflect this corpus, this evaluation framework, this interpretive domain and this definition of material gain.

This paper does not establish that R1→R2 is the universally correct stopping point, or that any specific numerical threshold determines when optimization should stop. The evidence supports declining material-gain rates across successive rounds, not a universal optimal round count.

The historical analysis does not establish what fraction of indeterminate transitions represented genuine material improvement. The 273 indeterminate transitions may contain real material gains the record cannot confirm—or they may predominantly represent score movement without material value. Both interpretations are consistent with the evidence.

The controlled experiments used EPS, a specific system in a specific interpretive domain. Their findings should not be generalized to all generative AI systems or all optimization tasks.

The bounded obligation architecture is the system's current response to these findings. It is supported by implemented contracts, tests and the experimental evidence. It is not established as the universally optimal stopping mechanism across all AI optimization contexts.

Conclusion

Better is not a stopping criterion

The central finding of this study is that evaluative improvement signals—score movement, evaluator preference, identified opportunities for change—do not by themselves provide adequate authority for continued optimization in interpretive AI work.

The historical evidence showed declining rates of explicitly demonstrable material gain across successive optimization rounds. The controlled experiments showed that score movement and material value diverged in practice, and that regression occurred in roughly one-third of post-Round-1 optimization attempts.

Together, these findings reject the assumption that an AI system's ability to identify another improvement is evidence that another improvement is warranted.

The architectural response was to redesign the optimization obligation. Rather than asking the evaluator to find the best way to improve a successor, the architecture requires the evaluator to identify specific bounded repair obligations, carry those obligations forward into the successor, and evaluate whether the successor has addressed them—not whether it has maximally improved on all dimensions.

This changes what the system is doing when it optimizes. It is not trying to maximize an open-ended evaluator score across unconstrained attempts. It is trying to satisfy a bounded definition of the work that remains.

When that bounded work is complete, the system has a principled basis for stopping: the obligations have been addressed, regression has been checked, and the best valid state has been retained.

That basis is grounded in evidence rather than preference. It is observable rather than assumed. And it gives the human decision-maker an accurate representation of what the optimization established—not merely a score and a round count.

The contribution is therefore not a claim that iterative refinement fails, nor that optimization metrics are inherently unusable. Prior research establishes that iterative refinement can improve LLM outputs [1] and that strong optimization against imperfect proxies can diverge from underlying objectives [2,3,5]. NR-05 contributes a governed stopping architecture for deciding when additional interpretive optimization is actually warranted.

In interpretive AI work, the ability to identify another change is not sufficient evidence that another optimization round is warranted. Governed optimization requires a bounded definition of material improvement, a stable repair obligation against which intervention can be evaluated, and an explicit basis for stopping once that obligation has been satisfied without material regression.

Research Status and Evidence Basis

This paper combines historical operating evidence, controlled experiments, implemented architecture and implementation tests from the Executive Positioning System.

The historical retrospective analysis examined:

  • 84 unique historical runs;
  • 623 Stage 08 optimization trajectories;
  • 399 complete and interpretable adjacent round transitions;
  • 211 trajectories reaching at least three rounds;
  • 81 reaching at least four rounds;
  • 39 reaching at least five rounds;
  • 16 reaching at least six rounds.

The historical evidence is strongest for optimization trajectory, score movement and recorded material findings. It is weaker for retrospective candidate-facing materiality because 273 of the 399 interpretable transitions lack sufficient independent evidence for confident material-value classification.

Controlled Stage 08 experiments provide stronger comparative evidence. Experiment 1 found two material gains, four minor or sub-material gains and three regressions across nine post-Round-1 rounds. It also demonstrated that later rounds can regress and that Framework score should not be treated as independent quality ground truth.

Experiment 2 further examined feedback sensitivity and reinforced the distinction between evaluator preference, Framework score and candidate-facing quality. Earlier research consequently rejected style, wording or evaluator preference alone as sufficient justification for continuation and retained best-valid preservation as a core quality control.

The current bounded stopping mechanism is supported by implemented Stage 08 contracts and tests governing material-repair requirements, successor verification, evidence integrity, terminal budgets and best-valid retention.

The study is based on one implemented AI system operating in a specific interpretive domain. The findings should therefore be understood as empirical observations and architectural conclusions from that system, not as universal experimental proof of optimal stopping behaviour across all generative AI systems.

References
[1] Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., and Clark, P. "Self-Refine: Iterative Refinement with Self-Feedback." Advances in Neural Information Processing Systems 36, 2023.
[2] Manheim, D., and Garrabrant, S. "Categorizing Variants of Goodhart's Law." arXiv:1803.04585, 2018.
[3] Gao, L., Schulman, J., and Hilton, J. "Scaling Laws for Reward Model Overoptimization." Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research 202, pp. 10835–10866, 2023.
[4] Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." Advances in Neural Information Processing Systems 36, Datasets and Benchmarks Track, 2023.
[5] Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D. "Concrete Problems in AI Safety." arXiv:1606.06565, 2016.
[6] Boyle, S. Artifact-Oriented AI Architecture: Governing State, Authority and Lineage. Northline Research Series: Governed AI Systems, NR-02. Northline Advisory, 2026.
[7] Boyle, S. Evidence Coverage and Traceability: Knowing What Supports a Conclusion. Northline Research Series: Governed AI Systems, NR-03. Northline Advisory, 2026.
[8] Boyle, S. Governed Self-Optimization: How an AI System Can Improve Its Work Without Losing Control. Northline Research Series: Governed AI Systems, NR-04. Northline Advisory, 2026.
Next in the series
NR-06
Governed Learning
How AI Systems Can Learn Without Turning Experience Into Truth
Want a copy of this paper?

Receive a link to the complete paper by email.

Steven Boyle
Northline Advisory

Steven Boyle is the founder of Northline Advisory, a technology advisory and research practice focused on technology leadership, enterprise transformation, governance, data and governed AI. His work draws on more than two decades of executive and operational experience across higher education and public-interest organizations.

Related Insights
Research Paper · Data & AI

Governed Self-Optimization

How an AI System Can Improve Its Work Without Losing Control

A governed optimization loop that lets AI propose and evaluate revisions while deterministic controls retain authority over continuation and promotion.

Steven BoyleNR-04October 2026 · 28 min readRead →
Research Paper · Data & AI

Evidence Coverage and Traceability

Knowing What Supports a Conclusion

A deterministic evidence architecture for preserving and inspecting the governed relationships that actually support AI-generated conclusions.

Steven BoyleNR-03October 2026 · 25 min readRead →
Research Paper · Data & AI

Beyond the Model

A Governed Architecture for Consequential AI Systems

Consequential AI requires an architecture that governs evidence, artifacts, evaluation, authority and learning around the model.

Steven BoyleNR-01October 2026 · 18 min readRead →