Beyond the Model
A Governed Architecture for Consequential AI Systems
Consequential AI requires an architecture that governs evidence, artifacts, evaluation, authority and learning around the model.
How an AI System Can Improve Its Work Without Losing Control
Generative AI systems can increasingly evaluate and revise their own work, but the ability to create another candidate does not establish that another iteration is warranted or that the newest candidate should become authoritative. This paper examines a governed self-optimization architecture developed within the Executive Positioning System (EPS), where probabilistic generation and evaluation operate inside deterministic controls for persistence, validation, continuation, best-valid retention and promotion. The architecture separates candidate creation from authority: a model may generate a successor, assess quality, identify defects and recommend further work, while governed controls determine whether the candidate is admissible, whether another round is justified, which valid version remains strongest, and which exact artifact may be promoted for downstream use. The resulting loop treats evaluation scores as signals rather than authority, preserves earlier valid states when later rounds regress, and constrains autonomous optimization to an explicitly authorized objective and evidence boundary. Its central proposition is that autonomy over iteration does not require autonomy over authority. The optimizer may propose change; governance determines whether another alternative should be created and whether any resulting change is allowed to become authoritative.
Receive the complete paper as a PDF by email.
A straightforward AI optimization loop looks simple:
Generate → Evaluate → Revise → Repeat
Each iteration appears to improve on the previous one. The model produces a candidate. An evaluator identifies weaknesses. The candidate is revised. The process continues until a score reaches a threshold or a round limit is reached.
That describes iteration. It does not yet describe governed optimization.
Iterative feedback and refinement have established precedent in language-model research. Self-Refine demonstrated that LLMs can use generated feedback to revise their own outputs iteratively and improve performance across a range of tasks without additional supervised training or reinforcement learning [1]. NR-04 begins from that capability but asks a different engineering question: what controls are required when iterative refinement is allowed to operate autonomously inside a consequential workflow?
Several questions remain unanswered.
Who determines whether the evaluator's criticism is material? Does the evidence needed to make the requested change actually exist? If the new version scores higher but introduces an unsupported claim, is it better? If a third round changes wording but not substantive quality, should the system continue? If the latest version is worse than the second, which version should survive? And when the system stops, what exactly makes the selected artifact authoritative?
These are architectural questions rather than prompting questions.
A self-optimizing system therefore requires more than a feedback loop. It requires a control structure around that loop.
Iteration creates alternatives. Governance determines whether another alternative should be created and whether it represents improvement.
That distinction became important while building EPS.
EPS was not initially engineered around minimizing model calls. The immediate objective was to make a governed end-to-end workflow work: authoritative evidence had to survive generation, outputs had to be evaluated, revisions had to remain traceable, and accepted results had to move through controlled promotion into later stages.
Once that workflow operated at scale, its computational behaviour became measurable.
A forensic analysis of one accepted Stage 08 run recorded 1,549,530 tokens across 54 provider/model interactions, with an elapsed time of 1:38:45. Approximately 75% of the tokens and 73% of measured provider elapsed time occurred after the first round. Evaluator calls consumed 1,047,473 tokens, more than twice the 502,057 tokens consumed by writer calls.
The problem was not merely that optimization was expensive. Some of that computation produced genuine improvement.
The more revealing problem was computational proportionality.
Three relatively simple sections illustrate it. Northline contained no selected evidence bullets, Career Break contained none, and McMaster contained one. Each nevertheless received three writer rounds and three evaluator rounds. Their Stage 08 costs ranged from roughly 127,000 to 152,000 tokens. Career Break moved from a score of 62.7 to 70.6 between its first two rounds, but remained at 70.6 in Round 3 despite another complete writer/evaluator cycle.
The system already possessed information about the relative complexity of these sections. It was not yet using that information to make its computation proportionate.
That changed the optimization question.
The useful question was no longer:
How do we make each model call cheaper?
It became:
When is another model call justified at all?
The research programme adopted an explicit governing principle: reduce API cost, token consumption and elapsed time while preserving or improving candidate-facing quality, evidence fidelity, governance, reproducibility and auditability.
That qualification matters.
A system can become cheaper by removing evidence, shortening evaluation, reducing validation or simply stopping after the first generation. None of those changes constitutes successful optimization if the result becomes less reliable.
Computational efficiency therefore has to be considered alongside the value produced by the computation.
Conceptually:
Computational value = material improvement obtained / computation consumed
This is not presented here as a formal EPS production metric. It describes the engineering objective.
The numerator is also difficult. Improvement cannot safely be reduced to movement in a single model-generated score. It may involve evidence fidelity, candidate-facing quality, structural clarity, preservation of important proof, removal of unsupported claims or resolution of a specific material defect.
Governed optimization is therefore inherently multi-dimensional.
The goal is not fewer calls. It is to spend additional computation where it produces material value without weakening the constraints that make the result trustworthy.
The first controlled Stage 08 experiment tested the sensitivity of first-round generation to different input packages.
One condition retained selected governed evidence while removing context that did not belong in individual employment-role optimization. Compared with the full-package control across the recorded Trent and WSPS terminal trajectories, that condition used 23.4% fewer tokens and 20% fewer calls and produced stronger final internal blind rankings.
It would be tempting to interpret this as evidence that smaller prompts are simply better.
The measurements do not support that conclusion.
External research provides a related caution. Shi et al. found that LLM reasoning can be degraded by irrelevant information included in the input context [2]. Their experiments do not establish the EPS result reported here, which concerns a different task and architecture, but they reinforce the distinction between providing more context and providing task-relevant context.
At Round 1 alone, the combined writer/evaluator token reduction was only 3.8%. Much of the larger trajectory-level difference arose because the control incurred an additional later round. The complete WSPS reduced-context trajectory was actually slightly slower than its control.
The more important finding was architectural.
The system had been sending some career-wide narrative material intended for broader synthesis into individual employment-role optimization. Removing inappropriate context improved the routing of information to the task being performed.
At the same time, a negative-control condition demonstrated the boundary of that optimization. When substantive achievement evidence was removed, evidence utilization fell to zero and evaluation scores collapsed into the mid-30s for both tested roles.
The research therefore did not support indiscriminate context reduction.
It supported a more precise principle:
Architecture can affect trajectory cost more than prompt size alone, but efficiency should remove unnecessary computation rather than necessary evidence.
A slightly smaller first-round request may save relatively little. A better-routed first round that avoids an unnecessary writer/evaluator successor pair can save much more.
And removing authoritative proof merely to make the request smaller can make the system cheaper and worse.
Evaluation emerged as one of the largest computational costs in EPS.
In the production forensic baseline, evaluator calls consumed more than twice the tokens used by writers.
The controlled experiments showed the same concentration. Across 70 successful measured experimental calls, evaluation accounted for 69.6% of tokens and 75.3% of provider elapsed time.
Those numbers identify a major optimization opportunity. They do not establish that evaluation is unnecessary.
Model-based evaluation itself has established value and known limitations. Zheng et al. found that strong LLM judges could achieve high agreement with human preferences while also exhibiting identifiable position, verbosity and self-enhancement biases and limitations in reasoning [3]. Independent evaluation can therefore provide useful evidence without becoming unquestioned authority.
Evaluators detected evidence-fidelity problems, unsupported claims, missing proof, structural weaknesses and regressions. Removing independent evaluation simply because it is expensive would eliminate one of the mechanisms through which the system discovers that a candidate should not become authoritative.
The architectural question is therefore not:
Can we eliminate evaluation?
It is:
How much evaluation is required for this artifact, at this point in its trajectory, given the defects and uncertainty that remain?
That leads toward selective evaluation, more compact evaluation contracts and evidence-aware continuation. But those are different from abandoning independent review.
The second controlled experiment examined a deeper assumption: if an evaluator recommends a change, should the system make it?
Prior iterative-refinement research establishes that feedback can improve generated output [1]. It does not follow that every feedback item has equal value, that every criticism warrants another generation, or that feedback should automatically acquire continuation authority. Experiment 2 tested that narrower question within EPS.
The experiment decomposed reviewer output into 70 feedback atoms — 36 for Trent and 34 for WSPS — and grouped revision instructions into controlled feedback conditions. Ten independent revisions and ten subsequent evaluations were executed, consuming 490,394 tokens across 20 provider calls.
The results were heterogeneous.
For Trent, the revision using all reviewer feedback ranked first in blind candidate-facing assessment.
For WSPS, presentation feedback alone ranked first, while all feedback ranked second. The style/rubric-only condition ranked last.
Reviewer feedback had real value, but the value depended on the defect and the artifact.
The experiment provided stronger justification for continuation when feedback identified concrete problems involving evidence fidelity, recoverable missing proof, metric or outcome preservation, and material hierarchy, structure or compression.
By contrast, style or wording preferences, rubric preferences and ATS terminology did not demonstrate sufficient standalone value to justify another writer call merely because an evaluator raised them.
This rejected a simple continuation rule:
Evaluator requested change → authorize another round
The stronger rule is:
A governed self-optimizer should not continue because an evaluator found something it could change. It should continue when the evaluator identified a material, supportable defect whose repair is expected to create value.
This converts feedback from an instruction into evidence considered by a continuation policy.
The experiments also exposed a problem with treating evaluator scores as objective functions.
In the second experiment, Trent's style/rubric-only revision gained 7.24 Framework points but ranked only fourth of six in blind candidate-facing assessment. WSPS's material-content condition received the highest Framework score, 90.54, but ranked third. Another WSPS condition lost 1.9 Framework points while ranking second in blind assessment.
The broader experiment therefore treated Framework scoring as useful diagnostic telemetry rather than independent proof of candidate-facing quality.
That matters because a conventional optimization system wants a scalar objective:
maximize score
But if the score is itself produced by probabilistic evaluation and captures only part of the desired outcome, optimizing against it too aggressively risks optimizing the measurement rather than the artifact.
The broader problem has established precedent in machine learning. Gao, Schulman and Hilton showed experimentally that when a learned reward model is an imperfect proxy, stronger optimization against that proxy can eventually reduce performance against a separate gold-standard objective [4]. EPS Framework scoring is not a reinforcement-learning reward model, so the mechanisms should not be treated as equivalent. The relevant parallel is narrower: a measurable evaluation signal can be useful without being sufficient authority for optimization.
EPS therefore treats score movement as one signal among several.
A score can inform an optimization decision without being allowed to become the decision.
These findings led to a more explicit architecture.
Conceptually, Stage 08 became:
Generate → Persist → Evaluate → Validate → Decide → Target → Re-generate → Re-evaluate → Retain or Stop
Each transition has a different purpose.
Generation creates a candidate. Persistence gives the candidate identity and preserves the round. Evaluation produces semantic judgments about the candidate. Validation checks deterministic requirements and integrity conditions. Decision determines whether another round is warranted. Targeting constrains what the successor round is expected to repair. Re-generation produces another candidate rather than silently replacing the prior one. Re-evaluation determines what actually changed. Retention preserves the strongest valid state. Stopping ends the trajectory when continuation is no longer sufficiently justified.
This architecture extends a principle developed earlier in this research series [6,7]:
Generation creates a candidate. Validation determines admissibility. Promotion creates authority.
Self-optimization inserts additional candidates between generation and promotion. It does not change the authority model.
A system capable of finding imperfections can almost always find another possible revision.
That makes "the evaluator suggested improvements" a poor stopping policy.
The experimental evidence instead supports distinguishing between material defects and optional refinements.
A material defect might involve unsupported or overstated evidence, omission of available material proof, loss or distortion of an important metric, a significant hierarchy or compression problem, or another issue that materially affects fidelity or candidate-facing value.
An optional refinement might involve stylistic preference, wording alternatives or rubric-oriented adjustments that do not materially improve the artifact.
Materiality alone, however, is not enough.
Suppose an evaluator says an employment section needs stronger evidence of a particular capability. The system may possess authoritative unused evidence that can support the repair.
Or that evidence may not exist.
Without distinguishing these cases, an optimization loop can encourage the writer to satisfy the evaluator anyway. That creates pressure toward unsupported inference or increasingly polished restatement of the same limited evidence.
The evidence diagnostics developed in NR-03 therefore become part of the optimization architecture [8].
A material defect alone is not sufficient. The system must also establish whether governed evidence capable of supporting the repair exists [8].
Conceptually:
Material defect + available governed evidence + bounded repair opportunity → candidate for continuation
The model can make the semantic judgment that a defect exists. The framework can determine whether the governed evidence relationship required for a repair is actually available.
A model should not be asked to solve an evidence deficiency as though it were a writing deficiency.
Once continuation is authorized, the next round should not simply receive the previous output and an instruction to "make it better."
A governed successor has a specific relationship to its predecessor.
It should know which artifact it is revising, which findings authorized continuation, which defects it is expected to repair, which evidence is available for those repairs, what valid content must be preserved, which claims or evidence relationships must not be weakened, and which round identity the new candidate will receive.
This creates a bounded optimization target.
The successor round is not a fresh attempt at the entire problem. It is a controlled intervention against identified deficiencies.
That matters because autonomous iteration increases the opportunity for drift. A broad rewrite can solve the identified problem while weakening something that was already correct.
A targeted contract narrows the change surface.
The more autonomous the iteration becomes, the more specific the authorization to change should become.
This separation between autonomous execution and governed authorization is consistent with broader AI risk-management principles that distinguish system operation from the governance, measurement and management of its risks across the lifecycle [5]. The successor contract described here is the EPS implementation response; it is not prescribed by NIST.
As first-pass quality improves, successor instructions should also become narrower. The system should not repeatedly pay to reconstruct decisions that are already valid.
Iterative systems often assume that the newest version is the current version.
That is unsafe when generation is probabilistic.
Three concepts need to remain separate [7]:
Latest — the most recently generated candidate.
Best-valid — the strongest candidate that currently satisfies the governing requirements.
Promoted — the exact candidate that has been explicitly given authority for downstream use.
These may point to the same artifact. They do not have to.
A later round can fail validation. It can introduce a material regression. It can improve one dimension while weakening another. It can score higher while becoming less faithful to evidence.
The system therefore needs to preserve optimization history independently from artifact authority.
Optimization lineage answers what happened. Artifact authority answers what the system is allowed to rely on.
This distinction prevents iteration itself from silently changing authoritative state.
A successor round is not guaranteed to dominate its predecessor.
The second experiment demonstrated this directly. Some revisions improved Framework scores without producing the strongest blind result. Others lost score while remaining highly ranked by blind assessment. One WSPS style/rubric-only revision was blind-worst and also produced an unsupported-claim signal.
This means rollback is not merely an exception-handling feature.
It is part of normal optimization.
The system must be capable of saying:
Round 3 exists, but Round 2 remains better.
That is why best-valid retention is more than version history. It is a governing mechanism.
The experimental work ultimately treated best-valid retention as non-negotiable, and the later implementation tests enforce the same principle: a later candidate should not displace the incumbent merely because it is newer or receives a higher aggregate score when it introduces a material regression.
A self-optimizer that cannot preserve an earlier stronger state is not safely optimizing.
It is merely moving forward.
Not every weakness can be solved by additional reasoning.
Sometimes the available evidence places a ceiling on legitimate improvement.
An evaluator may prefer a more compelling example, a larger outcome or more direct proof, but none may exist in the governed evidence base. Another round cannot create new evidence. It can only restructure what exists, omit the claim, acknowledge the limitation or risk implying more than the evidence supports.
That condition should be visible to the optimizer.
An evidence ceiling is not computational failure. It is a boundary on legitimate optimization.
Diminishing returns are different.
The evidence may remain adequate and further improvement may still be possible, but successive rounds may be producing progressively less value.
The early forensic trajectories illustrate the distinction:
| Section | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| Northline | 58.8 | 62.7 | 63.5 |
| Career Break | 62.7 | 70.6 | 70.6 |
| McMaster | 70.9 | 76.0 | 77.1 |
These observations do not establish a universal stopping threshold. They show why optimization must consider trajectory rather than only the current artifact.
A governed optimizer can ask whether a material defect was resolved, whether a new one appeared, whether evidence utilization improved, whether candidate-facing quality improved, whether the remaining weakness is evidence-resolvable, and whether marginal improvement is shrinking.
An evidence ceiling says the system lacks legitimate material with which to make the requested improvement.
Diminishing returns say additional computation is producing progressively less value.
They require different explanations even when both result in stopping.
Bounded autonomy requires computational limits.
But a maximum round count is a safety control, not evidence that the optimization process has converged.
If the system stops because the permitted number of rounds has been exhausted, all that has been established is that the circuit breaker fired.
The system may have converged earlier. It may still contain a material repairable defect. It may have reached an evidence ceiling. It may simply have exhausted its computational allowance.
These conditions should not be collapsed into the same state.
A mature optimizer therefore needs to distinguish among conditions such as:
Shortlist ready — the artifact satisfies the governed readiness criteria for its intended use.
Diminishing returns — additional computation is no longer expected to produce sufficient material improvement.
Evidence ceiling — substantive improvement is constrained by the authoritative evidence available.
Integrity halt — continuation is inappropriate because a governing condition has failed.
Hard limit — computation has reached an externally imposed boundary.
The last condition protects the system. It does not explain the system.
A safety ceiling can stop computation. It cannot prove that optimization is complete.
One of the more consequential findings from the research was that later-round improvement often involved information the system already possessed.
The second experiment found that a substantial portion of revision value came from evidence recovery and hierarchy repair rather than pure prose polishing. This suggested a clear architectural direction: make evidence and proof priorities more explicit before the first write rather than repeatedly paying to recover them later.
Subsequent first-round analysis reached a compatible conclusion. In the case examined, approximately 62.5% of material Round 2 directives were either explicit before Round 1 or reasonably derivable from information already available. Approximately 85–90% of that intelligence was technically present in the original writer payload, but only about 30–40% had been surfaced as explicit actionable instruction.
These figures are observational rather than controlled experimental findings. They do not establish the magnitude of improvement that a redesigned first-round architecture would produce.
Research on LLM context use reinforces the general caution that the presence of information in a model's context does not guarantee that it will be used effectively. Liu et al. found substantial performance sensitivity to the position of relevant information in long contexts, while Shi et al. showed that irrelevant contextual information can distract model reasoning [2,9]. These findings do not establish the EPS first-round result, but they support the distinction between context availability and effective use.
Available context is not necessarily actionable context.
Giving a model more information does not guarantee that the model will assign every piece of information the intended priority.
A stronger first-round architecture may therefore transform governed evidence into a concise pre-write agenda: proof that must be used, metrics that must be preserved, claims that must remain bounded, evidence priorities, known evidence ceilings, section-specific objectives and preservation constraints.
The objective is not to make the writer deterministic. It is to make predictable requirements explicit before probabilistic composition begins.
The cheapest repair is often the one the architecture prevents from becoming necessary.
Independent evaluation still remains necessary because some defects only become visible after composition. A bullet may become overpacked. A wording choice may overclaim. Two valid points may become redundant when combined. A hierarchy may prove weak only when the section exists as a whole.
The objective is therefore not to eliminate later evaluation. It is to reserve later computation for problems that actually emerge later.
Governance does not require a human to approve every sentence or every model call.
It requires clear decision rights.
A human can authorize an optimization scope: improve this artifact within these evidence, integrity and round constraints.
Within that scope, the system can generate, evaluate and perform bounded continuation according to approved policy.
The human does not need to manually choose every wording change. Nor does the model receive unrestricted authority to continue indefinitely or promote whatever it produces.
This creates a useful separation:
This allocation is compatible with lifecycle-oriented AI governance rather than continuous human micromanagement. NIST's AI RMF treats governance as a cross-cutting function while measurement and management continue throughout operation [5]. NR-04's specific allocation of decision rights is an EPS architecture, not a NIST-prescribed operating model.
The architecture therefore allows greater autonomy at the level of execution while retaining governance at the level of authority.
Autonomy over iteration does not require autonomy over authority.
Iterative AI systems also fail operationally.
Provider calls can time out. Responses can be malformed. A writer may complete while its evaluator fails. A process can restart between rounds. A retry can accidentally duplicate expensive work.
A governed optimizer therefore needs recoverable checkpoints.
The system should know whether it is waiting for generation, evaluation, deterministic validation, continuation or promotion. Recovery should resume from the last valid governed state rather than reconstructing state from whichever output happens to be visible.
The computational architecture review identified recovery conformance as a prerequisite to more aggressive optimization changes. It also distinguished generic recovery tests from actual adapter-level recovery behaviour.
This matters because computational efficiency and recoverability interact.
A system that saves tokens during normal execution but repeatedly replays expensive completed work after failures is not computationally efficient.
A governed optimizer needs a recovery model, not merely a retry policy.
The experiments did not produce a single optimization rule.
They changed the architecture's understanding of where optimization value comes from.
First, context should be routed according to the task rather than accumulated into one universal package. The first experiment showed that section-specific context could outperform inappropriate broader context without sacrificing governed evidence.
Second, predictable evidence-use requirements should move earlier. A governed role-local projection and pre-write evidence-use agenda emerged as ways of reducing avoidable evidence recovery after first generation.
Third, continuation should be based on typed material findings and evidence-aware repair opportunities rather than evaluator preference generally.
Fourth, evaluator cost should be reduced by improving evaluation representation and routing rather than assuming evaluation can simply be removed.
Fifth, semantic and deterministic responsibilities should remain separate. Evidence authority, independent review, best-valid recovery and governed promotion remain necessary even while their computational representations are optimized.
The sequence matters.
The sequence below summarizes the engineering progression observed in EPS; it is not presented as an established software-development methodology.
Make it work → make it governable → instrument it → understand the cost → optimize the architecture → measure again
That sequence reduced the risk of making the system cheaper before understanding what its expensive components were actually contributing.
The measured production data shows why continuation policy matters.
For the analysed CPA case, normal Stage 08 Round 1 required 16 calls and 387,626 tokens, with 18.48 minutes of measured provider time.
Post-Round-1 work required 36 calls and 1,161,904 tokens, with 41.38 minutes of measured provider time. Two additional unbound network failures added 7.82 minutes.
Across the application, the 18 post-Round-1 Stage 08 section rounds represented 46.1% of all recorded application tokens.
But those figures should not be interpreted as showing that 46.1% of application computation was waste.
Some successor rounds materially improved output. Some repaired fidelity. Some recovered evidence. Others regressed or produced little additional value.
That distinction is precisely why the economics of self-optimization cannot be separated from governance.
If another round costs computation but fixes a material unsupported claim, its value may be high.
If another round costs the same amount to satisfy a stylistic preference while introducing a regression, its value may be negative.
The relevant question is therefore not simply the price of another round. It is:
What is the expected material value of another round, given the remaining defect, available evidence, current trajectory and risk of regression?
That question points beyond architecture toward stopping science.
Existing AI governance frameworks establish the broader need to govern, measure and manage AI-system risks throughout operation [5]. NR-04 addresses a narrower architectural question inside that problem: which decisions may be delegated to an autonomous optimizer and which must remain outside its authority?
The resulting architecture gives the AI system meaningful autonomy.
Within an authorized optimization sequence, it can generate alternatives, evaluate outputs, identify repair opportunities, perform targeted revisions, compare results and determine whether continuation remains justified under the governing policy.
But that autonomy has boundaries.
The optimizer should not be permitted to invent evidence to overcome an evidence ceiling. It should not ignore an integrity failure because an aggregate score improved. It should not assume that latest means best. It should not promote its own output merely because it prefers it. It should not silently expand the scope of an existing authorization. And it should not describe a hard computational limit as evidence of convergence.
These constraints do not undermine autonomy. They define where autonomy belongs.
The optimizer may propose change. It should not control the conditions under which change becomes authority.
The evidence supports governed iterative optimization. It does not establish a universal optimal policy.
The production observations are case-specific. The measured Stage 08 runs reveal real system behaviour, but they are not a controlled sample from which universal workload percentages can be inferred.
Experiment 1 was deliberately narrow. Its 23.4% trajectory-level token difference should not be extrapolated across the entire system. The Round 1 difference was much smaller, and one tested trajectory was not faster.
Experiment 2 does not establish a universal hierarchy of reviewer feedback. All feedback performed best for one role while presentation feedback performed best for another. The experiment supports selective continuation, not a fixed global ordering of feedback classes.
Evaluator cost is not equivalent to removable cost. The measured 69.6% token concentration identifies where optimization research is warranted. Evaluators also detected material defects.
Framework scores are not ground truth. Blind candidate-facing assessment disagreed with score movement in multiple experimental conditions.
The first-round intelligence findings are observational. They identify a plausible opportunity to make existing intelligence more actionable before first generation, but they do not experimentally establish the magnitude of benefit.
Experiment 3 was designed to test aspects of stronger first-pass intelligence. It was not executed. Its proposed treatment effects, quality tolerances and hypotheses are not findings.
Finally, the current stopping architecture is governed but has not been established as empirically optimal. It can distinguish stopping conditions, preserve best-valid state and bound continuation. The question of when another round ceases to justify its cost and risk remains open.
These limitations define the boundary between what the system implements, what the experiments observed and what subsequent research still needs to establish.
A governed optimization architecture makes a more rigorous next question possible.
Once the system can identify the current best-valid artifact; the latest candidate; material unresolved defects; available repair evidence; evidence ceilings; quality trajectories; regression history; model-call cost; and the reason for continuation or stopping — it can begin to investigate whether another round is likely to be worth performing.
This is different from an arbitrary round limit.
It is also different from asking the model whether it thinks it is finished.
The stopping decision becomes an observable and testable policy.
The EPS research harness provides a mechanism for investigating that decision. A governed run can be forked, the optimizer's stop or continue decision can be observed, and a stop can deliberately be challenged by executing an additional round. If the challenge produces material improvement without unacceptable regression, the original stop may have been premature. If the challenge produces no material improvement or introduces a material regression, the stop gains empirical support.
That changes the question from:
Did the system stop?
to:
Did the system stop at the right time?
That is the transition from governed self-optimization to empirical stopping research — the subject of NR-05 [10].
Building an iterative AI system initially makes self-optimization look like a generation problem.
It is not.
Generation is the easiest part of the loop.
The harder problem is deciding when another generation is justified, what it is allowed to change, what evidence it can use, how its result should be evaluated, whether it is actually better, which earlier state must be preserved, and when the system has sufficient reason to stop.
The EPS evidence shows why those distinctions matter.
A production Stage 08 run consumed most of its optimization computation after the first round. Controlled experiments showed that reducing context did not produce value merely because prompts became smaller; better routing changed the trajectory. Removing substantive evidence made the system worse. Reviewer feedback materially improved some outputs, but its value varied by defect and role. Evaluator scores sometimes diverged from blind candidate-facing judgment. Later revisions could regress. And some later-round work recovered evidence the system already possessed.
Together, these findings point toward a different conception of self-optimization.
A governed self-optimizer is not a system that keeps rewriting its work until a model says it is better.
It is a system that preserves authoritative evidence, evaluates candidates independently, authorizes additional computation for material and supportable reasons, constrains what successor rounds are permitted to repair, retains stronger prior states when successors regress, and keeps optimization separate from authority.
The system should be able to answer four different questions:
Those questions should not collapse into a single model decision.
Prior work establishes that iterative feedback can improve LLM outputs [1], that context composition affects model performance [2,9], that model-based evaluators are useful but imperfect [3], and that optimizing strongly against an imperfect proxy can diverge from the underlying objective [4]. NR-04 contributes the governance architecture around those capabilities: continuation, repair, retention and promotion remain distinct decisions, and autonomous optimization does not acquire authority merely because it can generate or evaluate another candidate.
That separation is what makes autonomous improvement governable.
The optimizer may propose change. It should not control the conditions under which change becomes authority.
This paper draws on the current EPS system architecture, implemented Stage 08 optimization and closed-loop services, deterministic validation and promotion controls, implementation tests, production runtime analysis, two completed controlled Stage 08 experiments, subsequent first-round intelligence analysis, and the closed-loop research harness.
Implemented. EPS separates generation, evaluation, deterministic validation, continuation, optimization lineage, best-valid retention and promotion. It represents explicit continuation and stopping conditions, including evidence ceilings, diminishing returns and integrity-related termination, and bounds closed-loop optimization through governed authorization.
Implemented and tested. The system can retain an earlier valid artifact when a later candidate materially regresses; optimization state can remain distinct from promoted authority; targeted successor work is bound to governed source and target state; and the research harness can challenge stopping decisions without modifying the production run.
Measured in production. Runtime analysis exposed substantial computational concentration after first generation. In the principal Stage 08 forensic baseline examined here, 1.55 million tokens were consumed across 54 provider/model interactions, with approximately three quarters of token use occurring after Round 1. A later application-level analysis found that post-Round-1 Stage 08 section work alone represented 46.1% of recorded application tokens. These observations describe the measured runs; they are not presented as universal EPS workload ratios.
Experimentally observed. Experiment 1 demonstrated that task-appropriate context and retained governed evidence could produce stronger terminal trajectories with fewer calls and tokens than the tested full-package control, while removing substantive proof materially damaged quality. Experiment 2 demonstrated that reviewer-feedback value varied by defect and role, and that evaluator score movement did not reliably correspond to blind candidate-facing preference.
Observed but not experimentally established. Subsequent first-round analysis indicated that substantial later-round intelligence in the examined case was already available before initial generation but was not always represented as explicit actionable instruction. This supports further investigation of pre-write evidence-use agendas but does not establish their causal effect.
Designed but not executed. Experiment 3 proposed a controlled evaluation of stronger first-pass intelligence and governed evidence-use projection. It was not executed and contributes no measured findings to this paper.
Unresolved. The empirical accuracy of current stopping decisions across a broader corpus, the expected material value of additional rounds across different task types, and the conditions under which further optimization ceases to justify its computational cost and regression risk remain open research questions.
The claim made here is therefore bounded. The research supports an architecture for governing self-optimization. It does not establish that the current optimization policy is universally or mathematically optimal.
Receive a link to the complete paper by email.
Steven Boyle is the founder of Northline Advisory, a technology advisory and research practice focused on technology leadership, enterprise transformation, governance, data and governed AI. His work draws on more than two decades of executive and operational experience across higher education and public-interest organizations.
A Governed Architecture for Consequential AI Systems
Consequential AI requires an architecture that governs evidence, artifacts, evaluation, authority and learning around the model.
Governing State, Authority and Lineage
An implemented architecture for treating durable, versioned and governed artifacts—not transient model responses—as the unit of enterprise AI control.
Knowing What Supports a Conclusion
A deterministic evidence architecture for preserving and inspecting the governed relationships that actually support AI-generated conclusions.