Governed Learning
How AI Systems Can Learn Without Turning Experience Into Truth
An evidence-grounded architecture for using decision history as an advisory signal without allowing experience or frequency to become factual authority.
Engineering Reliable Software When the Implementer Is Probabilistic
As software development becomes increasingly AI-assisted, reliability cannot depend on prompt quality or close supervision alone. The governing problem is architectural: what is allowed to change, what must remain authoritative, how completeness is verified, and what conditions must be satisfied before runnable code becomes accepted system state. This paper presents the architectural-governance model implemented in the Executive Positioning System (EPS), showing the progression from prompt guidance and execution logging to machine-validatable schemas, repository authority, certified architecture, bounded implementation, and frozen baselines. It introduces architecture as a governing contract that constrains probabilistic implementation through deterministic scope, validation, certification, promotion and baseline controls. The paper also documents EPS-ARCH-010, an authority-correction program that resolved six conflicts before certification: family membership, section identity and ordering, optimization-unit boundaries, publication-readiness interpretation, Section Family completeness, and rendering in the execution path. A governed implementation lifecycle is then described, together with a decisive fail-closed example in which a release candidate that executed successfully from a working tree was denied baseline status because release-governance conditions were not satisfied: the manifest exposed untracked runtime modules and the repository could not reproduce the running application. By contrast, the accepted path required verification, certification, closure and approved promotion after 74 defined tests passed and identified observations were resolved. The central result is conceptual but operationally grounded: implementation, testing, certification, closure, promotion and acceptance are distinct states. A governed AI engineering system becomes reliable not by trusting the implementer, but by constraining implementation within an architecture that retains authority over what becomes system truth.
Receive the complete paper as a PDF by email.
Generative AI has become a practical tool for producing working software.
In recent development practice, AI-assisted coding has progressed from suggestion to generation to substantial implementation. Developers using AI coding assistants routinely accept generated code that would have taken significant time to write manually. In more integrated workflows, AI agents generate, revise, and iterate on software components with minimal per-line human inspection.
Repository-level software engineering has become an active research problem for language-model agents because it requires more than isolated code generation. Real-world repository tasks can require understanding existing code, coordinating changes across multiple functions and files, navigating the development environment, and executing tests and other tools [7,8].
The reliability problem this creates is not adequately described as a code-quality problem.
Generated code can run correctly. It can pass the tests it was asked to pass. It can satisfy the task it was given. And it can still be wrong—not because the model produced bad code, but because the code violates something the model was not adequately required to respect: the architecture of the system it is modifying.
This distinction between local correctness and systemic correctness is not new. Engineering disciplines that require high reliability have long maintained it. A component can satisfy its local specification while failing to conform to system-level constraints. Local verification is necessary but not sufficient for systemic acceptance.
What changes when implementation is substantially performed by a probabilistic AI agent is the mechanism by which that distinction must be maintained.
A human engineer carrying architectural intent in working memory is subject to forgetting, misunderstanding, and inconsistency. A probabilistic AI agent is subject to different but related problems: it may resolve architectural ambiguity plausibly but incorrectly, it may not reproduce identical decisions on repeated identical prompts, and it may express confidence in an architectural interpretation that is not authoritative.
Neither human nor AI implementation is guaranteed to preserve architectural intent through instruction alone.
The engineering response is not better instruction. It is better governance.
The model was being instructed to respect authority. The system was not yet enforcing authority.
This paper examines that problem through a longitudinal engineering case study of the Executive Positioning System (EPS), an AI-mediated software system whose architecture evolved during development from prompt-carried instructions into repository-enforced governance. The study reconstructs 14 supported engineering episodes using architecture documents, implementation records, revision logs, accepted architecture decision records, tests, certification reports, release evidence, Git checkpoints, closure records, and documented recovery incidents.
The central claim is not that architecture eliminated errors or causally reduced defect rates. The evidence does not support that conclusion, and the study does not make it.
The evidence supports a narrower and more consequential conclusion: in EPS, architecture evolved from guidance into a governing contract. It assigned singular authority, bounded permissible implementation, and separated code production from acceptance through repository state, validation, certification, promotion, and frozen baselines.
NR-07 is a longitudinal, evidence-triangulated single-system engineering case study.
The study reconstructs the engineering history of EPS using primary artifacts produced during development: architecture documents, requirements specifications, implementation records, revision logs, accepted architecture decision records, tests, certification reports, release evidence, Git checkpoints, closure records, and documented recovery incidents. These artifacts were coded into 14 supported engineering episodes.
Five episodes provide particularly strong evidence of the distinction between local functional correctness and systemic architectural acceptance. These are analyzed in detail in Section 4.
The most complete episode, EPS-ARCH-010, is analyzed separately in Section 5. It provides the strongest evidence of architecture functioning as an enforced governing contract rather than as advisory guidance.
The study is observational and retrospective. It does not establish causal relationships between architectural governance and defect rates. It does not claim that EPS's architecture is universally optimal or that the findings generalize beyond this system without further study.
Quantitative anchors from the evidence record:
| Measure | Observed result |
|---|---|
| Supported engineering episodes | 14 |
| Strong local-versus-systemic cases | 5 |
| EPS-ARCH-010 architecture corrections | 6 |
| Accepted ADRs in consolidated set | 8 |
| Runtime integration tests | 74 passed / 0 failed |
| S01 implementation baseline | 22/22 targeted + 28/28 governed regression |
| S02 implementation baseline | 21/21 targeted + 35/35 governed regression |
| H04 focused verification | 460 passed / 7 environment-dependent skips / 0 reported failures |
| Repository tags observed during evidence review | 9 |
| Runtime-certification observations subsequently resolved | 2/2 |
These figures are research findings within the scope of the evidence. The 74 passed / 0 failed runtime integration tests establish what the certification covered; they do not establish universal system correctness. The 14 episodes are the supported record; they are not an experimental sample from a larger designed study.
Software architecture has long been understood as more than a diagram of system structure. Standards formalize the concepts and relationships represented in architecture descriptions, while software-architecture research has emphasized the importance of explicit architectural decisions, assumptions, rationale and accumulated architectural knowledge [1–3].
NR-07 begins from that established foundation but examines a different question: what changes when substantial implementation is delegated to a probabilistic agent?
Early EPS development followed a pattern common in AI-mediated software projects: architectural intent was expressed through prompt instructions, documentation, and conventions that the model was expected to understand and respect.
This model has genuine utility. A capable model can apply architectural guidelines consistently within a session, follow naming conventions, respect documented constraints, and produce code that appears to conform to the stated architecture.
The limitations emerged incrementally.
Architectural intent carried only through instruction is subject to loss, misinterpretation, and inconsistent application. As EPS grew more complex, it became progressively harder to ensure that AI-generated implementation consistently respected authority boundaries, artifact ownership, promotion rules, and the relationships between components.
The problem was not model capability. The model could apply instructions accurately when they were well-formed, unambiguous, and present in the active context.
The problem was authority. An instruction is not an enforcement mechanism. A model that has been told to follow an architectural rule may follow it, but there is no system-level guarantee that it will do so, no mechanism to detect when it has not, and no consequence when it violates the rule while still producing code that runs.
The early EPS architecture was guidance. The engineering system needed to evolve toward governance.
Five episodes from the EPS engineering record provide particularly strong evidence of the distinction between locally correct implementation and systemically acceptable implementation.
A release candidate was functionally complete. The application ran. User-facing behavior appeared correct. The candidate was denied release status.
The reason was not a runtime defect. The committed repository state could not reproduce the running application. The relationship between committed code and the running application was not fully traceable through the repository record.
Runnable ≠ releasable.
This episode established that acceptance criteria for a release were not simply functional. The repository state had to be the authoritative source of the application—reproducibly, verifiably, as a condition of release rather than as a subsequent documentation task.
A runtime defect was identified. A local hotfix resolved the immediate problem. The fix was technically correct and addressed the failure that triggered it.
The fix also exposed distributed architectural authority. The hotfix had been applied to a component whose authority boundaries were not clearly established in the architecture. A similar defect could arise from a different path governed by a different authority assignment—or by no clear authority assignment at all.
The hotfix was useful. It was not architecturally complete. Accepting it without resolving the authority question would have carried the underlying structural problem forward into subsequent implementation.
This episode became EPS-ARCH-010, analyzed in detail in Section 5.
An optimization stage was implemented and functional. During review, it became apparent that the implementation reflected one architectural model while a separate specification document described a different architectural model for the same stage.
Both architectures were internally consistent. Neither was incorrect in isolation. They were incompatible: they assigned authority differently, made different assumptions about artifact ownership, and would produce different behavior under conditions that the current test coverage did not exercise.
The working implementation could not be accepted while two incompatible architectural models coexisted for the same component. The conflict had to be resolved at the architecture level before implementation could be accepted.
A component passed its focused unit and integration tests. Test coverage was appropriate for the component's stated scope. The tests were well-designed for what they measured.
End-to-end defects were subsequently identified that the focused tests had not covered. The component's behavior under conditions involving other components, promotion paths, and downstream artifact use was not established by the focused test results.
This episode established an explicit engineering principle: focused tests establish behavior within their scope. They do not establish end-to-end acceptance. Certification and release required separate end-to-end verification.
A research output was produced. The output was technically correct. Its content was valid. The output was rejected.
The rejection was based on provenance: the output had been produced through a path that did not conform to the established authority chain. Its source, derivation, and approval history could not be verified against the governed provenance requirements.
Correct output with invalid provenance is not an accepted output in a governed system. Provenance is not a documentation convention. It is an acceptance condition.
Provenance itself is an established technical concept. The W3C PROV family provides a formal model for representing relationships among entities, activities and agents so that origin and derivation can be represented and exchanged across systems [6]. The EPS finding is narrower and operational: in this case, provenance became part of the acceptance decision rather than merely descriptive metadata. NR-03 established the earlier Northline distinction between preserved support relationships and plausible after-the-fact reconstruction [10].
EPS-ARCH-010 is the most complete episode in the evidence record. It illustrates the transition from architecture as guidance to architecture as an enforced governing contract.
The episode began with the hotfix described in Section 4.2. What the hotfix revealed was not simply a defect in one component. It revealed that the architecture itself contained distributed authority—multiple locations where the same decision could be made, potentially differently, without a mechanism to ensure consistency.
A formal review was initiated. Recording consequential architecture decisions and their rationale has established precedent in software architecture [2,3]. In EPS-ARCH-010, however, the architecture record performed an additional operational function: it determined which component was permitted to own a consequential decision before implementation authority was released.
The review identified six authority and completeness conflicts in the existing architecture:
These were not defects in the code. They were defects in the architecture. Correcting them required formal architectural work: requirements were frozen, the six conflicts were resolved, the resulting architecture was documented, and the architecture was then independently certified before implementation authority was released.
The sequence matters.
In EPS-ARCH-010, implementation did not proceed while architectural authority was distributed or conflicted. Requirements were frozen first. Architecture was corrected and made singular. The corrected architecture was certified. Implementation authority was then released—bounded by the certified architecture rather than by prompt instructions.
This sequence is the governing contract in operation:
Architecture must be authoritative before implementation is permitted to proceed. Implementation is bounded by the certified architecture. Deviation from the architecture is not a style preference to be resolved later; it is an acceptance failure to be resolved before promotion.
The architectural evolution of EPS made implementation progressively more subordinate to architecture rather than guided by it.
This distinction is significant.
Guidance suggests. Subordination constrains. A developer who is guided by an architectural principle may choose to deviate if the implementation seems to require it. A developer—or AI agent—whose implementation is bounded by a certified architecture does not have that discretion. The architecture is the constraint, not the recommendation.
In EPS, this subordination was implemented through several mechanisms:
Repository authority. The repository is the authoritative source of system state. Implementation that cannot be reproduced from the repository is not accepted system state, regardless of whether it runs. The EPS artifact model—governing state, authority and lineage across artifact types—was established in earlier Northline work and provides the underlying framework for this authority assignment [9].
Bounded implementation baselines. Implementation baselines specified not only what was to be implemented but also the governed regression scope. The S01 baseline required 22/22 targeted tests and 28/28 governed regression tests. The S02 baseline required 21/21 targeted and 35/35 governed regression. These baselines made implementation bounded and verifiable rather than open-ended.
Certified architecture as a precondition for implementation authority. In the most consequential episodes, implementation did not proceed until architecture was certified. Certification was not a review of completed implementation. It was a condition that preceded implementation authority.
Together these mechanisms converted architectural intent from something carried in instructions into something enforced by the engineering system.
Implementation freedom inside a contract is desirable. Implementation freedom to redefine the contract is not.
One of the most persistent failure modes in software engineering is the conflation of adjacent states: treating a working implementation as a tested implementation, a tested implementation as a certified implementation, and a certified implementation as a closed program.
The use of controlled baselines and formal change management is established software-engineering practice. Configuration-management guidance treats controlled system configurations, review, change control and retained prior states as mechanisms for maintaining system integrity across change [4]. EPS did not invent those mechanisms; the relevant question in this study is how they functioned when implementation itself could vary probabilistically.
EPS maintained these as distinct states with explicit transitions between them.
The most instructive instance of this distinction in the EPS evidence record is the separation between runtime certification and formal program closure.
Runtime integration was certified against 74 automated tests, all passing. That certification established what it established: the runtime behavior covered by those 74 tests, at that point in the repository, under the conditions exercised by those tests.
Formal program closure was separate and occurred later. The runtime certification had produced two documentation observations. Program closure required that these observations be resolved. They were resolved. Closure was then recorded.
The 74 passing tests were necessary for certification. Certification was necessary but not sufficient for closure. Closure required the resolution of documented observations that certification had not yet addressed.
A program that is certified is not necessarily closed. Closure is a separate state with its own requirements.
In EPS, generated output does not automatically become system state. Promotion is the mechanism by which accepted output becomes part of the governed baseline.
Promotion requires that the output have passed verification, that its provenance be valid, that it conform to the certified architecture, and that it be accepted through the repository—not merely produced and used. NR-05 provides the earlier Northline evidence that a newer generated output is not necessarily the strongest valid output: later optimization rounds demonstrated regression, and best-valid retention was established as a necessary protection [12].
This separation between generation and promotion is not administrative. It is the mechanism by which the governing contract is enforced at the artifact level. A generated artifact that has not been promoted is not governed system state. An artifact that has been promoted through the governed path is authoritative—its state, source, and acceptance history are part of the record.
A governance mechanism that cannot reject non-conforming work is not a governance mechanism. It is a documentation convention.
In EPS, the engineering system was designed to be capable of refusing work that did not satisfy its acceptance criteria. The release candidate described in Section 4.1 was refused because its repository state could not reproduce the running application. The output described in Section 4.5 was refused because its provenance was invalid. The architectural work in EPS-ARCH-010 was refused as complete until the six authority conflicts were resolved.
These refusals are informative precisely because they were enforced. A system that records non-conformance but accepts non-conforming work regardless provides weaker evidence about its actual governance than a system that refuses non-conforming work and documents the refusal.
The evidence record in EPS therefore includes not only what was accepted but what was rejected and why. The rejections are part of the evidence base. They demonstrate that the governance mechanisms had actual authority over acceptance decisions.
Bounded implementation baselines are a direct engineering response to a characteristic of probabilistic implementation: an AI agent asked to implement a change will not necessarily limit itself to the change requested. It may refactor related code, update naming conventions, modify adjacent components, or apply improvements it infers are desirable.
These changes may be individually reasonable. In aggregate they create a significant governance problem. If the scope of an implementation is open-ended, it is not possible to certify what has been changed, verify that the change is bounded, or be confident that the regression scope covers everything the implementation touched.
Bounded change addresses this directly [4]. An implementation baseline specifies not only the target behavior but also the permitted scope of change and the regression scope that covers it. Implementation outside the bounded scope is not accepted.
This is a governance mechanism, not a prompt instruction. The bound is enforced through the acceptance criteria, not through the agent's judgment about what is in scope.
The underlying configuration-management mechanisms are established practice [4]. What changed in EPS was their role in relation to probabilistic implementation: once an intermediate state had been accepted, the engineering system preserved it deterministically rather than asking a model to reconstruct the same decision in a later attempt.
The result is a system where the scope of every accepted change is known, documented, and covered by a specific regression scope—not inferred from what the agent was asked to do. The earlier Northline work on governed self-optimization established the complementary principle: that the system, not the model, must determine when iterative improvement is bounded and when it constitutes regression [11]. NR-05 further established that successor generation does not automatically yield a better result, and that earlier accepted states may represent the strongest valid artifact [12].
The progression from prompt-carried architectural intent to repository-enforced governance in EPS was not driven by changes in the AI model's capability.
The same class of AI model that implemented early EPS stages under prompt-guided architecture also implemented later stages under repository-enforced bounded baselines. The reliability improvement that the architecture achieved was not a consequence of using a more capable model. It was a consequence of surrounding the model with a stronger governing system.
This is a significant point for engineering teams considering reliability improvements in AI-mediated development.
The expected response to reliability problems in AI implementation is often to improve the model: use a newer model, write better prompts, provide more context, apply chain-of-thought reasoning. These improvements have genuine value. They are not the primary mechanism by which governed reliability is achieved.
Governed reliability requires a system that constrains what the model may decide, verifies what it has produced, and controls what its work becomes. That system is deterministic. Its reliability does not depend on the model's consistency. It enforces constraints regardless of how the model resolved its implementation choices. NR-04 established the complementary principle within the EPS workflow: that the optimization process must be governed by the system rather than by the model's own assessment of improvement [11]. NR-06 extended that principle to historical experience: that a system's previous outputs must not silently acquire the authority to direct future decisions [13].
The central finding of NR-07 is that EPS architecture evolved from guidance into a governing contract.
A governing contract has three properties that distinguish it from guidance:
A governing architecture assigns singular, explicit authority for every consequential architectural decision. It does not leave authority implied, distributed, or ambiguous. When a decision must be made about component ownership, artifact authority, promotion rules, or acceptance criteria, the governing architecture specifies who or what holds that authority—and that authority is singular.
Distributed authority is an architectural defect. It is correctable, as EPS-ARCH-010 demonstrates, but it must be corrected before implementation authority is released. An architecture with distributed authority over a consequential decision cannot enforce consistent behavior because there is no single governing source.
A governing architecture does not simply describe what should be built. It bounds what may be built. Implementation that conforms to the architecture is permissible. Implementation that violates or extends beyond the architecture is not—regardless of whether it runs.
This bounding function is the primary mechanism by which architecture governs probabilistic implementation. An AI agent that is guided by architecture may respect its boundaries. An AI agent whose implementation is bounded by a certified architecture cannot accept its own output as governed system state if that output violates the architecture. The acceptance decision is made by the governance system, not the implementing agent.
A governing architecture separates implementation from acceptance. Producing code is not the same as producing governed code. Acceptance requires that the implementation conform to the architecture, that its provenance is valid, that it passes the required verification, and that it be promoted through the repository into the governed baseline.
This separation is what makes the governing contract real. Without it, the architecture is advisory. With it, the architecture is enforced—not by asking the implementer to remember it, but by requiring conformance before acceptance.
The EPS evidence record supports a general model for governed AI-mediated software development. This model does not eliminate probabilistic implementation. It governs the conditions under which probabilistic implementation may produce authoritative system state.
The model has eight stages:
Intent. The purpose, requirements, and constraints of the work are established. This stage produces the governing intent that all subsequent stages serve.
Architecture authority. The architecture governing the implementation is established, singular, and certified before implementation authority is released. Competing authorities are resolved. The architecture is the contract.
Implementation contract. The bounded implementation scope, target behavior, regression scope, and acceptance criteria are specified. The contract makes the implementation verifiable rather than open-ended.
Bounded change. Implementation proceeds within the bounded scope. The implementing agent—human or AI—does not have authority to redefine the contract. Changes outside the bounded scope are not accepted.
Verification. The implementation is verified against the specified acceptance criteria. Focused verification establishes behavior within its scope. It does not establish end-to-end acceptance.
Certification. End-to-end verification establishes that the implementation conforms to the architecture within the certification scope. Certification is separate from focused verification. It produces a certification record.
Promotion. The certified implementation is promoted into the governed baseline through the repository. Promotion is a formal state transition, not a documentation step. The promoted state is authoritative.
Governed baseline. The promoted implementation becomes part of the frozen baseline. It is authoritative system state. Subsequent implementation is bounded relative to this baseline.
The governing sequence is:
Intent → Architecture Authority → Implementation Contract → Bounded Change → Verification → Certification → Promotion → Governed Baseline
The eight-stage model in Section 12 describes a workflow. But the reliability of governed implementation depends less on whether the workflow is followed and more on whether the state transitions are enforced.
A workflow can be bypassed. States cannot—if the state machine is correctly implemented.
In EPS, the architecture implements a state machine over implementation artifacts. An artifact in an unverified state cannot be certified. A certified artifact with unresolved observations cannot be closed. A closed artifact that has not been promoted through the repository is not governed baseline state. These transitions are enforced by the system, not by the engineer's judgment about whether the transition is appropriate.
The RC1 case (Section 4.1) illustrates what happens when the state machine is enforced: a runnable release candidate in a state that could not satisfy the repository-reproducibility requirement was held at the candidate state. It could not be promoted to release state because the state transition required repository-reproducibility, and that requirement was not satisfied.
The accepted runtime path required verification, certification, resolution of observations, closure, and promotion before baseline status. The rejected path produced runnable code that could not complete the required state transitions.
Most of the mechanisms described in this paper are not novel inventions. Architecture descriptions and architectural decisions are established software-engineering disciplines [1–3]. Configuration baselines and controlled change are established engineering controls [4]. Provenance has formal representation models [6]. Verification and disciplined software-development practices are also well-established parts of software assurance [5].
The stronger conclusion from the EPS case is that probabilistic implementation changes the operational importance of controls software engineering already understands.
A probabilistic AI agent resolves architectural ambiguity. When a specification is incomplete or an architectural decision is underdefined, the agent does not ask for clarification—it makes a decision, and the decision is plausible. The generated code runs. The resolution appears reasonable.
Plausible ambiguity resolution is more dangerous than an obvious error. An obvious error fails and is detected. A plausible ambiguity resolution succeeds locally and may not be detected until it has propagated through subsequent implementation.
The response is not to write better prompts. The response is to eliminate architectural ambiguity before implementation authority is released—as EPS-ARCH-010 demonstrated.
A human developer asked to implement the same change twice will generally produce similar results. A probabilistic AI agent asked to implement the same change twice may produce different results—both correct in isolation, different in their architectural implications.
This characteristic makes implementation verification more important, not less. If the same prompt does not guarantee the same implementation, the acceptance decision cannot rely on the implementation being identical to a previous accepted implementation. It must verify the current implementation against the current requirements.
AI agents express confidence. They generate explanations, justifications, and architectural rationale for their implementation choices. These expressions are not authoritative. A model that confidently implements an architectural violation is not less wrong because it expressed confidence.
In a governed development system, authority is not conferred by confidence. It is conferred by the architecture, the repository, and the acceptance record. An agent's justification for its implementation choices is neither an acceptance record nor an architectural authority.
The 14 supported episodes in the EPS evidence record support the following conclusions:
Architecture evolved from guidance into a governing contract in EPS. Earlier operating concepts relied on model-carried architectural intent. Later operating concepts enforced architectural authority through repository state, bounded baselines, and certified architecture as a precondition for implementation authority.
Local functional correctness was not sufficient for systemic acceptance. Five episodes provide direct evidence of implementations that were locally correct—functional, running, passing focused tests—but could not be accepted as governed system state because they violated systemic acceptance criteria.
Certification and closure are different states. The evidence record distinguishes them explicitly. Runtime integration was certified against 74 tests; closure required subsequent resolution of two documentation observations. Treating these as the same state would have miscounted program completion.
Bounded change is a governance mechanism that addresses the scope indeterminacy of probabilistic implementation. The S01 and S02 implementation baselines defined both target behavior and regression scope, making each accepted change verifiable and bounded.
Repository authority is a precondition for release, not a documentation convention. The RC1 case established that a runnable application is not releasable if its repository state cannot reproduce it.
The evidence does not establish that architecture eliminated defects in EPS.
The evidence does not support a causal claim that governed architecture reduces defect rates. This study did not design or execute a controlled comparison. The evidence record covers a single system. Defect rates before and after architectural governance were not measured under controlled conditions.
The evidence does not establish that the EPS architecture is optimal or universally applicable. The architectural decisions made in EPS reflect the requirements, constraints, and development context of that system. They are not presented as a template.
The evidence does not establish that the 14 supported episodes are representative of all EPS engineering work. They are the supported episodes in the reviewed record. The record may be incomplete.
The 74 runtime integration tests establish behavior within their certification scope. They do not establish universal system correctness or the absence of defects outside that scope.
The EPS evidence record suggests several practical implications for teams using AI agents for substantial implementation.
Architecture must be singular before implementation authority is released. Distributed architectural authority is not resolved by writing better prompts. It must be resolved at the architecture level, with explicit corrections documented and certified before implementation proceeds.
Focused tests do not establish systemic acceptance. Teams that rely on component-level tests to certify AI-generated implementation will encounter integration failures that the focused tests could not detect. End-to-end verification is a separate requirement with separate acceptance criteria.
Repository state is the authoritative record. Running code is not a substitute for reproducible repository state. Engineering teams using AI implementation should verify that the repository state can reproduce the running system—not as a documentation courtesy but as an acceptance condition.
Bounded implementation scope controls regression risk. Open-ended implementation requests to AI agents produce open-ended changes. Bounded implementation baselines specify scope, define regression coverage, and make accepted changes verifiable. The cost of defining the baseline is lower than the cost of discovering that the regression scope was insufficient.
Promotion gates protect baseline integrity. An artifact that has been generated is not governed system state. Promotion through the governed path—verification, certification, repository acceptance—is the mechanism by which generated artifacts become authoritative. Bypassing promotion degrades baseline integrity.
The dominant mental model for AI-assisted software development is supervision: a human developer reviews, accepts, or rejects AI-generated code on a per-change basis. The quality of the outcome depends primarily on the quality of the human review.
This model is workable at small scale. It does not scale reliably as AI-generated implementation becomes more substantial, more complex, and more structurally consequential.
EPS represents a different mental model: governed implementation. In this model, the quality of the outcome depends not primarily on reviewing each change but on surrounding every change with a governing system that constrains what may be produced, verifies what has been produced, and controls what the produced work becomes.
The human role in this model shifts from reviewing each implementation to governing the system within which implementation occurs. The governing system—architecture authority, bounded baselines, certification, promotion, repository state—provides reliability that does not depend on the consistency of per-change review.
This shift is not cost-free. Building and maintaining a governed implementation system requires engineering investment. The EPS evidence record suggests that this investment becomes more valuable as AI-generated implementation becomes more substantial—not because AI agents become less capable, but because the consequences of undetected architectural non-conformance become larger.
AI agents can produce substantial, functional, plausible software. They can pass tests, satisfy requirements, and implement changes that appear architecturally sound. And they can do all of this while violating the architecture of the system they are modifying.
The EPS evidence record documents this problem and the engineering response to it. Over 14 supported episodes, the governing mechanisms of EPS evolved from prompt-carried architectural intent to repository-enforced governing contracts. Five episodes demonstrated directly that local functional correctness was insufficient for systemic acceptance. One episode—EPS-ARCH-010—demonstrated that architecture itself must be governed before implementation authority can be released.
The evidence supports a conclusion that is narrow in scope and broad in implication:
In EPS, architecture became a governing contract. It assigned singular authority, bounded permissible implementation, and controlled acceptance. It did this not by expecting the implementing agent to remember it, but by requiring conformance before promotion.
The 74 passing tests, the bounded baselines, the certified architecture, the repository tags, the closure records—these are not documentation artifacts. They are the evidence that the governing contract was real.
When the engineer doing much of the implementation is probabilistic, reliable software depends not only on governing what the AI produces, but on governing how its work becomes part of the system.
NR-07 is a longitudinal, evidence-triangulated single-system engineering case study. The evidence record comprises 14 supported engineering episodes reconstructed from primary artifacts produced during EPS development.
Primary evidence sources
Architecture documents and requirements specifications; implementation records and revision logs; accepted architecture decision records (8 in consolidated set); tests and test results including 74-test runtime integration certification; certification and closure reports; release evidence; Git checkpoints and repository tags (9 observed); documented recovery incidents; bounded implementation baselines (S01, S02, H04).
Evidence classification
| Evidence class | What it establishes | What it does not establish |
|---|---|---|
| Historical reconstruction | Chronology, problem evolution, earlier operating concepts | Current implementation authority |
| Design and requirements | Intended architecture and engineering rationale | That the design was implemented |
| Repository implementation | Existence and scope of code/change | Complete end-to-end correctness |
| Tests | Behaviour covered by the stated tests | Behaviour outside test scope |
| Architecture certification | Conformance of architecture to defined requirements | Implementation completion |
| Implementation certification | Conformance within certification scope | Universal system correctness |
| Closure records | Formal completion of defined program obligations | Absence of future defects |
| Accepted baseline / tag | Authoritative repository state at that point | Independent proof of every runtime behaviour |
Limitations
This study does not establish a causal relationship between architectural governance and defect reduction. It does not claim that the EPS architecture is universally optimal. The 14 supported episodes represent the recoverable supported record; they are not a designed experimental sample. The 74 runtime integration tests establish behavior within their certification scope; they do not establish universal system correctness. The study covers a single system; findings may not generalize without further study.
The study was conducted by the same engineer responsible for EPS development. The evidence record is primary and the analysis is transparent, but independent replication is not yet available.
NR-03 examined evidence authority.
NR-04 examined repeated optimization.
NR-05 examined the authority to continue or stop optimization.
NR-06 examined the authority granted to historical experience.
NR-07 examines how architecture governs what probabilistic implementation may become.
What may the system treat as evidence?
How should the system improve an artifact?
When should it stop changing the artifact?
When should previous experience influence the next decision?
How does the architecture govern what the implementation may become?
Receive a link to the complete paper by email.
Steven Boyle is the founder of Northline Advisory, a technology advisory and research practice focused on technology leadership, enterprise transformation, governance, data and governed AI. His work draws on more than two decades of executive and operational experience across higher education and public-interest organizations.
How AI Systems Can Learn Without Turning Experience Into Truth
An evidence-grounded architecture for using decision history as an advisory signal without allowing experience or frequency to become factual authority.
Material Improvement and the Limits of Iterative Optimization
Historical and controlled evidence on material improvement, diminishing returns and the governance of stopping decisions in iterative AI optimization.
How an AI System Can Improve Its Work Without Losing Control
A governed optimization loop that lets AI propose and evaluate revisions while deterministic controls retain authority over continuation and promotion.