← Back to Research
NORTHLINE RESEARCH SERIES · GOVERNED AI SYSTEMS · NR-01

Beyond the Model

A Governed Architecture for Consequential AI Systems

Steven BoyleNorthline Advisory
|October 2026 · 18 min read
Abstract

Consequential uses of generative artificial intelligence require more than capable models and well-constructed prompts. This paper reconstructs the engineering evolution of a governed AI system and examines the architectural problems that emerged as model outputs moved from conversational assistance toward evidence-based, reviewable, and publishable artifacts. The reconstruction found that the recurring constraints were primarily architectural: establishing authoritative evidence, separating evidence selection from generation, preserving state and provenance, controlling artifact promotion, maintaining optimization lineage, and assigning explicit human decision rights. The resulting pattern places probabilistic model reasoning inside a deterministic system of record that governs evidence, artifact identity, validation, evaluation, promotion, and publication. The model remains responsible for interpretation and generation, but it does not independently determine what is authoritative or what becomes institutional record. The paper argues that increasingly capable models make this surrounding architecture more important, not less. The principal engineering implication is that trustworthy consequential AI depends on separating model capability from institutional authority and making that separation explicit, persistent, and traceable across the system lifecycle.

Get this research paper

Receive the complete paper as a PDF by email.

1. The Model Is Only Part of the System

Organizations adopting generative AI understandably spend considerable attention on model selection. Better models can improve reasoning, instruction following, synthesis and writing. Larger context windows can make more information available. Retrieval can supply information that was not present in model training.

Improvements in model capability are important, but established AI governance frameworks already treat trustworthy AI as a broader system and organizational concern. NIST's AI Risk Management Framework is intended for organizations designing, developing, deploying or using AI systems [1]. ISO/IEC 42001 similarly establishes requirements for an organizational AI management system that is maintained and continually improved over time [2].

None of these capabilities, by themselves, establish governance.

Consider a system that asks a language model to make a recommendation from a collection of organizational records. Even if the model produces an excellent answer, several questions remain.

Which records was it permitted to treat as evidence? Which version of each record was authoritative? What evidence actually supports the conclusion? Was the output generated, reviewed or approved? If the system produces another version tomorrow, does the new version supersede the old one? What happens if the new version scores better but introduces a factual regression? Who has authority to approve the result? What exactly does that approval make authoritative? If the system learns from the decision, what is it permitted to learn?

These are not primarily model questions.

They are architecture questions.

Generative-AI risk guidance provides related support for this concern. NIST identifies confabulation, information-integrity risk and problematic human-AI configurations including automation bias and over-reliance among risks organizations may need to manage when using generative AI [4]. Those risks reinforce the need for controls outside the generated answer itself.

That distinction became increasingly apparent while building EPS. The original objective was practical: construct a working system capable of using a substantial body of career evidence and opportunity information to reason about executive positioning and produce high-quality application materials. As the system developed, the difficult engineering problems increasingly appeared outside the language model itself.

The challenge was no longer simply getting an AI to produce a good answer. It was engineering a system in which good answers could be produced, evaluated, governed, recovered, approved and reused without losing control of their evidentiary basis.

This shift is consistent with broader governance frameworks that treat AI risk as a system and organizational concern rather than solely a property of model output [1,2]. In NR-01, the specific engineering response comes from the EPS case.

2. What Building a Working System Exposed

EPS evolved into an eleven-stage governed workflow, with seven stages currently exposed in its production interface. It processes opportunity information, strategic interpretation, governed career evidence, employment-role evidence, resume optimization, whole-document validation, cover-letter development and deterministic final publication.

The current runtime deliberately separates two kinds of work.

The probabilistic layer performs work for which natural-language judgment is valuable: interpreting a mandate, developing positioning, evaluating evidence, writing, reviewing and making semantic assessments.

The deterministic layer owns things that should not depend on probabilistic judgment: artifact identity, persistence, versioning, schema validation, authority transitions, source references, lineage, workflow state, checksums, publication and readiness.

That boundary became one of the most important architectural decisions in EPS.

It is tempting to put more responsibility into the model as models become more capable. But capability is not the only criterion for deciding where responsibility belongs.

A model may be capable of deciding which version of an artifact is current. That does not mean version authority should be probabilistic.

A model may be capable of reconstructing which evidence probably supports a statement. That does not mean provenance should be inferred after the fact.

A model may be capable of deciding whether a document is ready to publish. That does not mean publication authority should depend on an unrepeatable semantic judgment.

The architectural question is therefore not simply:

"Can the model do this?"

It is:

"Should this responsibility be probabilistic?"

EPS assigns natural-language interpretation to models where that capability creates value, while keeping identity, authority and lifecycle control outside the model.

The current system architecture formalizes this explicitly: model output is not authoritative merely because it exists. Authority is created through governed persistence, validation and explicit approval or promotion [6–8].

Diagram of the Northline Governed AI Architecture showing Evidence, Meaning, Reasoning, Output, Evaluation and Learning surrounded by governance and authority controls, human decision points and deterministic infrastructure.
Figure 1. The Northline Governed AI Architecture
The Northline governed AI architecture separates the core reasoning flow—Evidence → Meaning → Reasoning → Output → Evaluation → Learning—from the governance, human authority and deterministic infrastructure that control how information and authority move through the system.

3. From Prompts to Governed Artifacts

A conventional generative-AI interaction tends to revolve around messages:

input → prompt → response

That model is useful for conversation. It is less satisfactory for sustained consequential work.

The distinction is architectural rather than a criticism of conversational systems. ISO/IEC/IEEE 42010 treats architecture description as a system-level discipline involving architectural concepts and their relationships [3]. NR-01 applies that broader systems perspective to generative AI: the model may be central to the reasoning task without being sufficient to define the surrounding system of authority.

EPS instead treats important model outputs as artifacts with state.

An artifact has an identity. It can have versions. It has source relationships. It can be generated without being approved. It can be evaluated without becoming authoritative. It can be superseded without disappearing. Its relationship to downstream artifacts can be preserved.

This produces a critical distinction [6]:

"Generation creates a candidate. Validation determines admissibility. Promotion creates authority."

Those states should not be collapsed.

A generated artifact may be syntactically valid but factually deficient. A valid artifact may not have been approved. An approved artifact may no longer be the latest artifact. And the latest artifact may not be the best artifact.

EPS Stage 09 does not assemble a resume from whatever Stage 08 happened to generate most recently. It consumes promoted Stage 08 section versions. Stage 10 similarly consumes the promoted Stage 09 resume. Stage 11 consumes promoted upstream artifacts and performs deterministic publication rather than asking the model to reconstruct the final package.

Authority therefore moves through the system deliberately rather than following generation automatically.

Diagram showing the governed artifact lifecycle from evidence through generated, validated and evaluated artifacts, human review, promotion and downstream publication.
Figure 2. From Generation to Authority
The artifact lifecycle in a governed AI system. Generation creates a candidate, deterministic validation establishes admissibility, evaluation assesses the candidate, and explicit promotion creates the authoritative version consumed downstream.

4. Latest Is Not the Same as Best

Iterative AI introduces another problem.

Suppose an initial artifact is generated and evaluated. An evaluator identifies a material weakness, so the system authorizes a second version. The second version addresses the weakness but accidentally removes an important qualification or introduces a factual regression.

Which version is now authoritative?

A naïve iterative architecture often assumes:

newer = better

EPS does not.

Its optimization architecture distinguishes several states, including latest_round, selected_round, promoted_round and the best-valid state. Optimization lineage is also kept separate from display and downstream authority.

That permits a later round to remain part of the historical optimization trajectory while an earlier round remains the strongest valid artifact.

This is not merely theoretical. The Stage 08 governance tests explicitly cover a successor that satisfies its requested repair while introducing a new material regression. The candidate successor is prevented from replacing the incumbent even where other evaluation signals improve.

This leads to a broader principle:

"Optimization history and artifact authority are different things."

An AI system needs to remember what it tried without assuming that everything it tried became better.

That distinction becomes increasingly important as systems begin generating their own successors [8].

Diagram comparing an optimization lineage containing multiple generated rounds with a separate authority track in which the best-valid promoted round remains authoritative despite later iterations.
Figure 3. Optimization Lineage and Artifact Authority
Optimization lineage records what the system tried, while artifact authority determines what is trusted downstream. A later artifact may remain part of the decision history without superseding an earlier best-valid or promoted version.

5. Evaluation Must Be Separate from Generation

Another lesson emerged from iterative development: generation cannot also be treated as evidence of success.

EPS Stage 08 therefore separates writing from independent evaluation.

A writer receives governed evidence and produces a candidate-facing artifact. A separate evaluator examines that artifact against the relevant framework and produces review findings, scorecards, diagnostics and an optimization decision. Deterministic validation then determines whether the required evaluation artifacts actually exist and are valid.

The system does not treat writer completion as workflow completion.

This creates a more useful sequence:

governed evidence → generation → persistence → independent evaluation → deterministic validation → decision

The distinction matters because generative systems are very good at producing outputs that look complete. A coherent document is not evidence that the underlying task has been completed correctly.

EPS therefore requires evaluation-required stages to demonstrate more than file existence. Review-unavailable or fallback artifacts may be retained for diagnosis and recovery, but they do not become evidence that evaluation succeeded.

This is an important architectural distinction for organizations beginning to operationalize AI.

A workflow should not ask only:

"Did the model return something?"

It should ask:

"What evidence establishes that the result satisfies the conditions required for the workflow to continue?"

6. Human Oversight Needs an Architectural Role

"Human in the loop" is frequently proposed as the answer to AI governance.

The phrase is directionally useful but architecturally incomplete.

A human can be placed almost anywhere in a workflow without it being clear what authority that person actually exercises. A user clicking "approve" is not meaningful governance unless the system defines what is being approved, which version is being approved, what evidence it depends upon, and what changes because approval occurred.

A better design question is:

"Which authority transition requires human judgment, and what exactly does that approval make authoritative?"

EPS uses explicit human approval gates before important authority transitions. The human is not expected to manually perform every operation or verify every deterministic relationship. Instead, human judgment is concentrated where judgment changes authority.

For example, promoting a resume section does not simply indicate that the user "likes" it. Promotion makes an exact validated version authoritative for downstream assembly.

That creates a much clearer relationship between automation and accountability.

The model can reason.

The framework can validate.

The system can maintain lineage.

But where the workflow requires human authority, the architecture must preserve that decision as an identifiable transition rather than an informal interaction.

NIST's Generative AI Profile identifies human-AI configuration risks including automation bias and over-reliance [4]. NR-01's response is architectural: human review is connected to identifiable evidence, artifacts, versions and decisions rather than treated as generic supervision of model output.

Broader AI governance frameworks place risk-management and accountability responsibilities around AI systems rather than assigning institutional governance to the model itself [1,2]. NR-01 makes that boundary operational through explicit rules for evidence authority, review, promotion and publication.

7. Evidence Must Remain Evidence After Reasoning Begins

One of the more difficult problems in AI systems is maintaining the relationship between evidence and conclusions as information moves through multiple reasoning stages.

Providing a model with evidence is not the same as governing evidence.

Once information is summarized, interpreted, rewritten and combined with other information, its origin can become increasingly difficult to reconstruct. A downstream model may receive an apparently authoritative statement that is actually an interpretation produced several stages earlier.

Provenance provides an established technical foundation for preserving information about origin and derivation. W3C PROV provides a formal model for representing and interchanging provenance information across systems and contexts [5]. NR-01 applies that concern to AI authority: generated interpretation may operate over governed evidence, but generation should not silently become a new factual source.

EPS addresses this by separating factual authority from interpretive authority [6].

Stage 07, for example, owns factual and evidence authority for employment evidence. Stage 08 can improve candidate-facing expression, emphasis and structure, but it cannot create new candidate facts simply because those facts would strengthen the document.

Later deterministic diagnostics extend this further [7].

The implemented Phase 2 graph-diagnostics service computes proof coverage, evidence reuse, complementary proof, evidence availability, selection traceability and explicit statement traceability from governed identifiers and declared relationships.

Importantly, it makes no model call.

Where explicit support relationships do not exist, the service reports them as unavailable rather than asking an LLM to reconstruct them from semantic similarity.

That is a deliberate architectural choice.

An AI model may be able to infer that a sentence probably came from a particular source. But inferred provenance is not the same thing as recorded provenance.

For consequential workflows, that distinction matters.

8. A Six-Part Architecture for Governed AI

The lessons from EPS can be generalized beyond executive positioning.

Existing standards and frameworks already address important parts of this problem: architecture description [3], AI risk management [1], organizational AI management [2], generative-AI risk [4] and provenance [5]. NR-01's proposition is not that these disciplines are absent. It is that a generative-AI implementation can still remain model-centric unless those responsibilities are assembled into an operational architecture around the model.

The resulting Northline model is:

Evidence → Meaning → Reasoning → Output → Evaluation → Learning

These are not simply processing stages. Each represents a different governance problem.

Evidence

What is the system allowed to know?

Evidence governance concerns source authority, provenance, versions, availability, identity and the distinction between recorded fact and derived interpretation.

Meaning

What does the evidence mean in the context of the decision?

This is where semantic interpretation becomes necessary. Models are particularly valuable here because many consequential problems cannot be reduced to deterministic rules.

But interpretation should remain distinguishable from its source evidence.

Reasoning

What conclusions or strategies follow from that interpretation?

Reasoning may be probabilistic. Its inputs, permitted scope and resulting artifacts should not be uncontrolled.

Output

What did the system actually produce?

Outputs need identity, lifecycle state and lineage. They should not become authoritative simply because generation completed.

Evaluation

Is the output sufficiently good, supported and safe to advance?

Evaluation should be distinguishable from generation, and evaluation itself does not necessarily create authority.

Learning

What, if anything, should this experience change?

This is the most dangerous transition to leave implicit.

An observed outcome is not automatically knowledge. A user preference is not automatically a universal rule. A successful decision in one context is not automatically evidence that the same decision should be made in another.

A governed learning system therefore needs another progression:

evidence → observation → inference → governed knowledge

Each transition requires stronger justification, not weaker controls.

Formal provenance models provide precedent for preserving origin and derivation as information is reused across systems [5]. In EPS, curation therefore reorganizes governed knowledge without allowing prior generated documents to become new factual authority.

EPS has begun implementing this distinction in its longitudinal-learning architecture, but the broader learning problem remains an active research area and is treated separately later in this series.

9. Governance Does Not Mean Freezing the System

There is an understandable concern that adding controls around AI will make systems rigid, slow or prohibitively expensive.

The EPS development history suggests a different interpretation.

The first objective was to establish a working governed system. Early implementations prioritized end-to-end correctness, evidence integrity, recoverability, traceability and authority over computational efficiency.

That resulted in real overhead.

Only once the workflow was functioning could its computational behaviour be observed meaningfully. ResearchCapture was introduced to record model operations, source state, framework execution, lineage and runtime information. Subsequent telemetry exposed duplicated context, expensive review operations, unnecessary repeated work, recovery inefficiencies and places where iteration needed tighter bounds.

Optimization then became an engineering problem supported by evidence.

Changes have included tighter model payloads, bounded review contracts, limits on successor generation, deduplication, improved recovery and more compact evaluation. Operational execution that once took hours has been reduced to minutes, while the current telemetry continues to identify further opportunities for reducing tokens and latency.

The sequence matters:

Make it work → Make it governable → Instrument it → Understand it → Optimize it → Measure again

This is different from optimizing prompts before the behaviour of the larger system is understood.

It also points to an important distinction between governance overhead and architectural inefficiency. They are not the same thing. Some controls are necessary to preserve authority and integrity. Other computational work may add little or no value.

Instrumentation is what allows the two to be separated.

The current EPS telemetry should therefore not be interpreted as the computational profile of a finished optimized architecture. It is evidence from a working system that has entered an active optimization phase.

Diagram showing the progression from a working system to a governed system, instrumentation, observation, selective optimization and repeated measurement, with continuous improvement feeding subsequent cycles.
Figure 4. From Working System to Measured Optimization
The EPS engineering sequence prioritized a functioning governed system before computational optimization. Instrumentation and telemetry then made cost and value visible, allowing targeted optimization while preserving evidence integrity, evaluation and authority controls. End-to-end execution has moved from hours to minutes, with further optimization opportunities remaining.

10. Observability Changes What Can Be Governed

ResearchCapture emerged partly from this need.

For each run, EPS can record architecture identity, Git state, source snapshots, model-call metadata, request and response hashes, token usage, generated artifacts, framework evaluation records and decision-lineage events.

At publication, the research source snapshot can be frozen.

This creates the possibility of reconstructing not merely what an AI system produced, but the conditions under which it produced it.

Conceptually:

source state → model operation → generated artifact → evaluation → governed decision → successor / retention / promotion / stop

This matters for both engineering and research.

Without that information, a team can know that its AI system is expensive but not necessarily why. It can know that outputs changed but not which intervention caused the change. It can know that a later artifact exists but not whether it became authoritative.

Observability therefore becomes more than an operational concern.

In a governed AI system, observability is part of accountability.

11. What This Means for an Organization Starting Now

An organization does not need an eleven-stage system or the EPS architecture to apply these principles.

But it should resist starting with prompts alone.

Organizational AI governance likewise depends on persistent processes rather than one-off interactions. ISO/IEC 42001 requires an AI management system to be established, implemented, maintained and continually improved [2]. NR-01 addresses the corresponding software-architecture question: consequential AI decisions require state that survives independently of an individual model exchange.

Before allowing AI to participate in consequential work, the more useful architectural questions are about boundaries.

What information constitutes authoritative evidence? Which decisions can be probabilistic? Which state transitions must remain deterministic? What is a draft? What makes something valid? What makes it approved? Which exact version becomes authoritative? How can a downstream process prove which upstream version it consumed? How are model outputs independently evaluated? What happens when a later output is worse? What happens when a model call fails halfway through? Which decisions require human authority? What can the system learn from prior work, and what must it never infer automatically?

The answers will differ by organization and risk level.

The architectural requirement does not.

Consequential AI needs an explicit system of authority around probabilistic reasoning.

This does not require removing judgment from the model. It requires being deliberate about where model judgment begins and ends.

12. What Remains Unresolved

EPS provides an implemented example, not proof that one architecture is universally correct.

Several questions remain open.

The appropriate balance between deterministic and probabilistic control will vary by domain. The amount of provenance required for an executive document is not necessarily sufficient for medicine, finance or public administration.

Independent model evaluation also creates its own problems. Evaluators are probabilistic. Their judgments can vary, and additional evaluation consumes computation. The presence of an evaluator should therefore not be mistaken for objective truth.

Similarly, iterative optimization raises an unresolved economic question: when does another round of reasoning create enough value to justify its cost? EPS now has a research harness capable of challenging continuation and stopping decisions, but the empirical evidence should be accumulated before making broad claims about optimal stopping behaviour.

Learning introduces still greater difficulty. A governed system can record decisions and outcomes without being entitled to turn every observation into future authority. Determining when repeated experience becomes sufficiently reliable to influence subsequent decisions remains a significant research problem.

These limitations are not arguments against governed AI.

They are reasons to make the boundaries visible.

Conclusion

Building EPS began as an effort to make AI perform a difficult piece of work reliably. It ultimately exposed a larger architectural problem.

The most consequential failures were not necessarily failures of model intelligence. They could occur when the right evidence was present but its authority was unclear; when a good output existed but its lifecycle state was ambiguous; when a later version was assumed to be better; when evaluation and generation became conflated; when approval lacked an explicit authority effect; or when experience risked becoming "knowledge" without sufficient justification.

Those problems cannot be solved by a better prompt.

Nor are they solved simply by choosing a better model.

A governed AI system needs an architecture capable of deciding what the model is allowed to reason from, preserving what happened during that reasoning, evaluating what was produced, controlling what becomes authoritative, and governing what the system is permitted to learn from experience.

Established standards already make clear that trustworthy AI depends on more than raw model capability: AI risk must be managed across the lifecycle [1], organizations need persistent AI-management processes [2], architecture is a system-level concern [3], generative AI presents information-integrity and human-over-reliance risks [4], and provenance can be represented explicitly [5].

NR-01's contribution is the engineering synthesis derived from EPS: the model can remain responsible for interpretation, comparison, reasoning and generation while evidence authority, persistent state, deterministic orchestration, human approval and publication authority remain properties of the surrounding system.

The resulting principle is simple:

"Model capability is not system capability."

For consequential AI, the system surrounding the model increasingly determines whether intelligence can be used reliably.

And that shifts the engineering question.

The question is no longer only "how capable can the model become?"

It is also:

"How do we design the system around probabilistic intelligence so that evidence, judgment, authority and accountability remain connected?"

That is the architecture problem governed AI must solve.

Research Status and Evidence Basis

This paper derives its architectural argument from the implemented EPS system rather than treating EPS as a controlled validation of all claims. Current architecture, artifact authority, lifecycle and deterministic/probabilistic boundaries are documented in EPS_ARCHITECTURE.md and docs/architecture/EPS_Governed_V1.0_Architecture.md. Iterative optimization and best-valid behaviour are governed by docs/contracts/v1.0/Stage08_Contract.md and supporting implementation and tests. Deterministic evidence diagnostics are implemented under docs/architecture/EPS_Phase_2_Graph_Diagnostics_Implementation_Specification.md. Research observability is implemented in eps_api_dev/research_capture.py. The production telemetry corpus is stored outside Git under gitignored runtime directories; quantitative telemetry findings should therefore be cited to the preserved telemetry analysis rather than reconstructed from the source repository. Repository references for this paper were checked against the codex/eps-protected-enhancements-20260908 line through commit 651612d8239c85431ffaf8004e103609add0e3a6.

References
[1] Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. National Institute of Standards and Technology, 2023. DOI: 10.6028/NIST.AI.100-1.
[2] ISO/IEC. ISO/IEC 42001:2023 — Information technology — Artificial intelligence — Management system. International Organization for Standardization, 2023.
[3] ISO/IEC/IEEE. ISO/IEC/IEEE 42010:2022 — Software, systems and enterprise — Architecture description. International Organization for Standardization, 2022.
[4] Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., and Roberts, K. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. National Institute of Standards and Technology, 2024. DOI: 10.6028/NIST.AI.600-1.
[5] Lebo, T., Sahoo, S., and McGuinness, D., eds. "PROV-O: The PROV Ontology." W3C Recommendation, 30 April 2013.
[6] Boyle, S. Artifact-Oriented AI Architecture: Governing State, Authority and Lineage. Northline Research Series: Governed AI Systems, NR-02. Northline Advisory, 2026.
[7] Boyle, S. Evidence Coverage and Traceability: Knowing What Supports a Conclusion. Northline Research Series: Governed AI Systems, NR-03. Northline Advisory, 2026.
[8] Boyle, S. Governed Self-Optimization: How an AI System Can Improve Its Work Without Losing Control. Northline Research Series: Governed AI Systems, NR-04. Northline Advisory, 2026.
Next in the Series NR-02 Artifact-Oriented AI Architecture Governing State, Authority and Lineage Read NR-02 →
Want a copy of this paper?

Receive a link to the complete paper by email.

Steven Boyle
Northline Advisory

Steven Boyle is the founder of Northline Advisory, a technology advisory and research practice focused on technology leadership, enterprise transformation, governance, data and governed AI. His work draws on more than two decades of executive and operational experience across higher education and public-interest organizations.

Related Practitioner Insights
Practitioner Insight · Governance

Artifact-Oriented AI Architecture: Closing the Enterprise AI Architecture Gap

An architectural pattern for high-consequence enterprise AI systems in which probabilistic reasoning operates within deterministic governance and authority resides in governed artifacts rather than generated responses.

Steven BoyleAugust 4, 2026 · 7 min readRead →
Practitioner Insight · Governance

Governed AI with Gated Learning: Lessons from Building EPS

A governed architecture pattern for introducing learning into consequential AI workflows, illustrated through the development of EPS.

Steven BoyleSeptember 11, 2026 · 8 min readRead →
Practitioner Insight · Governance

From Decision History to Governed Intelligence

How governed AI systems can learn from decision history without allowing accumulated experience to silently become authority.

Steven BoyleSeptember 18, 2026 · 6 min readRead →