Article · Governance

Governed AI with Gated Learning: Lessons from Building EPS

A governed architecture pattern for introducing learning into consequential AI workflows, illustrated through the development of EPS.

Steven BoyleNorthline Advisory
|September 11, 2026 · 8 min read
Governed AI with Gated Learning: Lessons from Building EPS

This article presents a governed architecture pattern for introducing learning into consequential AI workflows, illustrated through the development of EPS.

Building EPS has made me pay closer attention to what happens after an AI model produces an answer. A draft can read well and still overstate the evidence. A review can recommend changes without improving the result. A useful decision in one case can become a poor assumption in the next.

These problems have shaped how I think about AI architecture, particularly as systems begin to retain experience and use it in later work. The architecture needs to establish which information the system can rely on, who can approve a result, and under what conditions past experience may influence a new decision.

EPS is an executive positioning and application system. It brings together a job opportunity, a candidate's career evidence, and human review to develop a resume and cover letter. The task is familiar: understand what the employer needs and explain the candidate's relevant experience clearly and accurately.

That makes EPS a useful example of a broader engineering problem. The system must interpret information, make recommendations, and improve its writing while preserving the facts it was given. Adding learning introduces another responsibility: deciding which parts of previous work deserve to influence future work.

I describe this approach as governed architecture with gated learning. Governance establishes who or what has authority to change information and approve its use. Gated learning applies that discipline to information retained from experience.

The building blocks have established precedents. Source traceability is formalised in standards such as W3C PROV. Frameworks such as LangGraph document checkpoints, human review, and recovery from interrupted execution. EPS provides a practical setting in which to apply these ideas and examine the decisions they require.

The first decision concerns evidence. A candidate's verified career record and an employer's requirements both matter, but they serve different purposes. The career record establishes what the candidate has done. The job description establishes what the employer is looking for. The model's job is to explain the relationship between them

Candidate claims begin with verifiable evidence. Employer requirements establish relevance; they do not supply facts about the candidate.
Candidate claims begin with verifiable evidence. Employer requirements establish relevance; they do not supply facts about the candidate.

Consider a hypothetical candidate who established governance committees and clarified decision rights within an organisation. A new opportunity involves AI governance. The earlier work may provide relevant evidence of the candidate's ability to establish accountability and put governance into practice. It does not, by itself, establish that the candidate has governed AI systems.

A useful system should be able to make that connection and explain its limits. Restricting it to matching identical words would miss relevant experience. Allowing it to turn related experience into a claim of direct experience would misrepresent the candidate. The engineering task is to support reasoning while keeping the resulting claims grounded in their sources.

This requires a clear division of responsibility. The model interprets the opportunity, assesses relevance, and proposes language. Application software manages records, versions, workflow controls, and approval status. People review the resulting argument and decide whether it represents them accurately.

Each has limits. Software can check whether a referenced record exists and whether a version was approved. Those checks do not establish that a paragraph is persuasive. A model can explain why an achievement appears relevant, but that explanation cannot authorise a change to the underlying facts. Human review provides judgment and accountability, although approval itself is not proof that every claim is correct.

EPS's current workflow reflects this separation through governed sources, review, and approval. A generated draft remains distinguishable from a version accepted for subsequent work. Retaining that distinction makes it possible to revisit a decision without losing the record of what was previously approved.

It also gives governance a practical purpose. A user should be able to identify the version they reviewed and the material used to produce it. If work is interrupted, the recovery process should preserve completed work and make the remaining task clear. These are ordinary expectations of dependable software, including software that uses probabilistic models.

EPS's governed production workflow and the proposed selection-learning extension. The current knowledge backbone supports evidence relationships; learning from selection history remains a future capability.
EPS's governed production workflow and the proposed selection-learning extension. The current knowledge backbone supports evidence relationships; learning from selection history remains a future capability.

Review introduces its own engineering problem. During EPS development, repeated review sometimes led to further rewriting with little visible improvement. A model could identify something else to change, even when the existing text already made the relevant point. The additional work consumed time and tokens without necessarily producing a stronger application.

That experience changed how I think about optimisation. Feedback is useful when it identifies a substantive weakness. Another score or a preference for different wording is weaker evidence of progress. An improvement process needs a reason to continue and a way to recognise when a later version has lost something valuable.

EPS already applies a feedback-driven optimization process. Before another round, it estimates the possible improvement and the risk of regression. After the round is reviewed, it compares the result with the earlier prediction. That outcome becomes part of the next decision. The system can continue with a targeted repair, stop when further improvement appears unlikely, or retain an earlier valid version.

This resembles a problem studied in rational metareasoning: whether another computation is likely to improve the final decision enough to justify performing it. Research on the value of computation provides a useful scientific frame for this question. EPS currently applies a governed heuristic controller rather than a formal value-of-computation algorithm. Its estimates guide a bounded workflow; they are not learned probabilities.

The same discipline applies to learning. A system's history contains much more than reliable examples. It includes incomplete drafts, rejected suggestions, misunderstandings, and decisions made under circumstances that may no longer apply. Storing that history is useful. Treating all of it as equally credible would create a new source of error.

A knowledge graph can represent the relationships that give an achievement meaning. It occurred within a particular role, involved particular responsibilities, and may demonstrate several capabilities. Its relevance to an opportunity is an interpretation of those facts. The graph can preserve the distinction between the source information and the conclusions drawn from it.

For EPS, the proposed learning extension would use experience from previous applications to inform future recommendations. An achievement that proved useful in explaining governance experience might deserve consideration for a later opportunity. The system would still need to establish why it matters in that new context.

This is also related to learning to rank. Previous selections, changes in order, and human decisions could help EPS recommend evidence for a later opportunity. Those observations are biased by what the system presented and what the user had an opportunity to review. Research on learning to rank from biased feedback shows why that bias has to be addressed before historical choices are treated as training evidence.

There is a distinction between learning that something was selected and learning that it was effective. A user may approve a passage because it is accurate and clear. That does not show that the passage caused an interview invitation. Even an interview invitation would not isolate the contribution of that passage from the candidate's wider experience, the applicant pool, or the employer's preferences.

A learning system therefore needs to preserve what its observations actually establish. Repeated use can suggest a pattern worth examining. It cannot, on its own, establish relevance, quality, or success. Otherwise, familiar choices may become more influential simply because the system keeps repeating them.

Gated learning places a boundary around that influence. Before retained experience affects a recommendation, its relevance and reliability need to be assessed. It must remain open to correction, and the current task must still receive its own analysis. Historical information can help the system make a judgment while remaining subject to the same evidence and approval requirements as other inputs.

This is the proposed direction for EPS. The application already uses governed evidence, AI reasoning, review, and approval. The extension that would make selection experience reusable across applications remains a design proposal. Its expected benefits include less repeated analysis and more useful recommendations, but those benefits need to be demonstrated through evaluation.

Further ahead, the accumulated prediction and outcome records could support a learned decision policy. In reinforcement-learning terms, the action might be to stop, revise, or retain an earlier version. The reward would have to represent more than a reviewer score. It would need to account for whether a material problem was resolved, whether evidence integrity was preserved, whether the user accepted the result, and what the additional work cost. Sutton and Barto's treatment of reinforcement learning provides the established foundation for thinking about policies, actions, and rewards.

Distributional reinforcement learning suggests another possible direction. Instead of estimating only an average return from another action, a system can model a distribution of possible returns. Bellemare, Dabney and Munos showed why that distribution can matter in reinforcement learning. For EPS, the practical interest would be uncertainty and asymmetric risk: a revision may offer a modest chance of improvement while carrying a smaller but more consequential chance of losing evidence, clarity, or an already strong argument.

These are research directions rather than descriptions of the current optimizer. EPS does not yet learn its decision policy from rewards or apply a distributional reinforcement-learning algorithm. The immediate work is to build reliable, governed experience across applications. A learned policy becomes credible only when its observations, rewards, and evaluation method are strong enough to support it.

Keeping that distinction clear is part of the engineering discipline. A graph does not establish that a system learns effectively. The value has to appear in the quality of subsequent work and in the effort required to produce it.

EPS has given me a practical way to examine these questions because the result is easy to inspect. A reader can compare the application with the evidence, judge whether the argument is clear, and decide whether another revision helped. The same questions apply when AI systems prepare reports, assess proposals, or support operational decisions.

As those systems retain more experience, their architecture needs to preserve the difference between what happened, what was inferred, and what was approved. Future recommendations should remain explainable against the current evidence and open to human correction. That is the standard I am applying as I develop EPS's learning capabilities.

Originally published on LinkedIn on September 11, 2026.

Steven Boyle
Northline Advisory

Steven Boyle is the founder of Northline Advisory, a technology advisory and research practice focused on technology leadership, enterprise transformation, governance, data and governed AI. His work draws on more than two decades of executive and operational experience across higher education and public-interest organizations.