Investigating how consequential technology decisions can become more governable.
Northline Research investigates the architecture, governance and engineering of AI-enabled systems used for consequential work.
The programme combines system architecture, implementation, deterministic controls, testing, production telemetry and controlled experimentation to examine how evidence, reasoning, authority, evaluation and learning can remain governable as AI systems become more capable.
Research authored by Steven Boyle, Northline Advisory · ORCID 0009-0006-1607-7362
Steven Boyle's research and publication record spans engineering systems, data and decision support, digital transformation, and governed AI systems, with peer-reviewed work dating to the 1990s.
Governed AI Systems
Northline's initial research programme examines a central engineering problem: how AI systems can perform increasingly complex reasoning and optimization while evidence, authority, evaluation, learning and accountability remain governed.
The programme is grounded in the architecture and implementation of EPS, a working governed AI system, and progressively extends from system architecture into traceability, optimization, experimental evaluation, learning and AI-assisted software engineering.
Beyond the Model
A Governed Architecture for Consequential AI Systems
Consequential AI requires an architecture that governs evidence, artifacts, evaluation, authority and learning around the model.

How do we engineer systems in which good decisions can reliably occur?
The individual papers examine different parts of that problem. Together they investigate how evidence becomes usable meaning, how reasoning remains connected to its sources, how outputs acquire authority, how quality is evaluated, when optimization should stop, and what a system should be permitted to learn from experience.
What evidence informed an output?
How did reasoning connect evidence to the conclusion?
How was the output evaluated?
What is the system permitted to learn from the result?
Governed AI Research Model
The Northline research model follows the movement from evidence through interpretation, reasoning, output, evaluation and learning. Governance surrounds that flow by determining what information is authoritative, what transitions are permitted and where human judgment remains required.
into Evidence
What information is available?
What information is available, its provenance, quality and relevance.
What does the evidence mean in context?
How evidence is interpreted within the context of the decision.
How are conclusions formed?
How the system connects evidence, requirements and constraints.
What does the system produce?
What the system ultimately recommends, produces or communicates.
How do we know whether it is good enough?
How the quality, integrity and completeness of that output are assessed.
How should the system improve?
How evaluation informs subsequent improvement without allowing uncontrolled optimization.
Evidence
↓What information is available?
What information is available, its provenance, quality and relevance.
Meaning
↓What does the evidence mean in context?
How evidence is interpreted within the context of the decision.
Reasoning
↓How are conclusions formed?
How the system connects evidence, requirements and constraints.
Output
↓What does the system produce?
What the system ultimately recommends, produces or communicates.
Evaluation
↓How do we know whether it is good enough?
How the quality, integrity and completeness of that output are assessed.
Learning
↩ feeds backHow should the system improve?
How evaluation informs subsequent improvement without allowing uncontrolled optimization.
EPS provides a system in which these questions can be implemented and tested.
Northline's research is grounded in the development of EPS, a working governed AI system originally created to investigate executive positioning and application development from a substantial body of structured career evidence.
The value of EPS as a research environment is not the application domain itself. It is that the system creates a sufficiently complex setting in which probabilistic reasoning, governed evidence, deterministic controls, human authority, iterative optimization, evaluation and learning must operate together.
Research findings are drawn from different forms of evidence. Those forms should not be treated as interchangeable.
Architecture
What the system is designed and governed to do.
Implementation & Tests
What the software, contracts, schemas and automated tests actually enforce.
Production Telemetry
What the running system actually does during operation, including model calls, optimization trajectories, costs, timing, retries and governed outcomes.
Research Harness & Experiments
What happens when specific system decisions or architectural hypotheses are deliberately challenged under controlled conditions.
Architecture tells us what the system is designed to do. Implementation and tests tell us what is enforced. Production telemetry tells us what happens in use. The research harness lets us deliberately challenge the system's decisions.
Claims should not outrun evidence.
Northline distinguishes between what has been designed, what has been implemented, what has been observed, what has been experimentally measured and what remains a proposition. This distinction is important because a plausible architectural idea is not the same thing as an implemented control, and an operational observation is not the same thing as an experimental finding.
Implemented
Demonstrated in the current architecture, contracts, schemas, code or tests.
Observed
Seen in operational behaviour or production telemetry.
Experimentally Measured
Established through a defined experiment or controlled research harness.
Proposed
An architectural proposition, extension or hypothesis that has not yet been demonstrated.
Unresolved
A question for which the available evidence is insufficient to support a conclusion.
The purpose is not to force every question into an experiment. It is to make the evidentiary basis of each research claim explicit.
Research should improve the way consequential technology decisions are made.
Northline's research is intended to move in both directions. Practical problems expose questions worth investigating. Architecture, engineering and experimentation make those questions testable. The resulting evidence can then inform advisory work, organizational practice and subsequent research.
Better AI requires more than better models.
The challenge is to engineer the surrounding system so that evidence, reasoning, authority, evaluation and learning remain governable as model capability increases.
Northline Research is working on that problem.

