From factory to intelligence : The evidence plane, and the asset that compounds

SAMI
October 11, 2026 10 mins to read
Share

Post 4 of 8 in the series on rebuilding a bank’s software development factory for the agentic era.

Here is the architecture. Five layers, two cross-cutting planes, and one workstream that decides whether any of the rest of it works.

I will start at the bottom, because the bottom is the part that gets cut first when the programme comes under cost pressure, and cutting it is the single most reliable way to fail at everything above.

Layer 0, the evidence plane

This layer makes the factory’s own work machine-readable.

Decisions captured as structured records linked to the work they govern, which in practice means lightweight architecture decision records in git rather than minutes in SharePoint. Specifications as versioned artifacts with explicit deltas, so that what changed about a requirement is itself a reviewable object. Requirements traceable through tasks to commits to deployments. The obligations register expressed as data rather than as a spreadsheet maintained by one person in compliance. Meeting outcomes captured as linked decisions with owners.

Every transformation skips this. It looks like documentation, and documentation never gets funded.

The way to keep it funded is to refuse to describe it that way, because it is not documentation. It is the substrate every layer above consumes, and its business case is written entirely in the currency of those layers. Without it, agents work blind. The delivery world model models artifacts rather than decisions. Evidence generation stays a manual exercise performed three weeks before every review by someone assembling screenshots into a folder.

At XYZ this was roughly 1,400 person-hours a year of pure archaeology. That is the number I would put in the funding paper, not a paragraph about knowledge management.

The 2024 factory had a paper prototype of this plane: the FactSheet documents linking projects to business outcomes. The 2026 version is the same intent with a schema, an API, and a freshness target.

Layer 1, capabilities

The engineering primitives a change gets composed from. Golden paths and paved road templates. Reusable domain services. Controls as code, meaning policy checks that execute in the pipeline rather than living in a reviewer’s head. Synthetic test data provisioning that yields clean client records on demand. Ephemeral environments. Evidence generation that emits the audit artifact as a build output.

Block governs its financial primitives with reliability, compliance and performance targets, and that bar transfers without edits. Each capability has a named owner and stated objectives, which is what MaRisk AT 7.2 already implies for IT systems supporting material processes.

What changes from 2024 is the consumer. Back then, the consumer of a capability was a developer who could route around a rough edge by asking a colleague at the next desk. Now the consumer is increasingly an agent, and an agent cannot compensate for anything it cannot read.

So the test I apply throughout: a capability an agent cannot discover, invoke and verify through a machine-readable contract is not yet a capability. Agent-consumable or it does not count.

The measure of this layer is not how many templates exist. It is what fraction of changes are composed entirely from paved road elements, because that fraction is the fraction where autonomy is even discussable.

Layer 2, the two world models

These are different objects with different freshness requirements, and conflating them causes real design errors.

The delivery world model is how the factory understands itself. Flow, dependencies, capacity, quality, change risk, control posture. What a delivery manager used to hold in their head and relay in a status meeting, held instead by a system with read access to the repositories, the pipelines, the boards, the change and incident record, and the configuration database. It is worthless if it is a day stale.

It is also the object with the sharpest constraint on it, for the works council and AI Act reasons I set out in post three. Scoped to answer questions about flow and systems and controls. Demonstrably incapable of producing individual performance assessments. Signed off before the first index is built.

The domain world model is how the factory understands the bank. Business domains, system topology, data lineage, the obligations register, and the recovered semantics of the legacy core. Parts of it are valuable even when a year old, provided the age of each fact is visible.

I said in post one that I would answer Block’s framing question at the end of the series: what does your organisation understand that is genuinely hard to understand, and is that understanding getting deeper every day. This layer is where my answer lives. I will defend it properly in the final post, when the people side of it is on the table too, because the answer turns out to be about people as much as about data.

Layer 3, the intelligence layer

The agent fleet and its orchestration: planning, implementation, review, test generation, migration, incident triage, control evidence.

Plus four things that make agents survivable in a bank, which are what most of the remaining posts are about.

Context engineering, meaning a context file per repository, skills that encode the organisation’s recurring task shapes, and MCP servers exposing internal systems as tools rather than as screenshots pasted into a chat.

Evaluation harnesses that measure agent output against ground truth on your own code, before autonomy expands anywhere.

Guardrails, including the ones that stop an agent reaching production data.

And gate tiering, which belongs to the plane below.

Layer 4, interfaces

IDE, pull request, ticket, chat, terminal, dashboard. Where people meet the system.

Necessary, worth investing in, and explicitly not where value is created. That has a budgeting consequence worth saying out loud: when the programme comes under cost pressure, the instinct will be to cut Layer 0 and keep building dashboards, because dashboards demo well to a steering committee and an obligations register does not. That instinct is the failure mode, not the saving.

The human edge

The framework I started from had five layers and an assurance plane, and it treated humans as implicit inhabitants of the intelligence layer, present as human-in-the-loop gates inside Layer 3.

I think that is backwards for a bank, and it is the one correction I made to the design.

In this environment the placement of humans is a control design in its own right. Reviewed by governance, evidenced to internal audit and to supervisors, negotiated with works councils, and required to stay stable across model and vendor changes. Bury it inside the intelligence layer and it gets configured rather than designed, and configuration drifts.

So the human edge is an explicit cross-cutting plane: named roles, decision rights, and the map assigning every change class to decide, sample, or inform. Making it a first-class object forces “who signs, and where is that recorded” to be answered at design time instead of discovered during an inspection. It also gives compliance and the works council one artifact to review rather than a hunt through pipeline YAML.

The assurance plane

Model risk management applied to the factory’s own AI use, aligned with the institution’s existing model governance rather than invented beside it.

AI Act conformity tracking for two distinct populations: the factory’s own tools, and the AI systems the factory builds for the business. Different articles, different signatures, and conflating them is a governance error that surfaces at the worst possible moment.

DORA ICT risk integration, including the register of information under Art. 28(3) and concentration analysis under Art. 29 once agent capacity depends on one model provider.

The immutable audit trail. Named accountability. Load-bearing, not bolted on.

The roadmap inversion, wired in

Block’s most elegant idea is that when the intelligence layer cannot compose a solution because a capability is missing, that failure signal is the roadmap. Customer reality generates the backlog directly.

The factory version works and I endorse it without reservation. When an agent cannot complete a change because no golden path covers the case, because a piece of domain context is missing, or because a required control has no automated evidence path, that failure, logged with its cause and aggregated across the fleet, is the platform backlog. It beats any platform survey, because it is generated by real work failing in a specific and recorded way, and because it is counted rather than remembered.

The modification is that a bank has a second backlog no failure signal will ever generate. DORA did not emerge from a blocked agent task. It arrived with an application date of 17 January 2025 and a perimeter defined in Brussels. Same for every EBA guideline revision, every thematic review, every reporting change with a legal deadline.

So the architecture runs two backlogs with an explicit arbitration rule between them, and the arbitration is a human decision with a name attached, because deprioritising a regulatory deliverable is exactly the class of call where being wrong is existential. Block reserves precisely those calls for people, so this modification sits inside the spirit of the original rather than against it.

The diagram, described

If you are handing this to a designer, here is what I would ask for.

One landscape page. Five horizontal bands stacked bottom to top. Band 0, the evidence plane, drawn as a foundation slab, visually heavier than the rest, spanning the full width. Band 1, capabilities, with six small tiles: golden paths, domain services, controls as code, test data, ephemeral environments, evidence generation. Band 2, world models, split into two side-by-side compartments with a small clock glyph on the delivery side to signal freshness. Band 3, the intelligence layer, a row of agent tiles above a thin strip labelled orchestration, context engineering, evaluation, guardrails. Band 4, interfaces, six small tiles.

Two vertical planes flank the stack at full height. Left is the human edge, carrying role chips and a signature glyph. Right is assurance, carrying model risk, AI Act, DORA, audit trail.

Thick upward arrows from each band into the one above, because each layer consumes the one below. A thin arrow leaving band 3 on the right, looping down the outside and re-entering band 1, labelled failure signal becomes platform backlog. A second arrow entering the frame from outside at the top right, labelled regulatory backlog, exogenous. The two meet at a small diamond labelled human arbitration, positioned inside the left plane, so that the arbitration is visibly a human act.

Caption: layers consume downward, failure signals become the platform backlog, humans hold the gates and the signatures.

The sequencing that matters

Build Layer 0 first. Not as a phase two, not as a parallel workstream that catches up later.

The world model is never better than the plane beneath it. Point one at an organisation whose decisions live in Word files and it will faithfully model the artifacts and miss the decisions, which is worse than useless, because it will be confidently incomplete and people will trust it.

Next post: what Team Topologies becomes when boundaries get drawn around context ownership rather than communication bandwidth, a verdict on each Scrum ceremony, and a routine feature traced end to end through the target factory.

Leave a comment

Your email address will not be published. Required fields are marked *