Case 01 · AI harness / design tool

Boreal

Agents can change a codebase. Deciding what deserves to ship is the harder problem.

Boreal is a local-first desktop workbench for directing agents, inspecting their work, and verifying what changed.

Nothing is written to a project until the run has been checked and a person has approved it — and every write stays reversible until it is accepted.

Role
Product design · interface design · implementation
Year
2026
Status
Working Windows build · limited personal V1
Boreal's change review: a unified diff rendered from verified before/after state, with approve, apply, and roll back.
The change review: a unified diff rebuilt from the run's own before-and-after fingerprints, with approve, apply and roll back. Nothing is written until it is approved.

The premise

Execution is not the same thing as trust.

Most agent tooling asks for trust: a prompt goes in, files change, and the evidence arrives as a summary. Boreal treats every change as a transaction instead — the run, the check, the diff and the approval stay separate, visible states.

The workbench is built on the Ariadne engine, which runs scoped tasks and records what each one was allowed to touch.

Boreal's run record: the request, the agent's report, and the independent check, each with its own timestamp.
The run record: what was asked, what the agent reported, and the independent check that followed — each line written by a different party.
Detail of the run record: started working, reported complete, checked the result — with the disclosure lines naming what could not be verified.
Detail: the three states a run passes through, and the disclosure lines that name what the runtime could not report.

Principles

Three rules decide every interaction in the workbench.

  1. The agent proposes. The person approves.

    Approvals happen in a native confirmation dialog, so a forged page payload cannot change what is being approved.

  2. A run is judged against its scope.

    Every task records the profile, the scope and the requirement checks it used, so nothing has to be taken on faith afterwards.

  3. Failure stays a visible state.

    Interrupted runs, blocked approvals and checks that did not run remain states of the transaction rather than disappearing into a success summary.

The execution boundary

The agent proposes. The engine checks. The person approves. The project is written last.

That order is enforced structurally, not by convention. A task is a contract before it is a conversation: a profile, the scope it may write, the paths it may never touch, the checks it must satisfy, and the repair attempts it gets. A profile can demand more evidence or more review; it can never grant itself an approval or widen its own scope.

The agent also never works inside the project. A run receives a copy of the workspace, hashed as it was handed over, with symlinks and junction points refused before they can be followed. The worker gets its contract and the copy — no engine state, and no route to the approval path.

Approval is its own state, reachable only through the engine's trusted channel: the confirmation is assembled from engine records in the native host window, not from the page, so a forged payload cannot change what is being approved. The records are honest about what they are — the approval itself says the identity is asserted by the caller, not verified.

  • A run starts as a record

    The durable run record is written before any worker executes, so an interruption leaves a recoverable state rather than an untracked change.

  • The workspace is a copy

    The run works in a staged copy under the engine's own state root. The project is not the workspace, and the copy is fingerprinted as it is handed over.

  • The check is not the claim

    Validation is computed from the workspace bytes — what changed against the baseline, and whether the task's own checks can inspect it. A claimed success has no bearing on the result.

  • One thing at a time

    A second run is refused while one is executing; mutating calls fail with a busy state instead of queueing behind the lock, so the project is never written by two hands at once.

A claim is never validation — and a completed run is still not accepted work.

The run record

A run is judged against its record, not its summary.

When a worker reports complete, that claim is used for one thing only: deciding there is something to check. Everything else is recomputed. The validator compares the staged workspace to its baseline, enumerates what actually changed, and evaluates the task's declared checks against the bytes on disk.

That produces states that are deliberately unflattering. A workspace that moved after execution is refused rather than checked, because validating a different revision would be theatre. A complete claim that changed nothing fails its checks. A required check with no way to inspect its input is recorded as not run — never as passed.

Review sits above validation and says what it is: a recorded verdict, asserted by the reviewer and labelled as asserted rather than as a second independent check. Acceptance is a separate decision again, and it is never inferred from a pass.

Validation input
Baseline vs. workspace fingerprints — never the worker's summary
A check that cannot inspect
Recorded as not run, never as passed
A workspace that moved
Refused; the run is blocked as a repository conflict
Review
A recorded verdict, labelled asserted-not-verified
Acceptance
A distinct decision, never inferred

The review layer

A change is more than a chat message. It is a record you can walk.

The transaction record is collected from the run workspace's own fingerprints, never from a file list the worker supplies. It lists every created, modified and deleted path with the before and after it was computed from, and approval binds to a hash of exactly that content — change the bytes and the approval is stale by construction.

Applying is journaled: the entry for a path is written before that path is mutated, and each file is replaced atomically. The engine is equally clear about what it does not do — a multi-file apply is not atomic — so an interruption lands in a partial state that recovery can resolve and rollback can undo. A rollback restores the previous bytes, removes directories the change created, and, if a file changed after the apply, refuses to overwrite it: both versions are preserved under the transaction and the conflict is reported rather than smoothed away.

Chat can carry a proposal. It cannot carry a diff, a scope, an approval bound to content, and a way back. That is the surface this layer exists to be.

Detail of the change review's changes panel: created and modified counts, and the three steps the review walked through.
Detail of the change review's diff and its approve and apply controls, with the note that nothing is written until it is approved.
One review, two decisions: what the change set contains and how it was walked through, then the diff itself — with approve, apply, and roll back as the only ways out. Apply stays disabled until the approval is on record.

Execution wasn't the hard part. Knowing what deserved approval was.

Preview before acceptance

The change can be run before it is kept.

Boreal's live preview is a real development server, and it starts only from a preview config the project declares and a person explicitly trusts. Trust is bound to the connector and to a fingerprint of the project's build files, and a changed config makes it stale. The frame never loads that server directly: it loads a tokenized loopback address, and a request without the session token is refused before anything is forwarded. The frame itself is sandboxed with scripts only — no same-origin access back into the workbench.

Inspection is read-only by design. Clicking an element reports what the frame is rendering and where the source claims the line lives; the coordinator validates that claim against the approved project root — extension, containment, third-party directories — and returns a project-relative path and a digest, never file content. Where no mapping can be verified, the panel says exactly that instead of guessing.

Direct text and style edits take their own explicit path: the change is checked against the parsed source and the language server before it is written to the real file, with an undo ledger beside it, and the preview reloads on the next build signal or by the panel's own hand. Structural edits are not free-form — they are proposed as scoped tasks and reviewed like any other change. The canvas works the same way: drafts are revisioned, a stale save is refused rather than merged, and a saved revision can be promoted into a scoped implementation task.

Boreal's live preview: the project running in a sandboxed frame, with a selected element and its source mapping.
The sandboxed preview: the project runs behind a tokenized loopback gateway and can be reloaded, restarted and inspected without writing anything into the project.
Detail of Boreal's inspect panel: run controls, the selected element, and the note that inspecting is read-only.
Detail: inspecting is read-only. The source line it reports was validated against the approved root and returned with a digest — and where no mapping can be verified, the panel says so rather than guessing.
Editing the selected element in Boreal: the text change field and the computed style values beside the running preview.
A typed edit from the inspector: parsed against the source and checked by the language server before the file is written, with an undo ledger beside it. Structural edits take the proposed-and-reviewed path instead.
The whole loop in motion: request, run, check, preview, inspect, and the review that stands between the agent and the project.

Recovery is part of the interaction

An interrupted run is a state, not an error screen.

The store writes atomically and keeps an append-only event log, because the interesting failures happen between steps. If the engine restarts with a run still marked running, that run becomes interrupted with an evidence record and its task returns to an authorised state — it never silently becomes a failure, and it never counts against the repair budget.

The transaction side resolves the same way. An apply that stopped halfway is reclassified by comparing every destination to the hashes it was supposed to have, and a rollback then restores the baseline byte-for-byte. A rollback that meets edits made after the apply stops, preserves both versions and says so. Drafts refuse a save built on a stale revision and keep every revision rather than merging quietly.

What recovery is not: a resumed conversation. Each run still gets a fresh workspace and a fresh agent session, and continuity between attempts is a context package built from the project's own records rather than a replayed chat.

Interrupted run
Recovers as interrupted, with an evidence record; the task returns to an authorised state
Half-applied change
Reclassified by comparing every destination to its before/after hash
Rollback
Restores the baseline byte-for-byte; refuses to overwrite later edits and preserves both versions
Stale draft save
Refused; saved revisions stay immutable
After a restart
Recovery runs when the coordinator starts; atomic writes mean no half-written record survives

Current state

The build runs, packages, and states its own limits.

A packaged Windows build is produced by the project's own pipeline: a frozen sidecar, a protocol-parity check that fails the build if the app calls a method the sidecar does not serve, and a packaged-runtime gate that drives the agent runtime and the typed editor without a model or a network. The build installed and ran in a disposable state root as part of its acceptance record.

The counted evidence at the last verification: 541 Python engine tests, 155 frontend tests, 20 Rust tests, plus the packaged gates re-run against the build. One live model-backed run is recorded, and it is recorded honestly — the engine blocked it when the worker never returned a complete implementation claim. A completed live model edit is not yet proven, and neither is breadth: the preview and inspector work on projects Boreal scaffolds, or on fixtures carrying the same instrumentation, and the record says arbitrary real projects are not supported yet.

Built and working

  • A run works in a staged copy of the project, against a recorded scope, and the project is written last.
  • Every change is a transaction: proposed from workspace fingerprints, checked, approved against a content hash, applied with a journal, and reversible.
  • Approvals are confirmed in the native host window; the confirmation's text comes from the engine, not the page.
  • Validation computes what changed from workspace bytes; a check that cannot inspect its input is never reported as passed.
  • Interrupted runs and half-applied changes recover from durable records instead of becoming error screens.
  • The live preview runs behind a tokenized loopback gateway in a scripts-only sandbox, with read-only inspect and a validated source mapping.
  • A packaged Windows build is produced by a gate-checked pipeline and was installed and exercised in the project's own acceptance record.

Not yet proven

  • Boreal is a development-mode application: the packaged path is verified, but it is not a public release, and it is not code-signed.
  • The preview, the inspector and the source mapping are proven on scaffolded or instrumented projects; arbitrary real projects are explicitly not supported yet.
  • The canvas's promotion path is real, but its implementation worker is still a deterministic renderer rather than a model.
  • A completed live model-backed edit is not yet verified; one live run ended blocked when the worker never returned a complete implementation claim.
  • There is no OS-level containment: an approved command runs with the user's own privileges, and the record says so.
  • Review is a recorded verdict rather than a second independent check, and the acceptance record is written and owned by the build itself — no independent user study is claimed.

A runtime reports what happened. It never decides what is true. Boreal is arranged around that one sentence — the rest is deciding how a person sees the work before anything is written.