Skip to content
Xingchi Guo
All projects

Engineering Execution Control Plane

ExecRelay

Engineering execution and evidence infrastructure for developers and AI coding agents. ExecRelay runs work in a real terminal and records process outcomes, evidence, verification, and repository changes as separate, linked records.

Status
Developer preview · macOS
Stack
  • Rust
  • Tauri
  • TypeScript
  • React
  • PTY
  • Shell integration
  • Structured evidence
  • Verification
  • Git evidence
  • Agent adapters

The problem

A developer, a script, or an AI coding agent can all say the same things: the command succeeded, the tests passed, the build is clean, these files changed, the task is complete.

Each of those is a claim. Most of the time the claim is true, which is exactly why it is easy to stop checking. Terminal output scrolls past or gets truncated, a summary describes what was intended rather than what ran, and an agent's closing message is prose written by the same system whose work is being judged.

As more engineering work is done by agents, the useful question shifts from "what did it say?" to "what actually happened, and what does that support?" ExecRelay is built around answering the second question with records rather than narration.

Core principle

ExecRelay keeps four things apart that are usually collapsed into a single "it worked".

  • Agent claimis not the same asProcess outcome
  • Process outputis not the same asStructured evidence
  • Executionis not the same asVerification
  • Repository changeis not the same asChange attribution

AI-generated prose is never treated as process truth. This is a modelling discipline about where each fact comes from. It is not a claim of tamper-proofing or a general security guarantee.

Architecture

The product looks like a terminal, but the terminal is one input. Everything that runs — a command typed by a person, a verification check, a read-only Git collector, a command issued by an agent — becomes the same canonical Execution record. Evidence, verification, and attribution are separate records that point back to it.

  1. Who runs work

    • Developer

      Interactive terminal session in a real PTY.

    • Verification run

      Runs the commands of explicit criteria.

    • AI coding agentIn development

      AgentRun and AgentEvents link to executions; they do not replace them.

    every command becomes

  2. Process truth

    • Execution

      One record per command: lifecycle, outcome, exit code, and where the outcome came from.

    produces

  3. Evidence

    • EvidenceBlock

      Raw terminal bytes, kept alongside a cleaned text view derived from them.

    • Structured evidence

      Parsed test and typecheck results that cite their source blocks.

    • Repository snapshot

      Read-only Git state: branch, head commit, status, diff summary.

    supports

  4. Conclusions

    • VerificationRun

      Did trusted exit codes satisfy the criteria that were snapshotted for this run?

    • ChangeAttribution

      What changed in the repository between the before and after snapshots?

Each box is its own record type. Records reference each other by id; none of them overwrites another.

Underneath, the runtime is split along one boundary. The interface is React and TypeScript; process ownership lives in a native Rust runtime reached through Tauri's command and event layer.

Runtime boundary

  1. Workbench

    React / TypeScript. Presentation and domain state.

    then
  2. Tauri boundary

    Commands down, events up.

    then
  3. Rust runtime

    PTY sessions, process lifecycle, collectors.

    then
  4. Operating system

    Shell, processes, repository.

The split exists so that a rendering layer never owns a long-lived process. PTY allocation, exit status, and the read-only collector runner are native; the interface subscribes to what they report.

Canonical records

Execution

Represents
One concrete command or process run, with its lifecycle, outcome, exit code, working directory, and intent (interactive, verification, evidence collection, or agent).
Why separate
It is the single source of process truth. There is deliberately no second, agent-specific execution type.
Links to
Its evidence blocks, and optionally the verification run, change attribution, or agent run it belongs to.

VerificationRun

Represents
One evaluation of a set of criteria, with an immutable snapshot of those criteria and a result per criterion.
Why separate
A command finishing is not the same as a requirement being met. Editing a profile later must not rewrite history.
Links to
The execution behind each criterion result, the profile it came from, and an optional change attribution.

ChangeAttribution

Represents
The repository delta observed across a window, computed from a before and an after snapshot, with a baseline classification and stated limitations.
Why separate
What the repository shows is independent of what anyone reports having changed.
Links to
Its target (an execution, a verification run, or an agent run) and the two repository snapshots.

AgentRunIn development

Represents
A coding agent's attempt at a task, with its own lifecycle, outcome, and status.
Why separate
A task-level status is not a process outcome, and neither is a verification verdict.
Links to
Lists of execution, verification-run, change-attribution, and event ids. It references those records; it does not copy their state.

AgentEventIn development

Represents
One normalized step in an agent's timeline: a message, tool call, tool result, file change, or a link to an execution.
Why separate
An agent's messages are recorded as messages. They are a timeline of what was said and requested, not evidence of what ran.
Links to
Its agent run, and the raw adapter payload it was normalized from.

Execution truth

ExecRelay does not decide whether a command succeeded by reading text that looks like success.

  • Lifecycle and outcome are separate fields. Lifecycle says whether a command is pending, running, finished, or interrupted. Outcome is succeeded, failed, or unknown.
  • Outcome defaults to unknown. It only becomes succeeded or failed when backed by a real exit code from a trusted source. Exit codes are never inferred.
  • Every outcome records its source. An execution stores where its verdict came from and why it was considered complete — a shell-reported exit, a closed pane, a user interrupt, and so on.
  • Shell integration is correlated per session. The terminal uses standard prompt and command markers, tagged with a per-session nonce. Markers without the right nonce are treated as untrusted, so output that merely imitates a marker cannot set an outcome. This is session correlation, not cryptographic authentication.
  • Error-looking text is only a hint. Output heuristics can flag that diagnostics appeared. They cannot change an outcome to failed.

The terminal itself is a real PTY with split panes, and it is designed to keep working independently of capture: observing work should not be able to break the work.

Evidence

Evidence in ExecRelay has two layers, and neither is simply "the logs".

Raw evidence is the exact terminal output of an execution, including control sequences, stored as evidence blocks. A cleaned display version is derived from it for reading. The raw form is kept; the readable form is the presentation.

Structured evidence is derived by deterministic, versioned parsers. Parsers currently exist for Jest, Vitest, pytest, cargo test, and the TypeScript compiler. Each structured record names the parser and version that produced it and the evidence blocks it was read from, so a parsed "3 failed" can always be traced back to the bytes it came from. Unknown values stay unknown; counts are not filled in.

Structured evidence explains a result. It does not decide one.

What decides

  1. Command

    then
  2. Execution

    then
  3. Trusted exit code

    then
  4. Criterion result

What explains

  1. Execution

    then
  2. Raw evidence blocks

    then
  3. Parser

    then
  4. Structured evidence

Verification

Execution answers what actually happened when this ran? Verification answers does the trusted evidence satisfy a specific requirement?

A verification profile is a list of command criteria, each with the exit codes it expects. Running it creates a VerificationRun:

  • The criteria are snapshotted into the run, so later edits to the profile cannot change what a past run checked.
  • Each criterion runs as a normal execution and is matched back by explicit identity, not by "whichever command finished next". A command typed by a person mid-run cannot satisfy a criterion.
  • A criterion passes or fails only on a trusted exit code. If the outcome is unknown, untrusted, or interrupted, the result is error — unevaluable — rather than a guess.
  • Text never decides. There is no heuristic or model judgment in the verdict; it is a deterministic comparison.
  • Re-running creates a new run. Cancelling marks remaining criteria as skipped instead of inventing failures.

Criteria can carry their own working directory, which is what makes this usable in a monorepo.

Repository evidence

Repository state is collected by read-only Git collectors, and they are constrained on purpose:

  • No shell. A collector is a program plus an argument list; there is no shell string to interpolate into.
  • Allow-listed. Only git may run, only with approved subcommands, and arguments containing mutating keywords are rejected before anything executes.
  • Bounded. Collector output is capped and time-limited.
  • Summaries, not patches. Collectors gather branch and head commit, working-tree status, and per-file diff statistics. Full diff bodies are not stored by default.
  • Provenance on every record. Each collected item names the collector, its version, the collection run, and the repository root.

Change attribution compares a snapshot taken before a window with one taken after. The comparison is a pure function over the two snapshots. Its result is careful about what it can claim:

  • The baseline is classified — clean, already dirty, partial, or not a Git repository — because a dirty starting point weakens what can be attributed.
  • Each file is classified relative to that baseline: newly introduced, changed further, removed, or unchanged.
  • Limitations are recorded with the result, including that other processes may have touched the repository in the same window.

The record states that the repository changed during a window. It does not state that the execution was the sole cause. And a collector failing never alters an execution outcome or a verification verdict.

Designed for AI coding agents

This section describes architecture that is in active development. It is not part of the current public release.

ExecRelay's domain includes a vendor-neutral agent model. An AgentRun is a higher-level attempt at a task; it holds references to the executions, verification runs, and change attributions that happened under it. Agent activity is designed to reuse the same canonical execution and evidence pipeline rather than create a second source of truth:

  • A command run by an agent is an ordinary Execution, tagged with how it was linked to the run — a structured agent event, a dedicated pane, or timing correlation.
  • An agent saying "done" is an AgentEvent of type message. It does not complete the run, set an exit code, or pass a criterion.
  • Adapters declare a capability level — opaque, structured, or controlled — and an opaque agent is never given invented tool calls reconstructed from its terminal text.
  • Raw adapter payloads are preserved next to the normalized events derived from them.

Adapters for Claude Code and Codex exist in the development codebase, with normalization covered by fixture-based tests. Native supervision of live agent processes is still being built, and the public release does not include an agents interface. The boundary is intentional: the execution, evidence, and verification surface ships first, and agents attach to it.

Engineering decisions

Process truth over reported success

Why
A verdict should come from the process. Outcome stays unknown unless a trusted exit code backs it, and its source is recorded.
Tradeoff
More states to model and to show. The interface has to be comfortable displaying unknown instead of a green tick.

One canonical Execution

Why
Human commands, verification checks, collectors, and agent commands all produce the same record, so there is one place to look for what ran.
Tradeoff
Every source has to be translated into the shared contract, and each link to an agent run needs its own provenance.

Verification is exit codes against a snapshot

Why
A deterministic comparison is auditable and repeatable. Snapshotting the criteria keeps history honest when profiles change.
Tradeoff
Criteria are limited to commands with expected exit codes, and anything untrusted becomes an error rather than a pass.

Read-only, allow-listed repository evidence

Why
Observing a repository should not be able to change it, and attribution should describe what was seen without claiming causation.
Tradeoff
Collectors cannot fix or stage anything, store summaries rather than full diffs, and make a weaker claim than proof of cause.

Reliability boundaries

These are specific properties of the implementation, not general assurances.

  • Bounded collection. Collector output is capped and time-limited so a large repository cannot exhaust memory.
  • Explicit provenance. Outcomes, evidence, structured records, and agent links each record where they came from.
  • Failure containment. A parser or collector failing leaves executions and verdicts untouched; capture failing is designed not to affect the terminal.
  • Native process ownership. PTYs and process lifecycle live in the Rust runtime rather than in the rendering layer.
  • Local-first. No account is required, and execution, transcripts, and evidence stay on the machine unless a command you run sends them elsewhere.
  • Tested boundaries. Automated tests cover execution evidence, the shell-integration trust boundary, verification, structured-evidence parsers, Git collectors, and change attribution, alongside typecheck and lint.

What it does not do is equally deliberate: evidence is not cryptographically signed, attribution is not proof of causation, and the current release targets macOS on Apple Silicon.

Product

ExecRelay is available as a developer preview for macOS. The public release covers the terminal, capture, repository changes, and verification surfaces described above. Deeper agent-runtime integration remains in development.