Local executed evaluation
aip evaluate <state-id|HEAD> [--type test] [--timeout 5m] [--output-limit 1048576] [--format text|json] -- <program> [args...]
The -- delimiter is required. Every argument after it belongs to the child,
including --format, --help, spaces, and shell metacharacters. The default
kind is test, timeout is five minutes, and retained output is limited to one MiB
per stream. Limits must be positive; output_limit is at most eight MiB per stream.
No shell is inserted. An explicitly requested shell is an ordinary executable.
Stdin is EOF. The subprocess inherits the caller’s environment with PWD set to
the fresh working directory. Environment values are not stored in evidence.
The runner resolves the state once, verifies it, and materializes into a fresh
temporary directory outside the source repository. Execution always starts at
the materialized root, even when evaluate was invoked in a source subdirectory.
Bare executable names are resolved through PATH before execution; relative PATH
entries and relative explicit executable paths are interpreted relative to the
materialized root. The source working directory is never used to find ./program.
The executable path is recorded; this is not a digest of the executable or an
assertion that it is the same tool on another machine.
The reference runner supports Linux and macOS. It creates a new process group, kills that group on timeout/cancellation, and performs best-effort group cleanup after the main process exits. Pipe draining has a two-second bound to avoid a descendant keeping evaluation open indefinitely. Other platforms return an unsupported-runner error before executing anything. A subprocess that deliberately escapes the process group is outside this execution contract.
The runner has the current user’s permissions. A temporary directory and a process group are not a security sandbox. Execute only commands and projects you authorize. There is no CPU, memory, disk, network, or filesystem access isolation. The parent application can catch interruption and persist cancelled evidence; SIGKILL, machine failure, or a storage error may prevent a record from being saved.
After the subprocess and capture finish, the runner inspects the workspace,
stores stream artifacts, and publishes evidence v2. Inspection reads the
evaluated copy’s .aipignore and compares observed files with the state’s files
under those rules: output written to ignored paths is not a workspace change,
while an edited rules file always is. Source HEAD and working
files are not changed by AIP. Execution is outside the repository mutation lock;
only final evidence publication takes the lock. User code still has permission
to modify external paths. Lock contention or storage failure is a command error,
never a claim that evidence was saved. Temporary workspaces are retained for
inspection on both passing and failing runs; no automatic removal is performed
after executing arbitrary code. Their paths are returned as local diagnostic
data, outside the canonical evidence object.
Results
JSON retains the standard envelope. Once evidence is saved, ok: true means
the evaluation record was created, even when the evaluated command failed.
data contains id (evidence ID), state, result, termination, exit_code,
workspace, and directory. Command output is stored as artifacts, never mixed
into the CLI’s JSON stream. state show, object show, and object verify
handle both reported and executed evidence.
- Exit 0: saved evidence has result pass.
- Exit 2: saved evidence has result fail.
- Exit 3: saved evidence has result unknown, including timeout, cancellation, startup failure, signal termination, capture error, or failed inspection.
- Exit 1: invalid invocation or runner/materialization/storage failure without a successfully returned evidence record. The standard error envelope applies.
These process exit codes are distinct from the recorded child’s exit_code. Output truncation by itself does not change the result. Workspace changes by themselves do not change the result. Consumers needing stricter guarantees must inspect those fields. Execution evidence is an observation about the referenced initial state, not a promise that tests never changed their inputs.