Report Schema

Every explore strategy writes the same result shape (schemaVersion: 1), so a suite runner, the ledger or a report can read any run without knowing which strategy produced it. The same object is data in --json output, result in the persisted <stem>.result.json, and what MCP get_mission_result returns.

Exit codes

One table for every jevitate command:

ExitClassMeaning
0okclean, succeeded, check passed, fixed, or the command did what it was asked
1defectsdefects found, a gating finding (check), still reproduces (verify-fix, ledger verify, regression run), a success check that did not hold, a stale journey demo
2inconclusivethe run could not finish its work (inconclusive/crashed, a check item errored or the budget ran out, a queued mission could not run). It proves nothing.
3hangthe app hung, and the hang reproduced on replay
4intermittenta hang, or a verify-fix signal, fired on some but not every replay
64usagea usage or input error, and nothing ran: a bad flag or argument, an unreadable input file, an unknown id, missing keys, a target outside the allowlist, --headed without a display, a production environment for demo
130 / 143killedSIGINT / SIGTERM; the partial result is still written

Without --json, a command prints a human summary on stdout and a refusal prints error <CODE>: <message> on stderr. With --json, stdout is one line: {v, ok, data} or {v, ok: false, error: {code, message}}. The exit code is the same either way.

Mission outcome

missionOutcomeExitMeaning
clean0finished its budget and found nothing (goal: the success checks held)
defects-found1at least one confirmed defect (goal: a success check did not hold)
inconclusive / crashed2the run could not do its work (too little coverage, a vacuous check, an unreachable or unresponsive app, a starved host), so its silence proves nothing
hang3the app hung, and it reproduced on replay
intermittent4a hang was observed but did not reproduce on every replay

missionOutcome is always the canonical verdict, a goal run included. A goal's own ending is in goalOutcome (succeeded → clean; failed, exhausted, blocked → defects-found). outcome is strategy-specific (a goal's ending, a frontier run's stop reason) and not the portable verdict. When a run broke or proved nothing, failure.kind says why (insufficient-coverage, vacuous-check, target-unresponsive, degraded-environment, stalled, …). Full list: docs/outcomes.md.

Common fields

FieldMeaning
schemaVersion1. Increases when a common field changes incompatibly; new fields are additive.
strategy, missionOutcome, goalOutcome, exitCodeWhat ran and its verdict (above).
defectsEvery defect, whichever oracle found it, each with a 16-hex fingerprint and a kind. advisory: true marks one reported without gating (e.g. Jev's opinion alone); it never sets the outcome and check never gates on it.
defects[].evidence--evidence-video only: videoPath (the captioned repro clip), screenshots (just before and at the failing step), failingStep, signal, reproduced, and skipped / captureSkips when media was refused. verify-fix --record-video adds before / after.
hangsEvery hang, with its fingerprint and reproduction.
recordingPaths, transcriptPath, resultPathThe Recordings, the decision transcript, and the persisted result.
videoPaths--record-video only: every video the run recorded.
screenshotPaths, screenshotIndex, screenshotsSkipped--screenshots only: the masked images, the index.md contact sheet, and {step, reason} for each capture refused.
sideEffectsEvery write request the run fired (third-party ones marked).
target, scopeThe seed URL and allowlist; the route scope a frontier run was contained to.
usageModel calls, tokens and cost (totalUsd, priced).
hostHealth, environmentDegradedHow starved the host was, and advisories met while it was.
engineThe build that produced the result (below).
"defects": [{
  "kind": "http-5xx",
  "fingerprint": "b8b841bad287ceb7",
  "title": "HTTP 500 from /demo/api/profile",
  "evidence": {
    "videoPath": ".jevitate/logs/<date>/adversarial-<stamp>.evidence/b8b841bad287ceb7/clip.webm",
    "screenshots": ["…/before-step-4.png", "…/at-step-4.png"],
    "failingStep": 4,
    "signal": "server returned 500 (PUT /demo/api/profile)",
    "reproduced": true
  }
}],
"videoPaths": [".jevitate/logs/<date>/adversarial-<stamp>.videos/…webm"],
"screenshotPaths": ["…/adversarial-<stamp>.screenshots/01-profile.png"],
"screenshotIndex": "…/adversarial-<stamp>.screenshots/index.md"

Deprecated in 0.2.0, removed in the next minor: serverLogDefects (use the server-log entries in defects), recordingPath (use recordingPaths[0]), usage.usd (use usage.totalUsd).

The engine field

Every mission result (goal, feature, coverage/exploratory, adversarial, usability) carries an engine object identifying exactly which build produced it:

"engine": { "version": "0.2.0", "commit": "a1b2c3d", "builtAt": "2026-09-23T18:04:11Z" }

commit and builtAt come from a git-rev-parse computation baked in at build time (by esbuild for the published bundle, or a generated module for a tsc dev build). Either falls back to the literal string "unknown" — never a fabricated value — when neither injection path ran (e.g. a source tree with no git history available at build time). jevitate --version prints the same identity: the bare semver once that alone identifies the build, or version (commit …, built …) when it can determine more, so two results produced by different rebuilds in the same session are distinguishable.

Usability report: confidence

Each finding carries a single confidence number in [0,1] — this is the finding's own confidence, not a quality grade (see below). It is a product of four independent factors, each in [0,1]:

FactorWhat it measures
violationJev's probability the principle is violated, oriented by the flag rule for that question's kind (noul, score or choice).
applicabilityJev's probability the principle can sensibly be at issue on this screen type, job and app class.
groundingFrom independent code adjudication (never the model): 1.0 when the observation names a cited control or quotes verified on-screen text; a lower fixed value when evidence was cited but not named in the prose. No verified evidence at all means no finding — it is suppressed, not scored low.
agreementoccurrences ÷ screen-states on that route where the item was judged with the same evidence present — an issue flagged on 1 of 6 identical states is weakly supported.

confidence = mean(violation × applicability × grounding) × agreement, averaged over the finding's deduplicated occurrences (same rubric item × route × implicated controls/text). The product is deliberately conservative: a heuristic that doesn't apply, an unconfirmed violation, or a one-off among many observations can't score high. The report's cutoff, --min-confidence, defaults to 0.3.

Not the same number as quality. A finding may also carry quality: { label, confidence } — an independent grader's separate judgment of whether the finding is actionable, relevant-minor, generic or wrong, filtered by --show. While the grader is uncalibrated (a 0.2.0 preview), every grade is shown by default. That quality.confidence is a different number from the finding's own confidence above — don't conflate the two when reading a report.

Usability report: coverage

Coverage is anti-masquerade: a report with few findings means nothing if most rubric items were never evaluated. Every report carries:

FieldMeaning
totalItemsTotal (rubric entry × screen) items considered.
evaluatedItems actually judged (includes notApplicable items — a deterministic "does not apply" still counts as evaluated).
skippedItems that did not run, each with a reason — never a silently dropped rubric item.
budgetTruncatedScreen ids left un-analyzed because the judgment budget ran out.
notApplicableItems ruled out by a deterministic applicability gate before any model call (e.g. a "choice overload" heuristic can't apply to a two-button consent screen). Counted as evaluated, so they never make coverage read as incomplete.

Coverage is full only when totalItems > 0, evaluated === totalItems, skipped.length === 0 and budgetTruncated.length === 0. Below full coverage, the report's coverageWarning is present and says, verbatim, not to read it as good UX — findings reflect only the items that were actually evaluated.

Usability report: suppressed

A flagged judgment that doesn't survive adjudication is never silently dropped — it's counted and summarized in report.suppressed (total, byReason, byRubricItem, byRubricItemRoute, and the raw items list). Reasons:

ReasonMeaning
ungroundedThe specifics step named no control and no on-screen text.
rejected-evidenceThe specifics step cited a control or text that doesn't exist on the observed screen.
not-confirmedLooking for concrete evidence, the principle turned out not to be violated.
below-min-confidenceGrounded, but the finding's own confidence (above) fell below --min-confidence.
quality-policyThe independent grade (e.g. generic or wrong) isn't in --show's policy.
user-authored-contentA vocabulary-sensitive entry's quoted/cited evidence matches a value the run itself typed — a false positive on user-authored content, not the app's own copy.

What clean guarantees — two different fields

There are two cleans for a usability run, and they mean different things.
  • result.missionOutcome === "clean" (mission-level) means only that the run and its analysis completed — the page loaded, the model calls answered, and a report was produced. It is not a verdict on findings: usability findings are advisory by design and never flip missionOutcome to defects-found. A completed review with 12 major findings still reports missionOutcome: "clean" / exitCode: 0. (With --success checks, a check that did not hold makes it defects-found, as on a goal run.) Only a broken run (page never rendered, a required model call stayed unavailable, the engine crashed) reports inconclusive/crashed instead — including the case where the run finished but the analysis itself failed to produce a report.
  • report.clean (report-level, inside result.report) is the one that actually says something about findings. It is true only when coverage is full and there are zero kept findings and zero suppressed candidates. A report with suppressed-but-real signal (e.g. several findings dropped only for being below --min-confidence) is deliberately not clean — suppression is never allowed to read as "no issues".

Read result.report.clean, not result.missionOutcome, when you want to know "did this app pass its usability review."

Per-finding screen/route attribution

Every finding is attributed to exactly where it was observed:

FieldMeaning
screenIdThe representative screen-state the finding's evidenceRefs resolve in.
routeThe normalized route (URL pathname) the finding was observed on.
screenIdsEvery screen-state id the same (deduplicated) issue was observed on.
occurrencesHow many screen-states on this route exhibited the same issue — the dedupe count, and an input to confidence's agreement factor.
controlsHuman-readable identities of the implicated controls, e.g. button "Accept".
quotesVerbatim on-screen text excerpts, each independently verified present on the screen.
evidenceRefsResolved refs into the analyzed evidence (e.g. { id: "control:0" }) — only the evidence actually implicated, on screenId.

See also Usability Review and Flags Reference.