Report Schema
Every explore strategy writes the same result shape (schemaVersion: 1), so a suite
runner, the ledger or a report can read any run without knowing which strategy produced it. The
same object is data in --json output, result in the persisted
<stem>.result.json, and what MCP get_mission_result returns.
Exit codes
One table for every jevitate command:
| Exit | Class | Meaning |
|---|---|---|
0 | ok | clean, succeeded, check passed, fixed, or the command did what it was asked |
1 | defects | defects found, a gating finding (check), still reproduces (verify-fix, ledger verify, regression run), a success check that did not hold, a stale journey demo |
2 | inconclusive | the run could not finish its work (inconclusive/crashed, a check item errored or the budget ran out, a queued mission could not run). It proves nothing. |
3 | hang | the app hung, and the hang reproduced on replay |
4 | intermittent | a hang, or a verify-fix signal, fired on some but not every replay |
64 | usage | a usage or input error, and nothing ran: a bad flag or argument, an unreadable input file, an unknown id, missing keys, a target outside the allowlist, --headed without a display, a production environment for demo |
130 / 143 | killed | SIGINT / SIGTERM; the partial result is still written |
Without --json, a command prints a human summary on stdout and a refusal prints
error <CODE>: <message> on stderr. With --json, stdout is one
line: {v, ok, data} or {v, ok: false, error: {code, message}}. The
exit code is the same either way.
Mission outcome
missionOutcome | Exit | Meaning |
|---|---|---|
clean | 0 | finished its budget and found nothing (goal: the success checks held) |
defects-found | 1 | at least one confirmed defect (goal: a success check did not hold) |
inconclusive / crashed | 2 | the run could not do its work (too little coverage, a vacuous check, an unreachable or unresponsive app, a starved host), so its silence proves nothing |
hang | 3 | the app hung, and it reproduced on replay |
intermittent | 4 | a hang was observed but did not reproduce on every replay |
missionOutcome is always the canonical verdict, a goal run included. A goal's own
ending is in goalOutcome (succeeded → clean;
failed, exhausted, blocked → defects-found).
outcome is strategy-specific (a goal's ending, a frontier run's stop reason) and not the
portable verdict. When a run broke or proved nothing, failure.kind says why
(insufficient-coverage, vacuous-check, target-unresponsive,
degraded-environment, stalled, …). Full list:
docs/outcomes.md.
Common fields
| Field | Meaning |
|---|---|
schemaVersion | 1. Increases when a common field changes incompatibly; new fields are additive. |
strategy, missionOutcome, goalOutcome, exitCode | What ran and its verdict (above). |
defects | Every defect, whichever oracle found it, each with a 16-hex fingerprint and a kind. advisory: true marks one reported without gating (e.g. Jev's opinion alone); it never sets the outcome and check never gates on it. |
defects[].evidence | --evidence-video only: videoPath (the captioned repro clip), screenshots (just before and at the failing step), failingStep, signal, reproduced, and skipped / captureSkips when media was refused. verify-fix --record-video adds before / after. |
hangs | Every hang, with its fingerprint and reproduction. |
recordingPaths, transcriptPath, resultPath | The Recordings, the decision transcript, and the persisted result. |
videoPaths | --record-video only: every video the run recorded. |
screenshotPaths, screenshotIndex, screenshotsSkipped | --screenshots only: the masked images, the index.md contact sheet, and {step, reason} for each capture refused. |
sideEffects | Every write request the run fired (third-party ones marked). |
target, scope | The seed URL and allowlist; the route scope a frontier run was contained to. |
usage | Model calls, tokens and cost (totalUsd, priced). |
hostHealth, environmentDegraded | How starved the host was, and advisories met while it was. |
engine | The build that produced the result (below). |
"defects": [{
"kind": "http-5xx",
"fingerprint": "b8b841bad287ceb7",
"title": "HTTP 500 from /demo/api/profile",
"evidence": {
"videoPath": ".jevitate/logs/<date>/adversarial-<stamp>.evidence/b8b841bad287ceb7/clip.webm",
"screenshots": ["…/before-step-4.png", "…/at-step-4.png"],
"failingStep": 4,
"signal": "server returned 500 (PUT /demo/api/profile)",
"reproduced": true
}
}],
"videoPaths": [".jevitate/logs/<date>/adversarial-<stamp>.videos/…webm"],
"screenshotPaths": ["…/adversarial-<stamp>.screenshots/01-profile.png"],
"screenshotIndex": "…/adversarial-<stamp>.screenshots/index.md"
Deprecated in 0.2.0, removed in the next minor: serverLogDefects (use the
server-log entries in defects), recordingPath (use
recordingPaths[0]), usage.usd (use usage.totalUsd).
The engine field
Every mission result (goal, feature, coverage/exploratory, adversarial, usability) carries an
engine object identifying exactly which build produced it:
"engine": { "version": "0.2.0", "commit": "a1b2c3d", "builtAt": "2026-09-23T18:04:11Z" }
commit and builtAt come from a git-rev-parse computation baked in at
build time (by esbuild for the published bundle, or a generated module for a
tsc dev build). Either falls back to the literal string "unknown" —
never a fabricated value — when neither injection path ran (e.g. a source tree with no git
history available at build time). jevitate --version prints the same identity:
the bare semver once that alone identifies the build, or
version (commit …, built …) when it can determine more, so two results produced by
different rebuilds in the same session are distinguishable.
Usability report: confidence
Each finding carries a single confidence number in [0,1] — this is the
finding's own confidence, not a quality grade (see below). It is a product of
four independent factors, each in [0,1]:
Factor What it measures violationJev's probability the principle is violated, oriented by the flag rule for that question's kind (noul, score or choice). applicabilityJev's probability the principle can sensibly be at issue on this screen type, job and app class. groundingFrom independent code adjudication (never the model): 1.0 when the observation names a cited control or quotes verified on-screen text; a lower fixed value when evidence was cited but not named in the prose. No verified evidence at all means no finding — it is suppressed, not scored low. agreementoccurrences ÷ screen-states on that route where the item was judged with the same evidence present — an issue flagged on 1 of 6 identical states is weakly supported.
confidence = mean(violation × applicability × grounding) × agreement, averaged over
the finding's deduplicated occurrences (same rubric item × route × implicated controls/text).
The product is deliberately conservative: a heuristic that doesn't apply, an unconfirmed
violation, or a one-off among many observations can't score high. The report's cutoff,
--min-confidence, defaults to 0.3.
Not the same number as quality. A finding may also carry
quality: { label, confidence } — an independent grader's separate judgment
of whether the finding is actionable, relevant-minor,
generic or wrong, filtered by --show. While the grader is uncalibrated (a 0.2.0 preview), every grade is
shown by default. That quality.confidence is a different
number from the finding's own confidence above — don't conflate the two when
reading a report.
Usability report: coverage
Coverage is anti-masquerade: a report with few findings means nothing if most rubric items were
never evaluated. Every report carries:
Field Meaning totalItemsTotal (rubric entry × screen) items considered. evaluatedItems actually judged (includes notApplicable items — a deterministic "does not apply" still counts as evaluated). skippedItems that did not run, each with a reason — never a silently dropped rubric item. budgetTruncatedScreen ids left un-analyzed because the judgment budget ran out. notApplicableItems ruled out by a deterministic applicability gate before any model call (e.g. a "choice overload" heuristic can't apply to a two-button consent screen). Counted as evaluated, so they never make coverage read as incomplete.
Coverage is full only when totalItems > 0,
evaluated === totalItems, skipped.length === 0 and
budgetTruncated.length === 0. Below full coverage, the report's
coverageWarning is present and says, verbatim, not to read it as good UX — findings
reflect only the items that were actually evaluated.
Usability report: suppressed
A flagged judgment that doesn't survive adjudication is never silently dropped — it's counted and
summarized in report.suppressed (total, byReason,
byRubricItem, byRubricItemRoute, and the raw items list).
Reasons:
Reason Meaning ungroundedThe specifics step named no control and no on-screen text. rejected-evidenceThe specifics step cited a control or text that doesn't exist on the observed screen. not-confirmedLooking for concrete evidence, the principle turned out not to be violated. below-min-confidenceGrounded, but the finding's own confidence (above) fell below --min-confidence. quality-policyThe independent grade (e.g. generic or wrong) isn't in --show's policy. user-authored-contentA vocabulary-sensitive entry's quoted/cited evidence matches a value the run itself typed — a false positive on user-authored content, not the app's own copy.
What clean guarantees — two different fields
There are two cleans for a usability run, and they mean different things. -
result.missionOutcome === "clean" (mission-level) means only that the run and its
analysis completed — the page loaded, the model calls answered, and a report was
produced. It is not a verdict on findings: usability findings are advisory by
design and never flip missionOutcome to defects-found. A completed
review with 12 major findings still reports missionOutcome: "clean" /
exitCode: 0. (With --success checks, a check that did not hold makes it
defects-found, as on a goal run.) Only a broken run (page never rendered, a required model call stayed
unavailable, the engine crashed) reports inconclusive/crashed instead
— including the case where the run finished but the analysis itself failed to produce a report.
-
report.clean (report-level, inside result.report) is the one that
actually says something about findings. It is true only when
coverage is full and there are zero kept findings and zero suppressed
candidates. A report with suppressed-but-real signal (e.g. several findings dropped only for
being below --min-confidence) is deliberately not clean — suppression is never allowed to read as "no issues".
Read result.report.clean, not result.missionOutcome, when you want to
know "did this app pass its usability review."
Per-finding screen/route attribution
Every finding is attributed to exactly where it was observed:
Field Meaning screenIdThe representative screen-state the finding's evidenceRefs resolve in. routeThe normalized route (URL pathname) the finding was observed on. screenIdsEvery screen-state id the same (deduplicated) issue was observed on. occurrencesHow many screen-states on this route exhibited the same issue — the dedupe count, and an input to confidence's agreement factor. controlsHuman-readable identities of the implicated controls, e.g. button "Accept". quotesVerbatim on-screen text excerpts, each independently verified present on the screen. evidenceRefsResolved refs into the analyzed evidence (e.g. { id: "control:0" }) — only the evidence actually implicated, on screenId.
See also Usability Review and Flags Reference.