Coverage for src/ai_jury/ballots.py: 100%
392 statements
« prev ^ index » next coverage.py v7.16.1, created at 2026-09-30 06:29 +0000
« prev ^ index » next coverage.py v7.16.1, created at 2026-09-30 06:29 +0000
1"""Per-reviewer ballots: each panelist's own stance, with vendor provenance (issue #663).
3The consolidated ``findings``/``consensus`` views answer *what the panel found*.
4They do not answer *who said what*, which is the question a downstream gate asks
5when it wants the panel to **be** its review: one head-pinned verdict per
6panelist, attributed to the vendor and model that produced it. Until now that
7existed only as prose in the markdown report's vote block.
9Everything here is pure and deterministic — a function of the outcome, the config
10and the (optional) vote. No I/O, no wall-clock, no randomness. Two renderings are
11built on it:
13* :func:`reviewer_ballots` — the ``reviewers`` array of the JSON report.
14* :func:`keel_reviews` — the ``--format keel-reviews`` bundle, shaped for a
15 consumer that accepts a JSON array of ``{reviewer, verdict, scope, findings,
16 testing, vendor, model}`` records (keel's ``keel review --reviews``).
18Only legitimate, already-reported fields are emitted: severities, locations,
19claims, vendor/model provenance and durations. Never raw diff text, prompt text
20or secrets. Free-text lifted out of an agent's reply (``scope``, ``testing``) is
21attacker-influenced, so it is flattened to a single line and length-capped before
22it leaves this module.
24**A ballot has to name something** (issue #700). Every ballot of a real run read
25``scope: "Reviewed the supplied diff; named no specific file."``, ``testing:
26"not stated"``, ``model: ""`` — and on a tier whose review *is* the panel, those
27ballots were the whole review. The consumer refuses a hand-posted verdict shaped
28like that (:data:`_SCOPE_ANCHORS` mirrors the rule: a path, a ``path:line``, a
29backticked symbol, a called identifier, or a "checked …" clause), precisely
30because a verdict naming nothing cannot be told apart from one never performed.
32**And the name has to exist** (issue #710). Mirroring the consumer's *shapes* was
33not enough on its own: ``describe_scope`` backticked whatever the reviewer wrote
34on its ``Checked:`` line, so ``Checked: nothing`` was rendered as an anchor,
35passed the shape test and counted as a review — the form of naming something with
36none of the substance, satisfiable by an agent that read nothing. Every stated
37token is now resolved against the change the panel was shown
38(:class:`ai_jury.largediff.ChangeIndex`, carried on the outcome), and the shape
39test is what the *resolved* tokens then have to pass. This module is therefore
40deliberately **stricter** than the consumer on that one path and never looser:
41a scope it accepts is one the consumer accepts.
42So the placeholder is gone in three directions:
44* the review prompt asks the reviewer for its own ``Checked:`` / ``Tested:``
45 lines, rather than this module inferring coverage from whatever prose landed;
46* a scope that still names nothing is not written — the ballot is reported as an
47 :data:`ABSTAIN` whose scope states *why*, because an agent that returned
48 nothing useful must not be counted as one that reviewed;
49* ``model`` is never blank: it is the model id actually requested of the agent
50 that answered, or a statement that the CLI's default was used and the CLI does
51 not report which model that was. A ballot naming its vendor but not its model
52 cannot answer whether the same model sat twice, which is the whole product of
53 a cross-vendor panel.
55**And a ballot that named nothing is not a review** (#700, round 2). Recording the
56abstention was only half of it: the abstaining ballot still counted toward the
57number of reviews the run announced and toward ``--min-reviews``, so a panel of
58three "Looks good to me, no concerns." replies satisfied a gate that exists to
59refuse exactly that. :func:`ai_jury.panel.is_review` is now the single definition
60— a ``panelist`` record with a substantive scope and a voting verdict — every
61ballot carries the two fields it reads (``scope_substantive`` and ``verdict``)
62plus the answer itself (``counts_as_review``), and three further consequences
63follow from it:
65* every seat that ran gets a record, including one that returned nothing at all.
66 Dropping it left the bundle unable to say *which* agent fell silent; the record
67 is an abstention naming the seat and the reason, and the count excludes it.
68* a finding attached to no file is a scope only under ``--issue``, where a
69 finding carries no file by construction. In code-review mode a claim is not a
70 place in the code, and a ballot with nothing but claims abstains.
71* ``model_source`` rides along in the ``keel-reviews`` projection too, so a
72 machine consumer of that shape can tell a requested id from a statement about
73 the CLI's default without parsing English.
75**And what it says about itself has to be true of the run** (#709/#710, round 2).
76Two claims here were still stated more strongly than the evidence behind them:
78* the ``Checked:`` line was split into tokens **lexically**, on whitespace, so a
79 changed file whose name contains a space arrived as two tokens that name
80 nothing and the ballot abstained under ``not_in_change`` — the rule for
81 catching a review of nothing, refusing a real review over a space. Tokens are
82 now cut against the change index itself, and a span the reviewer quoted is one
83 token whatever is inside it;
84* a ``model`` id this module *derived* went out under ``model_source:
85 requested``, whose whole claim is that the id was on the wire. Deriving one is
86 still right where no invocation recorded one, but it is labelled
87 :data:`MODEL_RECOMPUTED` and it is never the answer for a stale result-cache
88 entry — the record format changed, so the cache refuses it rather than
89 recomputing over the gap.
90"""
92from __future__ import annotations
94import re
95from typing import Any
97from .findings import flatten_inline
98from .panel import (
99 ABSTAIN,
100 ADAPTER_FAILED,
101 CAUSE_FIELD,
102 CHAIR_ROLE,
103 NAMED_NOTHING,
104 NOT_IN_CHANGE,
105 PANELIST_ROLE,
106 REFUSED,
107 SILENT,
108 ballot_seats,
109 bundle_records,
110 is_review,
111 responded,
112 review_count,
113)
114from .voting import is_abstention, tally_votes
116# ``ABSTAIN`` is the stance recorded for a panelist that did not actually review
117# — nothing at all, an empty reply, a refusal, an adapter that failed, or a reply
118# naming nothing checkable. Counting such a seat as the "clear" stance
119# (APPROVE/READY) is precisely the bug :mod:`ai_jury.voting` refuses to have
120# (issue #251): a non-answer is not an approval. The vote tally drops those
121# reviewers entirely; a ballot list has to name every seat that ran, so it names
122# the abstention instead of inventing a stance for it. It is *defined* in
123# :mod:`ai_jury.panel` and re-exported here for the module's own callers, because
124# :func:`ai_jury.panel.is_review` — the one definition of what counts as a review
125# — has to test for it, and two copies of the token are two places for the count
126# and the record to disagree.
128#: ``reviewer`` / ``name`` used for the chair's entry.
129CHAIR_NAME = "chair"
131#: The ``mode`` that selects the issue-review vocabulary (``--issue``), and with
132#: it the single exception to "a scope must name a file": there a finding carries
133#: no file by construction, so the claims a reviewer raised are what it named.
134ISSUE_MODE = "issue"
136#: What ``testing`` says when the reviewer named no verification at all. It says
137#: it plainly rather than shrugging: "not stated" reads like a field nobody
138#: filled in, and a reader cannot tell that apart from a reviewer that ran
139#: nothing and said so (#700). Both are "no verification was run"; only one of
140#: them used to be legible.
141NOT_STATED = "Nothing run: this reviewer named no command, test run or reproduction."
143#: ``model`` for an agent whose CLI was invoked with no model id pinned. Filled
144#: in from the agent's command so the field states the situation instead of
145#: being blank — "which model answered" has an honest answer here, and it is
146#: "the CLI's default, which the CLI does not report".
147_CLI_DEFAULT_MODEL = "{command} default (the CLI does not report which model answered)"
149#: Where a ballot's ``model`` came from, as one machine token, so a consumer can
150#: tell a real model id from a statement about one without parsing prose.
151MODEL_REQUESTED = "requested" # an id was pinned (or mapped from `effort`) and sent
152MODEL_RECOMPUTED = "recomputed" # no invocation recorded one; derived from the config
153MODEL_CLI_DEFAULT = "cli_default" # nothing pinned; the CLI chose, and does not say
154MODEL_UNKNOWN = "unknown" # the answering slot has no spec in this config
155MODEL_NONE = "none" # there is no agent in this slot at all (an unchaired run)
157#: Caps on free text lifted out of agent output. A reviewer's reply is
158#: attacker-influenced, so a forged 100 kB "scope" must not become the bulk of a
159#: bundle posted downstream.
160_CLAUSE_MAX = 240
161_SCOPE_CLAUSES = 3
162_FILES_LISTED = 8
164# Sentence boundary: end punctuation followed by whitespace. Lines are split
165# first, so a bullet list yields one clause per bullet.
166_SENTENCE_SPLIT = re.compile(r"(?<=[.!?])\s+")
168#: Clauses that describe what a reviewer looked at (folded into ``scope``).
169#: Anchored on word boundaries, which is load-bearing rather than tidy: the most
170#: common finding wording in this tool's own fixtures is "**un**checked return
171#: value", and a bare substring test folds that claim into the coverage summary
172#: as though the reviewer had said it checked something.
173_COVERAGE_RE = re.compile(r"\b(checked|examined|inspected|reviewed|covered)\b", re.IGNORECASE)
175#: Clauses that describe what a reviewer did to verify a claim (``testing``).
176_TESTING_RE = re.compile(
177 r"\b(ran\s+(?:the\s+)?tests?|test\s+suite|unit\s+tests?|i\s+ran|reproduced|"
178 r"verified|verification|tested)\b",
179 re.IGNORECASE,
180)
182#: The two lines :mod:`ai_jury.prompts` asks every reviewer to open with. Matched
183#: after the line has been flattened and stripped of list/heading/emphasis
184#: markers, so ``- **Checked:** src/a.py`` and ``Checked: src/a.py`` both land.
185_STATED_SCOPE_RE = re.compile(r"^checked\s*\**\s*:\s*\**\s*(.+)$", re.IGNORECASE)
186_STATED_TESTING_RE = re.compile(r"^tested\s*\**\s*:\s*\**\s*(.+)$", re.IGNORECASE)
188#: A concrete thing a review points at, as the downstream consumer defines it —
189#: a path, a ``path:line``, a backticked symbol, or a called identifier. This is
190#: a test for *structure*, never a judgement about whether the review was any
191#: good; it distinguishes a review from a receipt. Kept deliberately identical
192#: to the consumer's own rule (keel's ``review-verdict-insubstantial``): a scope
193#: this tool is happy with but the consumer rejects is the defect in #700, and
194#: the only way the two cannot drift is for the same shapes to be listed here.
195#: Drift in the other direction is fine and intended: since #710 a stated
196#: ``Checked:`` token must also *resolve* against the change before it is
197#: rendered as one of these shapes, so this tool is stricter there and never
198#: looser — a scope it accepts is one the consumer accepts.
199_SCOPE_ANCHORS = (
200 re.compile(r"[\w./-]+\.[A-Za-z0-9]{1,5}:\d+"), # path/to/file.py:42
201 re.compile(r"[\w-]+/[\w./-]+\.[A-Za-z0-9]{1,5}\b"), # src/ai_jury/thing.py
202 re.compile(r"`[^`\n]{2,}`"), # `a_symbol`, `--a-flag`
203 re.compile(r"\b\w+\.\w+\(\)"), # module.function()
204)
206#: The escape hatch the consumer's rule keeps, and this one keeps with it: a
207#: genuinely clean review ("checked X, Y and Z; found nothing") is a real
208#: outcome and must not be forced to invent a file reference.
209_CHECKED_CLAUSE_RE = re.compile(r"\bchecked\b[^.\n]{8,}", re.IGNORECASE)
211# --- Resolving a stated `Checked:` line against the change (issue #710) -----
212#
213# `_tick` backticks any non-empty stated value, and a backticked token is an
214# anchor, so `Checked: nothing` — or `everything`, or `the diff` — was rendered
215# as an anchor, passed :func:`scope_is_substantive`, and made the ballot a
216# review. That is the *shape* of an anchor with none of the substance, and a rule
217# satisfied by an agent that read nothing is #700's own failure one layer up.
218#
219# So a stated token is resolved against the change the panel was shown — the
220# paths and symbols of :class:`ai_jury.largediff.ChangeIndex`, carried on the
221# outcome. The check is local and cheap: the jury holds the diff at that moment.
223#: List punctuation: a **hard** token boundary, kept apart from whitespace,
224#: which is not one. See :func:`_scope_tokens` — a changed file whose name
225#: contains a space is one token spelled with whitespace in the middle of it,
226#: and a lexical split on ``[\s,;]+`` cut it in half (#710, round 2).
227_LIST_SEP_SPLIT = re.compile(r"[,;]+")
229#: Whitespace: a *candidate* boundary, resolved against the change.
230_WHITESPACE_SPLIT = re.compile(r"\s+")
232#: A span the reviewer itself marked as one token — backticks, or double quotes,
233#: straight or curly. Whatever is inside is one token even when it contains a
234#: space or a comma: the reviewer drew the boundary, and re-splitting it is the
235#: same defect the joining below exists to fix. Single quotes are deliberately
236#: absent: ``'`` and ``’`` are also apostrophes, so ``the reviewer's own file``
237#: would open a span at ``'s`` and swallow the rest of the line.
238_QUOTED_SPAN_RE = re.compile(r"`([^`\n]+)`|\"([^\"\n]+)\"|“([^”\n]+)”")
240#: The same span shapes, anchored: a run that is *entirely* one mark, marks
241#: included. See :func:`_unmarked` — the marks a reviewer nested inside its own
242#: span come off too, so a path that is quoted *and* backticked resolves the way
243#: either alone does.
244_NESTED_MARK_RE = re.compile(r"\A`([^`\n]+)`\Z|\A\"([^\"\n]+)\"\Z|\A“([^”\n]+)”\Z")
246#: How many nested layers of marks are peeled off a span. Two, because two is as
247#: deep as nesting goes: there are three mark shapes and none can sit inside
248#: itself (``[^`\n]+`` cannot hold a backtick), so once :func:`_marked_spans` has
249#: taken the pair the span opened with, at most two remain. Anything a deeper
250#: line left on can only fail to resolve, which is reported.
251_MAX_NESTED_MARKS = 2
253#: How many whitespace-separated pieces may be joined while looking for a path.
254#: A cap rather than the whole line: the search is quadratic in the window, the
255#: line is attacker-influenced, and a filename with five spaces in it is already
256#: an outlier. Exceeding it can only make a token fail to resolve, which is the
257#: safe direction — an unresolved token is reported as unresolved.
258_MAX_JOINED_PIECES = 6
260#: Punctuation a token may be **wrapped** in: brackets and quotes, plus the
261#: markdown emphasis mark. Stripped from either end, so ``(src/a.py)`` and
262#: ``src/a.py`` are one token. Nothing here can begin a file name, which is why
263#: taking it off either end is unconditional. See :func:`_edge_stripped`.
264_WRAP_EDGE = "`'\"“”‘’()[]{}<>*"
266#: Sentence punctuation a token may be **trailed** by, stripped from the end
267#: only — a full stop ends a sentence far more often than it opens a filename,
268#: but a *leading* dot is part of the name and is never touched. A trailing
269#: ``:`` goes too, but a ``:42`` does not — that is a line. ``_`` is in neither
270#: set: it is part of the identifier, and stripping it turned ``_stated_line``
271#: into a symbol the change does not have.
272_TRAIL_EDGE = ".,;:!?"
274#: ``path:line`` / ``path:line-line`` — the line part is dropped before the path
275#: is looked up, because a diff index knows files, not line numbers.
276_PATH_LINE_RE = re.compile(r"^(?P<path>.+?):(?P<line>\d+(?:-\d+)?)$")
278#: A token that *claims* to name something: a path, a ``path:line``, a call, a
279#: dotfile, or an identifier (an ``_`` or an interior case change). Ordinary
280#: connective prose in the same sentence — "lines", "and", "the tests" — claims
281#: nothing, so it is neither counted as an anchor nor reported as a broken one. A
282#: token the reviewer backticked counts too: the reviewer marked it as a name.
283#:
284#: The **leading-dot** alternative is what makes ``.gitignore`` a name (#710,
285#: round 4). The trailing-extension one never reached it — ``gitignore`` is nine
286#: characters, not a suffix — so an *absent* dotfile was dropped as connective
287#: prose and the ballot said the line named nothing, for a path the reviewer
288#: named exactly. Prose is untouched by it: a full stop that ends a sentence
289#: sits at the end of the piece before it, never at the start of the next.
290# A bare dotted word — ``foo.proto``, ``e.g``, ``Ph.D`` — is *not* name-shaped.
291# Rounds 10 and 11 tried to tell a file from a Latin aside by the shape of the
292# suffix and by the length of the segments, and each rule had a family it
293# missed. The reviewer's own punctuation settles it instead: a slash, a leading
294# dot, or marks around the token make it a claim that is reported when absent;
295# a bare dotted word is read as prose. (#711 round 12)
296_NAME_SHAPED_RE = re.compile(r"[/\\]|\A\.[A-Za-z0-9]|\(\)$|_|[a-z0-9][A-Z]")
298#: Anchor-forming characters, removed before an unresolved token is quoted in an
299#: abstention. See :func:`_deanchor`.
300_DEANCHOR = str.maketrans(dict.fromkeys("`/\\:()[]{}<>", " "))
302#: ...and the one word that forms an anchor without punctuation, under
303#: :data:`_CHECKED_CLAUSE_RE`. Elided rather than dropped silently.
304_CHECKED_WORD_RE = re.compile(r"\bchecked\b", re.IGNORECASE)
307def _deanchor(text: str) -> str:
308 """A reviewer's token, quoted so it cannot pass for an anchor (pure).
310 An unresolved token is named in the abstention — a reader has to see that
311 the reviewer claimed ``src/made/up.py`` — but the abstention is
312 *deliberately anchorless* (see :func:`abstention_scope`): a consumer applying
313 the same substance rule has to reach the same conclusion this module did,
314 and it would not if the sentence explaining that nothing was checked itself
315 contained a path. So the shapes that make an anchor are removed: the
316 punctuation of :data:`_SCOPE_ANCHORS` and the ``checked …`` clause. The
317 reviewer's words survive; only their ability to masquerade as evidence does.
318 """
319 flat = _CHECKED_WORD_RE.sub("…", flatten_inline(text))
320 return " ".join(flat.translate(_DEANCHOR).split())[:_CLAUSE_MAX]
323def _marked_spans(value: str) -> list[tuple[str, bool]]:
324 """*value* as ``(text, quoted)`` runs (pure).
326 ``quoted`` runs are the spans the reviewer wrapped in backticks or double
327 quotes; the rest is everything between them. Splitting on the marks first is
328 what lets a quoted ``docs/my file.py`` survive as one token: a reviewer that
329 quoted a name has already said where it begins and ends, and no rule below
330 is entitled to a second opinion about that.
331 """
332 runs: list[tuple[str, bool]] = []
333 pos = 0
334 for match in _QUOTED_SPAN_RE.finditer(value):
335 if match.start() > pos:
336 runs.append((value[pos : match.start()], False))
337 # Exactly one alternative of :data:`_QUOTED_SPAN_RE` can have matched.
338 runs.append((match.group(1) or match.group(2) or match.group(3), True))
339 pos = match.end()
340 if pos < len(value):
341 runs.append((value[pos:], False))
342 return runs
345def _unmarked(text: str) -> str:
346 """One marked span's content, with nested marks peeled off (pure).
348 :func:`_marked_spans` removes the pair the span *opened* with, and a
349 reviewer that both quoted and backticked a path leaves the inner pair on:
350 ``resolve_stated_scope("“`docs/my file.py`”", changed)`` looked up the
351 backticked string, matched nothing, and reported the reviewer's own
352 ``docs/my file.py`` as a claim that failed (#710, round 3). Quoting a
353 backticked path is the same claim about the same file as either mark alone,
354 so the marks come off before the lookup — as many layers as
355 :data:`_MAX_NESTED_MARKS` allows, whichever mark each layer used.
356 """
357 inner = text.strip()
358 for _ in range(_MAX_NESTED_MARKS):
359 match = _NESTED_MARK_RE.match(inner)
360 if match is None:
361 break
362 inner = (match.group(1) or match.group(2) or match.group(3)).strip()
363 return inner
366def _path_base(token: str) -> str:
367 """*token* with a trailing ``:line`` / ``:line-line`` dropped (pure)."""
368 match = _PATH_LINE_RE.match(token)
369 return match.group("path") if match else token
372#: A call marker ``()`` that may sit inside wrapping punctuation and before
373#: trailing sentence punctuation: ``(module.env()),`` names the call
374#: ``module.env()``. Group 1 is everything before the marker. The trailing class
375#: is *derived* from the same two edge sets the plain strip uses, so the two
376#: cannot disagree about what counts as wrapping (#711 round 9: the hand-written
377#: class lacked ``<>`` and the curly quotes, which ``_WRAP_EDGE`` has).
378_WRAPPED_CALL_RE = re.compile(r"^(.*?)\(\)[" + re.escape(_WRAP_EDGE + _TRAIL_EDGE) + r"]*$")
381def _edge_stripped(piece: str) -> str:
382 """*piece* with its edge punctuation removed (pure).
384 Wrapping punctuation — brackets and quotes, :data:`_WRAP_EDGE` — comes off
385 either end; trailing sentence punctuation (:data:`_TRAIL_EDGE`) comes off
386 the end. **A character that begins a name is never removed**: a dot followed
387 by a letter or a digit, a letter, a digit and ``_`` are in neither set, so
388 ``(.gitignore).`` is ``.gitignore``, ``src/a.py,`` is ``src/a.py``, and
389 ``all.`` is ``all``.
391 A leading dot is part of a file name, and the earlier rounds of #710 stripped
392 it with everything else and then tried to put it back: the trims that
393 restored the stripped characters were offered to the change index,
394 most-stripped first, and the first trim the index confirmed was the token.
395 That asked the diff a question the token had already answered — and the diff
396 answered it wrongly whenever it happened to contain the stripped remnant.
397 With ``env`` (or ``bin/env``) changed, ``Checked: .env`` resolved on the
398 fully stripped ``env``, matching on a component boundary in a *different*
399 file, and the ballot counted as a review of a file the reviewer never named;
400 ``.gitignore`` did the same against a changed ``src/gitignore`` (#709/#710,
401 round 5). Never stripping the dot removes the question rather than adding a
402 case to it: ``.env`` against a changed ``env`` is a name this change does not
403 have, reported unresolved under ``not_in_change`` — and one deterministic
404 strip is all a token needs, so no trim is enumerated against the index at
405 all.
406 """
407 # A trailing ``()`` is a call marker, not wrapping punctuation: it is what
408 # tells :func:`_token_resolves` that ``module.function()`` is a symbol claim
409 # rather than a file (#711 round 7), so it is kept whole.
410 # The marker is looked for *inside* any wrapping, not at the raw end: on
411 # ``(module.env()),`` the last characters are wrap and trail punctuation,
412 # and a check at the raw end missed the call and then ate its parentheses
413 # as wrapping (#711 round 8).
414 marked = _WRAPPED_CALL_RE.match(piece)
415 if marked:
416 stripped = marked.group(1).lstrip(_WRAP_EDGE)
417 return stripped + "()" if stripped else ""
418 return piece.lstrip(_WRAP_EDGE).rstrip(_WRAP_EDGE + _TRAIL_EDGE)
421def _joined_against_change(pieces: list[str], changed: Any) -> list[str]:
422 """One whitespace-separated run's pieces, with spaced paths rejoined (pure).
424 **The change index is the tokeniser, not the whitespace.** ``Checked:
425 docs/my file.py`` used to split into ``docs/my`` and ``file.py``, neither of
426 which is in a diff that changes ``docs/my file.py``: the ballot came back
427 ``scope_substantive: false``, ``counts_as_review: false``,
428 ``abstention_cause: not_in_change`` — a real review of a real file, refused
429 for a space in its name, by the rule that exists to catch reviews of nothing.
431 So adjacent pieces are offered to the index joined, longest window first, and
432 a join that names a changed path *is* the token. Longest-first matters: with
433 both ``my file.py`` and ``docs/my file.py`` changed, the reviewer named the
434 second. A join is only ever accepted when the index confirms it, so this can
435 turn a non-token into a token but never the reverse — the failure direction
436 is an unresolved token, which is reported as one.
437 """
438 tokens: list[str] = []
439 index = 0
440 while index < len(pieces):
441 joined = ""
442 width = 1
443 for size in range(min(_MAX_JOINED_PIECES, len(pieces) - index), 1, -1):
444 candidate = _edge_stripped(" ".join(pieces[index : index + size]))
445 if changed.has_path(_path_base(candidate)):
446 joined, width = candidate, size
447 break
448 tokens.append(joined or _edge_stripped(pieces[index]))
449 index += width
450 return tokens
453def _scope_tokens(value: str, changed: Any) -> list[str]:
454 """The distinct candidate tokens of one stated ``Checked:`` value (pure).
456 Tokenised **against the change**, not lexically (#710, round 2). Three rules,
457 in this order: a span the reviewer quoted is one token whatever is inside it;
458 list punctuation is a hard boundary; and whitespace is a boundary only where
459 joining across it does not name a changed path.
460 """
461 seen: list[str] = []
462 for text, quoted in _marked_spans(flatten_inline(value or "")):
463 if quoted:
464 # No edge punctuation is stripped: inside the reviewer's own marks
465 # there is none to strip, and stripping it would cost the leading
466 # dot of a `.gitignore` the reviewer took care to quote. Only marks
467 # the reviewer nested inside its own span come off — see
468 # :func:`_unmarked`.
469 candidates = [_unmarked(text)]
470 else:
471 candidates = [
472 token
473 for segment in _LIST_SEP_SPLIT.split(text)
474 for token in _joined_against_change(
475 [p for p in _WHITESPACE_SPLIT.split(segment) if p], changed
476 )
477 ]
478 for token in candidates:
479 if token and token not in seen:
480 seen.append(token)
481 return seen
484def _is_name_shaped(token: str, value: str) -> bool:
485 """Does *token* claim to name something? (pure — see :data:`_NAME_SHAPED_RE`)
487 A token the reviewer *marked* counts too, whichever mark it used: quoting a
488 token is the reviewer saying it is a name, and a claim that failed has to be
489 reported as one rather than dropped as connective prose.
490 """
491 if _NAME_SHAPED_RE.search(token):
492 return True
493 return any(f"{o}{token}{c}" in (value or "") for o, c in (("`", "`"), ('"', '"'), ("“", "”")))
496def _path_shaped(base: str) -> bool:
497 """Does *base* claim to be a path rather than a symbol? (pure)
499 A directory, a dotfile, or any dotted name written without ``()``. The
500 extension list this replaced could not be complete — ``foo.proto`` fell
501 through to the symbol index and matched a ``proto()`` call (#711 round 7) —
502 so the rule is now the reviewer's own punctuation: ``module.function()``
503 names code, ``module.function`` and ``foo.proto`` name a file.
504 """
505 if "/" in base or base.startswith("."):
506 return True
507 return "." in base and not base.endswith("()")
510def _token_resolves(token: str, changed: Any) -> bool:
511 """Is *token* a path or symbol that is actually in the change? (pure)
513 A path-shaped token — a directory, a dotfile, or a dotted name without
514 ``()`` — is a path claim and resolves only as a path. It is never split and
515 asked of the symbol index: ``.env`` against a hunk that adds ``env(1)`` used
516 to resolve on the remnant ``env`` and count as a review of a file the
517 reviewer never named (#711 round 6). A symbol claim is a token with no dot,
518 or one ending in ``()``; for ``module.function()`` the member ``function``
519 is asked, because the qualification is the reviewer's and the member is
520 the claim.
521 """
522 base = _path_base(token)
523 if changed.has_path(base):
524 return True
525 if _path_shaped(base):
526 return False
527 member = base.rstrip("()").rsplit(".", 1)[-1]
528 return bool(member) and changed.has_symbol(member)
531def resolve_stated_scope(value: str, changed: Any) -> tuple[list[str], list[str]]:
532 """Split one ``Checked:`` value into ``(resolved, unresolved)`` tokens (pure).
534 ``resolved`` are the tokens that name a path or a symbol present in
535 *changed*; they are what makes the scope substantive. ``unresolved`` are the
536 tokens that are shaped like a name and are not in the change — reported, so
537 a reader sees the claim that failed, but never anchoring anything.
539 **The split into tokens is itself made against the change** (see
540 :func:`_scope_tokens`), because a lexical one gets a changed file whose name
541 contains a space wrong in the direction that costs a real review.
543 **A mixed line is carried by its real tokens.** ``Checked:
544 src/ai_jury/ballots.py, src/made/up.py`` is a review of
545 ``src/ai_jury/ballots.py``, with the second path listed in the scope as not
546 in the change: the reviewer demonstrably read something a reader can go and
547 check, and abstaining over the extra token would discard a real review to
548 punish a typo. A line with *no* resolving token is not a review, whatever
549 else it says.
550 """
551 resolved: list[str] = []
552 unresolved: list[str] = []
553 for token in _scope_tokens(value, changed):
554 if _token_resolves(token, changed):
555 resolved.append(token)
556 elif _is_name_shaped(token, value):
557 unresolved.append(token)
558 return resolved, unresolved
561def normalize_verdict(verdict: str) -> str:
562 """Fold a display verdict into a single machine token.
564 ``REQUEST CHANGES`` → ``REQUEST_CHANGES``, ``NEEDS-INFO`` → ``NEEDS_INFO``,
565 ``NO QUORUM`` → ``NO_QUORUM``. The markdown report keeps the spaced form for
566 humans; a machine consumer keys on one word, and a verdict that changes shape
567 between the two renderings is a verdict that gets matched wrong.
568 """
569 token = flatten_inline(verdict or "").strip().upper()
570 return re.sub(r"[\s\-]+", "_", token)
573def _prose_lines(text: str) -> list[str]:
574 """The agent's prose with fenced blocks removed.
576 A review's fenced ``json`` block is the *structured* findings — it is already
577 parsed into :class:`~ai_jury.findings.Finding` objects and rendered as the
578 ballot's ``findings``. Left in, its serialized claim/evidence text is by far
579 the longest thing in the reply and swamps every prose clause after it.
580 Unterminated fences swallow the remainder, which is the fail-safe direction:
581 less lifted text, never more.
582 """
583 lines: list[str] = []
584 in_fence = False
585 for raw in (text or "").splitlines():
586 if raw.lstrip().startswith(("```", "~~~")):
587 in_fence = not in_fence
588 continue
589 if not in_fence:
590 lines.append(raw)
591 return lines
594def _clauses(text: str) -> list[str]:
595 """Split agent output into flattened, capped candidate clauses (pure)."""
596 out: list[str] = []
597 for raw in _prose_lines(text):
598 line = flatten_inline(raw).strip().lstrip("-*#> ").strip()
599 if not line:
600 continue
601 # The split pattern consumes the whitespace after the end punctuation and
602 # ``line`` is already stripped, so no part can be blank — hence no guard.
603 for part in _SENTENCE_SPLIT.split(line):
604 out.append(part[:_CLAUSE_MAX])
605 return out
608def _first_matching(clauses: list[str], pattern: re.Pattern[str]) -> str:
609 for clause in clauses:
610 if pattern.search(clause):
611 return clause
612 return ""
615def _matching(clauses: list[str], pattern: re.Pattern[str], limit: int) -> list[str]:
616 hits: list[str] = []
617 for clause in clauses:
618 if pattern.search(clause) and clause not in hits:
619 hits.append(clause)
620 if len(hits) >= limit:
621 break
622 return hits
625def _files_named(findings: list) -> list[str]:
626 """Distinct file paths a reviewer named, in first-reported order."""
627 seen: list[str] = []
628 for f in findings:
629 path = flatten_inline(getattr(f, "file", "") or "").strip()
630 if path and path not in seen:
631 seen.append(path)
632 return seen
635def scope_is_substantive(scope: str) -> bool:
636 """Does this scope name something a reader could go and check? (pure)
638 The gate between a ballot and an abstention. Note what it does *not* do: it
639 never asks whether the review was correct, thorough or agreeable — it cannot,
640 and trying would make this a critic. It asks only whether the text points at
641 anything, which is the single question separating "this agent reviewed the
642 diff" from "this agent returned a string".
643 """
644 text = (scope or "").strip()
645 if not text:
646 return False
647 return bool(_CHECKED_CLAUSE_RE.search(text)) or any(p.search(text) for p in _SCOPE_ANCHORS)
650def _tick(text: str) -> str:
651 """One concrete token, backticked for a scope line.
653 Two jobs, and the second is why this is not an f-string at the call site.
654 Backticks make the token an anchor under :data:`_SCOPE_ANCHORS`, so a scope
655 built from a bare filename (``notes.md`` — no directory, so it matches none
656 of the path shapes) still names something checkable. And the token is
657 attacker-influenced, so its own backticks are stripped first: otherwise a
658 crafted path could close the quoting and forge structure around it.
659 """
660 return "`" + flatten_inline(text).replace("`", "").strip() + "`"
663def _stated_line(result: Any, pattern: re.Pattern[str]) -> str:
664 """The value of the first ``Checked:``/``Tested:`` line in the reply (pure).
666 The reviewer's own statement of its coverage, which beats anything this
667 module can infer from prose — inference is exactly how a ballot that named
668 nothing still shipped a scope sentence (#700). Fenced blocks are skipped as
669 everywhere else here, and the value is flattened and capped before use.
670 """
671 for raw in _prose_lines(getattr(result, "output", "")):
672 line = flatten_inline(raw).strip().lstrip("-*#> ").strip()
673 match = pattern.match(line)
674 if match:
675 value = match.group(1).strip().strip("*").strip()
676 if value:
677 return value[:_CLAUSE_MAX]
678 return ""
681def _free_clauses(result: Any) -> list[str]:
682 """Candidate prose clauses with the reviewer's own stated lines removed.
684 ``Checked: …`` matches the coverage pattern, so without this the stated line
685 is folded in twice — once quoted as the reviewer's statement and once again
686 as an inferred clause — and the second copy is unquoted, which is how a
687 crafted path gets into the scope without :func:`_tick` seeing it.
688 """
689 return [
690 clause
691 for clause in _clauses(getattr(result, "output", ""))
692 if not _STATED_SCOPE_RE.match(clause) and not _STATED_TESTING_RE.match(clause)
693 ]
696def _claims_named(findings: list) -> list[str]:
697 """Distinct claims a reviewer raised, for a review that attached no file.
699 Not decoration, and **scoped to ``--issue``**: there, every finding carries
700 ``file: ""`` by construction — the panel is reading an issue's prose, not a
701 diff, so there is no file for a finding to name — and the file list is empty
702 for a reviewer that did real work. Its claims are the only thing it *can*
703 name, and naming them keeps an issue-mode ballot out of the abstention branch
704 it does not belong in. That is the one legitimate exception to keel's rule
705 that a scope must name a file, line or symbol, or carry a ``Checked …``
706 clause, and :func:`describe_scope` applies it in issue mode only.
708 In code-review mode the same fallback was a hole (#700, round 2): a finding
709 with ``file: ""`` is a claim about a diff that failed to say *where*, and
710 :func:`_tick` backticked it into an anchor, so a scope naming no place in the
711 code passed the substance test and the ballot cast a voting verdict.
712 """
713 seen: list[str] = []
714 for f in findings:
715 claim = flatten_inline(getattr(f, "claim", "") or "").strip()
716 if claim and claim not in seen:
717 seen.append(claim[:_CLAUSE_MAX])
718 return seen
721def _stated_scope_sentence(stated: str, changed: Any) -> str:
722 """The scope sentence for a reviewer's own ``Checked:`` line (pure).
724 ``""`` when the line resolved to nothing in the change — the caller then has
725 no sentence to add from this source, and the ballot falls through to the
726 next one exactly as a reply with no ``Checked:`` line does. The unresolved
727 tokens are not silently dropped: :func:`_no_scope_reason` names them in the
728 abstention, where they cannot be mistaken for evidence.
730 With ``changed`` as ``None`` there is nothing to resolve against — a
731 hand-built outcome, a caller that never had a diff — and the pre-#710
732 structural rule applies. "Not verifiable here" is not "does not exist".
733 """
734 if changed is None:
735 return f"Checked, as stated by the reviewer: {_tick(stated)}."
736 resolved, unresolved = resolve_stated_scope(stated, changed)
737 if not resolved:
738 return ""
739 sentence = f"Checked, as stated by the reviewer: {', '.join(_tick(t) for t in resolved)}."
740 if unresolved:
741 listed = ", ".join(_deanchor(t) for t in unresolved[:_FILES_LISTED])
742 sentence += (
743 f" The same line also named {listed} — not in this change, and so"
744 f" anchoring nothing; the rest of the line is what this ballot rests on."
745 )
746 return sentence
749def describe_scope(result: Any, findings: list, *, mode: str = "code", changed: Any = None) -> str:
750 """What this panelist named that it read — or ``""`` when it named nothing.
752 Pure and deterministic. Four sources, most authoritative first:
754 1. the reviewer's own ``Checked:`` line, **resolved against the change** —
755 a token counts only when it names a path or a symbol that is actually in
756 the diff (#710), so ``Checked: nothing`` contributes nothing;
757 2. the distinct files it attached to its structured findings;
758 3. up to three "checked / examined / reviewed" clauses from its prose;
759 4. **in ``--issue`` mode only**, failing a file, the claims it raised.
761 An empty return is the meaningful case and the reason this no longer emits a
762 sentence unconditionally: with nothing from any of the four, the honest
763 output is *nothing*, and the caller turns that into an abstention. The old
764 fallback — "Reviewed the supplied diff; named no specific file." — asserted
765 coverage from the absence of evidence for it, and read identically whether
766 the agent had reviewed all 17 files or returned an empty string.
768 ``changed`` is what the first source is resolved against (#710). Until then
769 :func:`_tick` backticked whatever the reviewer wrote and
770 :func:`scope_is_substantive` accepted any backticked token, so ``Checked:
771 nothing`` was rendered as an anchor and the ballot counted as a review — the
772 shape of naming something, satisfiable by an agent that read nothing. The
773 ``Checked:`` path is now held to the standard the findings-derived path
774 already met: it must name a place that exists in the change.
776 ``mode`` gates the fourth source, and that gate is the whole of #700's third
777 round. keel's rule is that a scope must name a file, line or symbol, or carry
778 a ``Checked …`` clause; a backticked *claim* is none of those, and letting one
779 stand as a scope meant a code-review ballot raising one ``major`` finding
780 against ``file: ""`` cast ``REQUEST_CHANGES`` while naming no place in the
781 code. Issue mode is the one legitimate exception — see :func:`_claims_named`
782 — because there a finding genuinely has no file to name.
783 """
784 parts: list[str] = []
785 stated = _stated_line(result, _STATED_SCOPE_RE)
786 if stated:
787 sentence = _stated_scope_sentence(stated, changed)
788 if sentence:
789 parts.append(sentence)
790 files = _files_named(findings)
791 if files:
792 listed = ", ".join(_tick(f) for f in files[:_FILES_LISTED])
793 more = len(files) - _FILES_LISTED
794 suffix = f" (+{more} more)" if more > 0 else ""
795 parts.append(f"Named {len(files)} file(s): {listed}{suffix}.")
796 parts.extend(_matching(_free_clauses(result), _COVERAGE_RE, _SCOPE_CLAUSES))
797 if not files and mode == ISSUE_MODE:
798 claims = _claims_named(findings)
799 if claims:
800 listed = ", ".join(_tick(c) for c in claims[:_FILES_LISTED])
801 more = len(claims) - _FILES_LISTED
802 suffix = f" (+{more} more)" if more > 0 else ""
803 parts.append(f"Raised {len(claims)} finding(s) against no file: {listed}{suffix}.")
804 scope = " ".join(parts)
805 return scope if scope_is_substantive(scope) else ""
808def _named_only_absent(result: Any, changed: Any) -> bool:
809 """Did this seat state a ``Checked:`` line that names only absent things? (pure)"""
810 if changed is None:
811 return False
812 stated = _stated_line(result, _STATED_SCOPE_RE)
813 if not stated:
814 return False
815 resolved, unresolved = resolve_stated_scope(stated, changed)
816 return not resolved and bool(unresolved)
819def abstention_cause(result: Any, changed: Any = None) -> str:
820 """Which of :data:`ai_jury.panel.ABSTENTION_CAUSES` this seat's ballot records.
822 **The** classification, and the only one: the ballot carries its answer under
823 :data:`ai_jury.panel.CAUSE_FIELD`, the two abstention sentences below are
824 written from it, and :func:`ai_jury.panel.abstention_buckets` counts by it. A
825 second reading of the raw result is a second place for the count and the
826 prose to part company, which is how a seat that named a file and then refused
827 came to be counted as one that named nothing (#700, round 5).
829 ``changed`` is the change under review, and it separates the last two
830 causes: without it a seat that named only absent things is indistinguishable
831 from one that named nothing, so the cause degrades to ``named_nothing``
832 rather than being guessed.
834 Silence is tested **first**, so ``silent`` keeps meaning exactly what
835 :func:`ai_jury.panel.responded` says and the metadata's ``silent`` is the
836 same number whether it is taken from the results or from the ballots. A seat
837 that produced no output at all is silent even when its adapter also reported
838 failure: "nothing came back" is the fact an operator acts on, and the adapter
839 status is on the record beside it.
841 Only ever asked of a ballot that did not review; a seat that reviewed has no
842 cause, and :func:`reviewer_ballots` records an empty string for it.
843 """
844 if not responded(result):
845 return SILENT
846 if not getattr(result, "ok", False):
847 return ADAPTER_FAILED
848 if is_abstention(getattr(result, "output", "")):
849 return REFUSED
850 # The two shapes of "its scope did not stand", kept apart because they send
851 # their reader to opposite places (#710). A seat whose ``Checked:`` line
852 # named `src/made/up.py` did not fail to say what it read — it said it read
853 # something this change does not contain, and "named nothing checkable"
854 # printed over that ballot is a description its own scope contradicts.
855 if _named_only_absent(result, changed):
856 return NOT_IN_CHANGE
857 return NAMED_NOTHING
860#: The sentence each cause contributes to the scope of a ballot that could not
861#: state one. Keyed by :func:`abstention_cause` so the reason and the bucket are
862#: one classification; ``named_nothing`` is refined below by whether the seat
863#: raised findings it attached to nothing.
864_NO_SCOPE_REASONS = {
865 SILENT: "it ran and returned nothing at all — an empty reply from the CLI",
866 ADAPTER_FAILED: "its adapter reported failure and what came back named nothing",
867 REFUSED: "it returned a refusal rather than a review",
868}
871def _no_scope_reason(result: Any, findings: list | None = None, changed: Any = None) -> str:
872 """Why nothing checkable could be lifted from this seat (pure).
874 Six reasons, kept apart because they ask for different fixes: a broken
875 adapter, a CLI that answered with nothing at all, a refusal, a ``Checked:``
876 line naming only things this change does not contain, a ``Checked:`` line
877 naming nothing at all, a code review whose findings named no location, and a
878 reply that reviewed nothing. The first three are :func:`abstention_cause`'s,
879 read from the one classifier rather than re-tested here; the rest are the
880 shapes of ``named_nothing`` and ``not_in_change``, and the splits matter
881 because "said nothing" sends its reader somewhere different from "named a
882 file that is not in the diff" and from "raised findings and attached none of
883 them to a file".
885 Every quoted token passes through :func:`_deanchor` first: this sentence
886 lands in the ballot's ``scope``, and a scope explaining that nothing was
887 checked must not itself read as an anchor to the consumer applying the same
888 rule.
889 """
890 cause = abstention_cause(result, changed)
891 stated = _NO_SCOPE_REASONS.get(cause)
892 if stated:
893 return stated
894 named = _stated_line(result, _STATED_SCOPE_RE)
895 if cause == NOT_IN_CHANGE:
896 listed = ", ".join(_deanchor(t) for t in resolve_stated_scope(named, changed)[1])
897 return (
898 f"it stated it read {listed}, and no such path or symbol is in this change — "
899 f"so the ballot names a place a reader cannot go to, which is not the same as "
900 f"naming none"
901 )
902 if named:
903 return (
904 f"it stated it read {_deanchor(named)}, which names no path, line or symbol in "
905 f"this change at all — the shape of a scope with nothing in it"
906 )
907 raised = len(findings or [])
908 if raised:
909 return (
910 f"it raised {raised} finding(s) but attached none of them to a file, and its "
911 f"reply named no file, symbol or coverage clause either, so the ballot points "
912 f"at no place in the code a reader could go and check"
913 )
914 return "its reply named no file, symbol, coverage clause or finding"
917def abstention_scope(result: Any, findings: list | None = None, changed: Any = None) -> str:
918 """The scope of a ballot that could not state one: the reason, in the field.
920 Deliberately anchorless — no path, no backticked symbol, no "checked …"
921 clause — so a consumer applying the same substance rule reaches the same
922 conclusion this module did instead of being talked past it. The record is
923 here to say *nothing was reviewed*; dressing it up to survive the gate would
924 reinstate the defect with better prose.
926 ``findings`` is this reviewer's own findings, and it is passed so the reason
927 can tell "said nothing" from "said something that named nowhere" — the second
928 is a reviewer that worked and skipped the locations, and a reason that called
929 it "named no finding" would send its author looking for the wrong problem.
930 """
931 name = flatten_inline(getattr(result, "agent", "") or "").strip() or "this reviewer"
932 return (
933 f"Abstention: no scope can be stated for '{name}' because "
934 f"{_no_scope_reason(result, findings, changed)}. Recorded as an abstention rather than an "
935 f"approval — an agent that named nothing did not review, and counting it "
936 f"as one that did is the difference between a panel and a receipt."
937 )
940def describe_testing(result: Any) -> str:
941 """What this panelist ran to verify its claims, or :data:`NOT_STATED`.
943 The reviewer's ``Tested:`` line first, then any verification clause in its
944 prose. Both are lifted verbatim (flattened and capped) rather than
945 summarized: a testing claim carried downstream as evidence must be the
946 reviewer's words, not this module's paraphrase of them. With neither, the
947 field says plainly that nothing was run — which is a statement about the
948 review, where "not stated" was a statement about the field.
949 """
950 stated = _stated_line(result, _STATED_TESTING_RE)
951 if stated:
952 return f"Tested, as stated by the reviewer: {stated}"
953 clause = _first_matching(_free_clauses(result), _TESTING_RE)
954 return clause or NOT_STATED
957def _spec_for(config: Any, agent_name: str) -> Any:
958 for spec in getattr(config, "agents", []) or []:
959 if spec.name == agent_name:
960 return spec
961 return None
964def sent_model(result: Any) -> str:
965 """The model id this seat's invocation recorded having sent (``""`` if none).
967 :attr:`ai_jury.adapters.AgentResult.model`, stamped by the path that ran the
968 adapter from :meth:`ai_jury.adapters.Adapter.resolved_model` — the same call
969 that put the id in the argv or the request payload. Reading it back is the
970 whole of #709's fix: the ballot quotes the id the run sent instead of
971 computing a second one that can differ from it.
973 Empty for a record no invocation produced — a hand-built result, a chair
974 slot with no round-1 seat — and :func:`requested_model` answers for those
975 instead, under :data:`MODEL_RECOMPUTED` rather than :data:`MODEL_REQUESTED`.
976 A stale result cache is deliberately *not* on that list: the record format
977 gained this field, so :data:`ai_jury.cache.CACHE_SCHEMA` refuses an entry
978 written without it rather than serving a recomputation in its place.
979 """
980 return (getattr(result, "model", "") or "").strip()
983def requested_model(spec: Any) -> str:
984 """The model id this agent's CLI would be asked for, from the spec alone.
986 The fallback for a record that carries no sent id (:func:`sent_model`), and
987 it computes the id the way the invocation path computes it — with
988 :func:`ai_jury.config.spec_adapter`, **not** ``spec.vendor``. What it returns
989 ships under :data:`MODEL_RECOMPUTED`: it is derived here, and the one thing
990 it cannot be called is the id that was sent.
992 That distinction is #709. ``spec.model`` is what the operator wrote down, but
993 it is not always what is sent: where reasoning effort is encoded *in the
994 model id*, the ``effort`` knob rewrites it. How effort is expressed is a
995 property of the protocol the seat is invoked through, so every adapter keys
996 :func:`ai_jury.adapters.effort_args` on the adapter; this keyed it on the
997 vendor, and since #705 the two can differ. A seat configured
998 ``vendor = google, adapter = cli, model = gemini-3-pro, effort = high`` was
999 invoked with ``gemini-3-pro`` and balloted ``gemini-3-pro-high``, under a
1000 ``model_source: requested`` that claims to be the id actually sent.
1002 One thing this cannot see, and the reason the sent id is preferred over it:
1003 the invocation may consult the vendor's live model listing and fall back when
1004 the mapped id is not offered. That is I/O, and this module is pure.
1006 An effort level :func:`ai_jury.adapters.effort_args` rejects degrades to the
1007 configured id — a bad config value is ``validate_config``'s to refuse, never
1008 a ballot's to crash on.
1009 """
1010 from .adapters import effort_args
1011 from .config import spec_adapter
1013 configured = (getattr(spec, "model", "") or "").strip()
1014 try:
1015 plan = effort_args(spec_adapter(spec), getattr(spec, "effort", None), configured)
1016 except ValueError:
1017 return configured
1018 return (getattr(plan, "model", "") or configured or "").strip()
1021def describe_model(config: Any, agent_name: str, result: Any = None) -> tuple[str, str]:
1022 """``(model, source)`` for the agent that answered in this slot.
1024 ``result`` is that seat's own round-1 result when there is one, and it is
1025 the first source: it carries the id its invocation sent (#709), so
1026 ``model_source: "requested"`` names the string that was actually on the wire
1027 rather than a second derivation of it.
1029 Never the empty string for a slot that has an agent. An empty ``model`` was
1030 the provenance half of #700: a ballot naming ``vendor: "openai"`` and
1031 ``model: ""`` cannot answer "was that the same model as the other seat", and
1032 provenance is the entire product of a cross-vendor panel. Where the CLI was
1033 invoked with no id pinned there *is* an honest answer — the CLI's own
1034 default, which the CLI does not report back — so the field says that instead
1035 of going blank and letting a reader guess which of the two it meant.
1037 ``source`` is the same fact as one machine token, so a consumer can tell an
1038 id from a statement about one without parsing English.
1040 **A derived id is never labelled as a sent one** (#709, round 2). Where no
1041 invocation recorded an id — a hand-built result, a chair slot with no
1042 round-1 seat, a library caller — :func:`requested_model` still answers, and
1043 that answer goes out under :data:`MODEL_RECOMPUTED` rather than
1044 :data:`MODEL_REQUESTED`. It is a real id and worth quoting, but it is this
1045 module's arithmetic over the config, not a reading of the wire: for a Google
1046 seat at ``effort = high`` whose live model listing forced the adapter back to
1047 ``gemini-3-pro``, this returns ``gemini-3-pro-high``. Under ``requested``
1048 that is #709 restated — a model the run did not send, under a token whose
1049 whole claim is that it did.
1050 """
1051 name = (agent_name or "").strip()
1052 if not name:
1053 return "", MODEL_NONE
1054 spec = _spec_for(config, name)
1055 if spec is None:
1056 return f"unknown (no agent named '{name}' in this run's config)", MODEL_UNKNOWN
1057 sent = sent_model(result)
1058 if sent:
1059 return sent, MODEL_REQUESTED
1060 recomputed = requested_model(spec)
1061 if recomputed:
1062 return recomputed, MODEL_RECOMPUTED
1063 command = (getattr(spec, "command", "") or getattr(spec, "vendor", "") or name).strip()
1064 return _CLI_DEFAULT_MODEL.format(command=command), MODEL_CLI_DEFAULT
1067def _vendor_for(config: Any, agent_name: str) -> str:
1068 for spec in getattr(config, "agents", []) or []:
1069 if spec.name == agent_name:
1070 return spec.vendor or ""
1071 return ""
1074def participating(outcome: Any) -> list:
1075 """Every round-1 seat that ran, in the stable panel order — one ballot each.
1077 It used to be "seats that returned output at all", which quietly dropped a
1078 silent agent from the bundle: an `alpha` result with empty output left no
1079 `alpha` entry, and the report could not say which seat had returned nothing
1080 (#700, round 2). It is recorded as an abstention naming the seat and the
1081 reason instead, and :func:`ai_jury.panel.is_review` — not this function's
1082 length — is what decides whether it counts as a review.
1084 That is the change from #699, where the length of this *was* the count.
1085 Recording a seat and counting it as a review are now two different questions,
1086 because a seat can ballot without reviewing; the count lives in
1087 :mod:`ai_jury.panel` and reads the produced records.
1088 """
1089 return ballot_seats(getattr(outcome, "reviews", []) or [])
1092def _stance_by_reviewer(outcome: Any, names: list[str], mode: str) -> dict[str, str]:
1093 """Per-panelist stance, derived exactly as the vote tally derives ballots.
1095 :func:`ai_jury.voting.tally_votes` is the single source of truth for turning
1096 "the worst supported finding this reviewer raised" into a stance, so it is
1097 called rather than reimplemented — a second copy of that mapping is a second
1098 place for the severity thresholds to drift.
1099 """
1100 result = tally_votes(getattr(outcome, "groups", []) or [], names, mode=mode)
1101 return {b.reviewer: normalize_verdict(b.vote) for b in result.ballots}
1104def _verdict_for(result: Any, stances: dict[str, str], *, scoped: bool) -> str:
1105 """This panelist's stance, or :data:`ABSTAIN` (pure).
1107 ``scoped`` is the third way to abstain and the one #700 added: a seat that
1108 exited 0, said something, and named nothing checkable — including, since
1109 round 2, one whose only "scope" was a claim raised against no file in a
1110 code review. The other two — a failed adapter, an empty reply or a refusal —
1111 were already here, and this is the same principle applied one step further
1112 out. A reviewer whose scope is empty raised no findings *with a location*
1113 either, so the tally would have handed it the clear stance
1114 (``APPROVE``/``READY``): an approval inferred from silence, which is
1115 precisely what :mod:`ai_jury.voting` refuses to do (#251).
1116 """
1117 if not getattr(result, "ok", False):
1118 return ABSTAIN
1119 if is_abstention(getattr(result, "output", "")):
1120 return ABSTAIN
1121 if not scoped:
1122 return ABSTAIN
1123 return stances.get(getattr(result, "agent", ""), ABSTAIN)
1126def chair_verdict(outcome: Any, vote: Any = None) -> str:
1127 """The run's final verdict as one machine token.
1129 The panel vote when voting; otherwise the label the chair opened its
1130 synthesis with. The headline lift is
1131 :func:`ai_jury.report._verdict_headline` — reused rather than duplicated, so
1132 the JSON verdict and the markdown TL;DR can never disagree. The headline is a
1133 label plus a sentence (``REQUEST CHANGES — one confirmed major issue.``); only
1134 the label is a verdict, so the sentence is dropped.
1135 """
1136 from .report import _verdict_headline
1138 headline = _verdict_headline(getattr(outcome, "synthesis", None), vote)
1139 if not headline:
1140 return ABSTAIN
1141 label = re.split(r"[—–:.]| - ", headline, maxsplit=1)[0]
1142 return normalize_verdict(label) or ABSTAIN
1145#: The clause each cause contributes to a *scoped* abstention's sentence. Only
1146#: two causes can reach it: a seat whose scope stands named something, so it was
1147#: neither silent nor a reply that named nothing.
1148_SCOPED_ABSTENTION_REASONS = {
1149 ADAPTER_FAILED: "its adapter reported failure",
1150 REFUSED: "it returned a refusal rather than a review",
1151}
1154def _abstained_because(result: Any) -> str:
1155 """Why a seat that *did* name something checkable still abstained (pure).
1157 The companion to :func:`_no_scope_reason`, and deliberately not that
1158 function: this ballot's scope is substantive, so every reason phrased around
1159 "it named nothing" would be false of it. Only two of :func:`_verdict_for`'s
1160 three gates can fire while the scope stands — a failed adapter and a refusal
1161 — and the other two causes never reach here, because a seat with no scope
1162 gets :func:`abstention_scope` instead. Read from
1163 :func:`abstention_cause` all the same, so this sentence and the bucket the
1164 same seat is counted in cannot name two different things.
1165 """
1166 return _SCOPED_ABSTENTION_REASONS.get(abstention_cause(result), "it cast no vote")
1169def _chaired_ballot_sentence(result: Any, ballot: dict) -> str:
1170 """What the chairing agent's own ballot is, read off the record (#700, round 3).
1172 Said on the ballot as well as on the chair record, because the two are read
1173 in different places: a consumer that posts one verdict per review shows this
1174 text on its own, with no chair record beside it.
1176 The sentence is a function of ``counts_as_review`` — the answer
1177 :func:`ai_jury.panel.is_review` already gave for this record — and never of
1178 "did a scope come back". Round 2 keyed it on the scope alone and appended it
1179 before the verdict was in, so a chaired seat that named a file and then
1180 refused shipped ``verdict: ABSTAIN``, ``counts_as_review: false`` and a scope
1181 telling the human reading it that this ballot was one of the panel's reviews.
1182 The count and the prose beside it are one statement, or the prose is a second
1183 definition of a review that nothing keeps honest.
1184 """
1185 lead = " This reviewer also chaired the run (verification and synthesis);"
1186 tail = " the chair record is the panel's consensus rather than a further review."
1187 if ballot.get("counts_as_review"):
1188 return f"{lead} this ballot is one of the panel's reviews, and{tail}"
1189 return (
1190 f"{lead} this ballot abstained — {_abstained_because(result)} — so it is"
1191 f" NOT one of the panel's reviews, and{tail}"
1192 )
1195def _verified_count(outcome: Any, name: str) -> int:
1196 """Consensus groups this reviewer contributed to that the verifier upheld."""
1197 return sum(
1198 1
1199 for g in getattr(outcome, "groups", []) or []
1200 if name in (getattr(g, "reviewers", []) or [])
1201 and (getattr(g, "status", "") or "") == "verified"
1202 )
1205def reviewer_ballots(
1206 outcome: Any, config: Any, *, vote: Any = None, mode: str = "code"
1207) -> list[dict]:
1208 """The JSON report's ``reviewers`` array: one ballot per seat, then the chair.
1210 Panelist entries carry ``name``, ``role: "panelist"``, ``chaired``,
1211 ``vendor``, ``model``, ``model_source``, ``verdict``, ``scope``,
1212 ``scope_substantive``, ``counts_as_review``, ``testing``, ``findings``
1213 (indexes into the report's top-level ``findings`` array), ``round1_ok``,
1214 ``verified_count`` and ``duration_s``. The chair's entry is the one carrying
1215 ``role: "chair"`` and is always last.
1217 ``role`` is what a consumer splits on: the ``chair`` entry is the panel's
1218 consensus record and every other entry is a ballot (#699). But a ballot is
1219 not automatically a review — ``scope_substantive`` and ``verdict`` are the
1220 two facts :func:`ai_jury.panel.is_review` reads, and ``counts_as_review``
1221 carries its answer so the consumer need not re-derive it. Every seat that ran
1222 gets an entry, a silent one included, because the report has to be able to
1223 say *which* agent returned nothing; the count is what excludes it.
1225 ``chaired`` and the chair entry's ``agent``/``ballot_counted`` exist because
1226 the chairing agent reviews too: without them a reader cannot tell that the
1227 ``claude`` ballot and the ``chair`` record are the same agent, nor whether
1228 that agent contributed a review at all — and a reader who guesses drops a
1229 review the panel cast.
1231 ``scope`` and ``testing`` live here rather than only in the bundle (#700) so
1232 that the two renderings are the same text by construction: the JSON report
1233 said who voted and the bundle said what they read, and a reader comparing
1234 them had no guarantee the second described the first.
1235 """
1236 seats = participating(outcome)
1237 # What the panel was actually shown (#710), so a reviewer's `Checked:` line
1238 # is resolved against the change rather than accepted on its shape. ``None``
1239 # on an outcome that was not built from a diff; the rule then falls back to
1240 # the structural test, because "not verifiable" is not "does not exist".
1241 #
1242 # ``--issue`` is the one mode that supplies no change to resolve against: the
1243 # panel is reading an issue's prose, and a reviewer naming "the acceptance
1244 # criteria section" has named exactly what it was asked to name. That is the
1245 # same exception :func:`_claims_named` documents — there a finding carries no
1246 # file by construction — and applying a diff rule to a document with no diff
1247 # would abstain over a clean triage that did its job.
1248 changed = None if mode == ISSUE_MODE else getattr(outcome, "changed", None)
1249 names = [getattr(r, "agent", "") for r in seats]
1250 stances = _stance_by_reviewer(outcome, names, mode)
1251 all_findings = list(getattr(outcome, "findings", []) or [])
1252 chair_name = getattr(outcome, "chair", "") or ""
1254 entries: list[dict] = []
1255 for r in seats:
1256 name = getattr(r, "agent", "")
1257 chaired = bool(chair_name) and name == chair_name
1258 indexes = [i for i, f in enumerate(all_findings) if f.reviewer == name]
1259 own_findings = [all_findings[i] for i in indexes]
1260 scope = describe_scope(r, own_findings, mode=mode, changed=changed)
1261 scoped = bool(scope)
1262 if not scoped:
1263 scope = abstention_scope(r, own_findings, changed)
1264 model, model_source = describe_model(config, name, r)
1265 entry = {
1266 "name": name,
1267 "role": PANELIST_ROLE,
1268 "chaired": chaired,
1269 "vendor": getattr(r, "vendor", "") or "",
1270 "model": model,
1271 "model_source": model_source,
1272 "verdict": _verdict_for(r, stances, scoped=scoped),
1273 "scope": scope,
1274 # The two facts the count is made of, stated structurally rather than
1275 # left to be inferred from the prose in ``scope`` — a consumer that
1276 # had to parse English for this is a consumer that will get it wrong.
1277 "scope_substantive": scoped,
1278 "testing": describe_testing(r),
1279 "findings": indexes,
1280 "round1_ok": bool(getattr(r, "ok", False)),
1281 "verified_count": _verified_count(outcome, name),
1282 "duration_s": round(float(getattr(r, "duration_s", 0.0) or 0.0), 3),
1283 }
1284 # Derived, never asserted: the answer this record carries is the one the
1285 # gate and the announcements use, computed by the same function.
1286 entry["counts_as_review"] = is_review(entry)
1287 # And, when the answer is no, *why* — the fact a reader of the record
1288 # alone cannot recover, because a seat that fell silent and one that
1289 # answered without naming anything leave the same two fields behind. It
1290 # travels with the ballot so every renderer counts and describes the same
1291 # seat the same way instead of subtracting one bucket from another
1292 # (#700, round 5). Empty on a ballot that reviewed: it has no cause.
1293 entry[CAUSE_FIELD] = "" if entry["counts_as_review"] else abstention_cause(r, changed)
1294 # After the answer, never before it: the chair's sentence *reports* that
1295 # answer, so it cannot be written while the answer is still unknown.
1296 if scoped and chaired:
1297 entry["scope"] += _chaired_ballot_sentence(r, entry)
1298 entries.append(entry)
1300 # The chair's own round-1 seat, when it has one: the chairing agent reviews
1301 # too, and every phase of it runs through the same adapter, so the id its
1302 # ballot recorded is the id its synthesis was produced with.
1303 chair_result = next((r for r in seats if getattr(r, "agent", "") == chair_name), None)
1304 chair_model, chair_model_source = describe_model(config, chair_name, chair_result)
1305 verify = getattr(outcome, "verify", None)
1306 chair_scope = _chair_scope(outcome, entries)
1307 entries.append(
1308 {
1309 "name": CHAIR_NAME,
1310 "role": CHAIR_ROLE,
1311 "agent": chair_name,
1312 # Whether the chairing agent's own ballot is one of the counted
1313 # reviews — the fact the bundle never stated (#699). A chairing agent
1314 # that ran and abstained has a ballot in the bundle and is not a
1315 # review, so "did it ballot" is the wrong question to answer here.
1316 "ballot_counted": any(e["counts_as_review"] for e in entries if e["chaired"]),
1317 "reviews_supplied": review_count(entries),
1318 "vendor": _vendor_for(config, chair_name),
1319 "model": chair_model,
1320 "model_source": chair_model_source,
1321 "verdict": chair_verdict(outcome, vote),
1322 "scope": chair_scope,
1323 # Measured, not asserted: the chair record is a verdict too, and a
1324 # consumer applying one substance rule applies it here as well.
1325 "scope_substantive": scope_is_substantive(chair_scope),
1326 # This synthesis record is NOT one of the reviews, whatever its
1327 # scope says: the consumer reads it as the panel's consensus.
1328 "counts_as_review": False,
1329 # Empty, and present rather than absent: the chair is not a seat that
1330 # failed to review, it is the record carried alongside the seats, and
1331 # a key missing here would read as a cause nobody wrote down.
1332 CAUSE_FIELD: "",
1333 "testing": describe_testing(verify) if verify is not None else NOT_STATED,
1334 }
1335 )
1336 return entries
1339def _keel_finding(f: Any) -> dict:
1340 """One finding in the consumer's shape: ``file`` → ``path``, ``claim`` → ``message``."""
1341 line = getattr(f, "line", None)
1342 return {
1343 "severity": getattr(f, "severity", "") or "",
1344 "path": getattr(f, "file", "") or "",
1345 "line": line if isinstance(line, int) else None,
1346 "message": flatten_inline(getattr(f, "claim", "") or ""),
1347 }
1350def _chair_findings(outcome: Any) -> list:
1351 """The chair's surviving evidence: group representatives the verifier did not reject."""
1352 return [
1353 g.representative
1354 for g in getattr(outcome, "groups", []) or []
1355 if (getattr(g, "status", "") or "") != "unsupported"
1356 ]
1359def _chair_role_sentence(chair_agent: str, chaired_ballot: dict | None) -> str:
1360 """State, in the record itself, whether this chair's ballot is counted (#699).
1362 A chair that reviewed and a chair that only synthesised produce records that
1363 are otherwise identical, and a reader who cannot tell them apart guesses —
1364 which is how a panel's ballots get handed on short. So the record says which
1365 agent chaired and where (or whether) its own ballot is in the bundle, and it
1366 says plainly that this synthesis record is not itself one of the reviews.
1367 """
1368 who = f"`{chair_agent}`" if chair_agent else "This run's chair"
1369 if chaired_ballot is not None and chaired_ballot.get("counts_as_review"):
1370 return (
1371 f"The chair is {who}, which also sat on the panel: its ballot is the "
1372 f"'{chaired_ballot['name']}' review in this bundle and counts as one "
1373 f"of the reviews. This synthesis record does not — it is the panel's "
1374 f"consensus, not a ballot."
1375 )
1376 if chaired_ballot is not None:
1377 # It balloted and the ballot is not a review. Saying "no ballot from it"
1378 # would be false and saying "its ballot counts" would be the defect, so
1379 # the record says both facts. It does *not* say why: "named nothing
1380 # checkable" was one of the three ways to abstain asserted as if it were
1381 # the only one, and it is plainly false of a seat that named a file and
1382 # then refused. The ballot's own scope carries the cause (#700, round 3).
1383 return (
1384 f"The chair is {who}, which also sat on the panel but abstained: its "
1385 f"'{chaired_ballot['name']}' ballot is in this bundle and is NOT counted "
1386 f"as a review — that ballot's own scope says why. This "
1387 f"synthesis record is not counted either — it is the panel's consensus, "
1388 f"not a ballot."
1389 )
1390 return (
1391 f"The chair is {who}, which returned no panel review of its own, so this "
1392 f"bundle carries no ballot from it. This synthesis record is not counted "
1393 f"as a review either — it is the panel's consensus, not a ballot."
1394 )
1397def _chair_scope(outcome: Any, panelists: list[dict]) -> str:
1398 """What the chair's synthesis covered (pure).
1400 Backticked file names, like every other scope here (#700): the chair record
1401 is a verdict too, and a consumer applying one substance rule applies it to
1402 this record as well. The chairing agent's name is already backticked by
1403 :func:`_chair_role_sentence`, so a chaired run anchors either way.
1405 The two numbers it quotes are deliberately different (#700, round 2): how
1406 many of the ballots are **reviews**, and how many **records** the bundle
1407 carries. They used to be the same integer, which is exactly how an abstaining
1408 ballot got announced as a review.
1409 """
1410 groups = getattr(outcome, "groups", []) or []
1411 files = _files_named([g.representative for g in groups])
1412 listed = ", ".join(_tick(f) for f in files[:_FILES_LISTED]) if files else "no specific file"
1413 more = len(files) - _FILES_LISTED
1414 suffix = f" (+{more} more)" if more > 0 else ""
1415 chair_agent = getattr(outcome, "chair", "") or ""
1416 chaired = next((p for p in panelists if p.get("chaired")), None)
1417 reviews = review_count(panelists)
1418 # "Further" than the chair's own ballot, which the sentence before this one
1419 # has already accounted for — and no cause is named, because these seats
1420 # abstained for whichever of the three reasons applied to each and their own
1421 # scopes say which (#700, round 3).
1422 abstained = sum(1 for p in panelists if not p.get("chaired") and not is_review(p))
1423 abstained_clause = (
1424 f" {abstained} further ballot(s) abstained; they are recorded here but"
1425 f" are not reviews, and each one's scope says why."
1426 if abstained
1427 else ""
1428 )
1429 return (
1430 f"Chair synthesis over {reviews} panel review(s) and "
1431 f"{len(groups)} consensus group(s), across {listed}{suffix}. "
1432 f"{_chair_role_sentence(chair_agent, chaired)}{abstained_clause} This bundle carries "
1433 f"{reviews} review(s) plus this record, "
1434 f"{bundle_records(len(panelists))} records in all."
1435 )
1438def keel_reviews(outcome: Any, config: Any, *, vote: Any = None, mode: str = "code") -> list[dict]:
1439 """The ``--format keel-reviews`` bundle: one record per seat, plus the chair.
1441 Each record is ``{reviewer, verdict, scope, findings, testing, vendor, model,
1442 model_source, counts_as_review}`` where ``findings`` are ``{severity, path,
1443 line, message}`` objects — the shape a consumer of head-pinned per-reviewer
1444 verdicts accepts. Pure: the caller serializes it.
1446 A projection of :func:`reviewer_ballots`, not a second derivation of the same
1447 facts (#700). Every field but ``findings`` is renamed or copied straight
1448 across, so the verdict the JSON report shows for a panelist and the verdict
1449 the consumer is handed for it cannot disagree — and neither can the scope
1450 that is supposed to justify it.
1452 Two of those fields are the projection catching up with the ballot (#700,
1453 round 2). ``model_source`` because ``model`` here changed meaning in the same
1454 release — a CLI default is now an English sentence — and a machine consumer
1455 of *this* shape could not tell a requested id from a default without parsing
1456 prose; the ``reviewers`` array grew the discriminator and this one did not.
1457 ``counts_as_review`` because the bundle now carries abstention records for
1458 seats that returned nothing, and a consumer counting the array would count
1459 them.
1460 """
1461 ballots = reviewer_ballots(outcome, config, vote=vote, mode=mode)
1462 all_findings = list(getattr(outcome, "findings", []) or [])
1464 records: list[dict] = []
1465 for b in ballots:
1466 chair = b.get("role") == CHAIR_ROLE
1467 own = _chair_findings(outcome) if chair else [all_findings[i] for i in b["findings"]]
1468 records.append(
1469 {
1470 "reviewer": b["name"],
1471 "verdict": b["verdict"],
1472 "scope": b["scope"],
1473 "findings": [_keel_finding(f) for f in own],
1474 "testing": b["testing"],
1475 "vendor": b["vendor"],
1476 "model": b["model"],
1477 "model_source": b["model_source"],
1478 "counts_as_review": b["counts_as_review"],
1479 }
1480 )
1481 return records