Coverage for src/ai_jury/panel.py: 100%

60 statements  

« prev     ^ index     » next       coverage.py v7.16.1, created at 2026-09-30 06:29 +0000

1"""Panel arithmetic: how many reviews a consumer actually receives (#699, #700). 

2 

3A downstream gate does not ask "how many agents ran"; it asks "how many reviews 

4did I get". Those two numbers were never the same and nothing in the tool ever 

5said so, which is how a healthy three-CLI bench was reported as ``cross-vendor 

6ready: yes``, exited 0, and then failed at the consumer with *supplied 2 

7review(s) but tier requires at least 3*. 

8 

9Four facts settle the arithmetic, and all four live here so that every renderer, 

10``--doctor`` and the run itself quote the same one: 

11 

12* **A review is a ballot that reviewed.** :func:`is_review` is the definition, 

13 and it is the only one: a ``panelist`` record whose scope names something a 

14 reader could go and check, and whose verdict is a vote rather than 

15 :data:`ABSTAIN`. Every seat that ran gets a ballot record — that is how the 

16 report can say *which* seat returned nothing — but a ballot that named nothing 

17 is the record of a seat that did **not** review, and counting it supplies the 

18 consumer a review it will refuse. The consumer's own 

19 ``review-verdict-insubstantial`` rule rejects exactly that shape, so a count 

20 including it would be #699's mismatch wearing better prose: a number ai-jury 

21 states and the consumer does not find. 

22* **The chair's synthesis is not a review.** The consumer splits the report's 

23 ``reviewers`` array on ``role``: the entry carrying ``role: "chair"`` becomes 

24 the panel's consensus record, and only a panelist entry can be a review. The 

25 chair's synthesis is carried *alongside* the reviews, never as one of them. 

26* **The chair sits on the panel.** :func:`ai_jury.orchestrator.resolve_chair` 

27 only ever picks a chair from the *usable* agents, and round 1 runs every usable 

28 agent, so the chairing agent reviews on every code path. That ballot is an 

29 ordinary panelist record carrying no chair role, and the consumer counts it 

30 like any other. What was missing was any statement of that: a reader of the 

31 bundle could not tell whether the chairing agent had also voted, and one who 

32 guesses drops a review the panel actually cast. 

33* **A seat can ballot without reviewing.** An agent that returns nothing, refuses, 

34 fails, or answers with prose naming no file, line or symbol is recorded as an 

35 abstention naming *which* seat and *why* — visible in the bundle, and excluded 

36 from the count. So the number of reviews is knowable only as an upper bound 

37 before the run, which is why the gate is checked on both sides of it. 

38* **Every seat is in exactly one bucket, and the bucket is the cause.** 

39 :func:`abstention_buckets` splits the non-reviewing seats five ways — silent, 

40 named nothing, named only what is not in the change, refused, adapter failed — 

41 by reading each ballot's own recorded cause. It replaces a subtraction (``ballots - silent - supplied``) that was 

42 arithmetically right and descriptively false: once a ballot could carry a 

43 substantive scope and still abstain, the remainder held seats that had named a 

44 file and then refused, and every renderer billed them as having named nothing 

45 checkable — a cause the ballot beside the number contradicted (#700, round 5). 

46 

47Everything here is pure: no I/O, no clock, no randomness. 

48""" 

49 

50from __future__ import annotations 

51 

52from collections.abc import Mapping 

53from typing import Any 

54 

55#: Stance recorded for a seat that did not review. Defined here rather than in 

56#: :mod:`ai_jury.ballots` because :func:`is_review` — the one definition of what 

57#: gets counted — has to test for it, and a second copy of the token is a second 

58#: place for the count and the record to disagree. ``ballots`` re-exports it. 

59ABSTAIN = "ABSTAIN" 

60 

61#: The two values of a ``reviewers`` entry's ``role``. A consumer splits on this. 

62PANELIST_ROLE = "panelist" 

63CHAIR_ROLE = "chair" 

64 

65#: Records the bundle carries that are **not** reviews: the chair's synthesis. 

66#: Named rather than inlined as ``+ 1`` so that prose about the bundle's *shape* 

67#: and prose about its *review count* cannot quietly become the same number. 

68CHAIR_SYNTHESIS_RECORDS = 1 

69 

70#: Why a seat that balloted did not review. Five causes, because five different 

71#: things happen to a seat and they ask for five different fixes. 

72SILENT = "silent" 

73NAMED_NOTHING = "named_nothing" 

74#: The seat named something and none of it is in the change (#710). Split out of 

75#: ``named_nothing`` because the two send their reader to opposite places: a 

76#: reviewer that named nothing did not say what it read, while one that named 

77#: `src/made/up.py` said it read a file that is not in the diff — a claim about 

78#: coverage that the change itself contradicts, and a much louder signal. 

79NOT_IN_CHANGE = "not_in_change" 

80REFUSED = "refused" 

81ADAPTER_FAILED = "adapter_failed" 

82 

83#: The causes in the order every renderer lists them, and the key set of 

84#: :func:`abstention_buckets` — so a renderer cannot iterate a cause the buckets 

85#: do not carry, nor miss one they do. 

86ABSTENTION_CAUSES = (SILENT, NAMED_NOTHING, NOT_IN_CHANGE, REFUSED, ADAPTER_FAILED) 

87 

88#: The clause each cause contributes to a count sentence, subject ``"N "``. One 

89#: table, because the markdown report and :func:`shortfall` describe the same 

90#: seats and used to phrase them apart — and the phrasing was where they lied. 

91#: Every clause is true of *every* ballot in its bucket; in particular none of 

92#: them says the seat named nothing, except the one bucket where that is what 

93#: happened. "N named nothing checkable" asserted over a seat that named a file 

94#: and *then* refused is the defect this table exists to make impossible 

95#: (#700, round 5). 

96CAUSE_PHRASES = { 

97 SILENT: "returned nothing at all", 

98 NAMED_NOTHING: "answered but named nothing checkable", 

99 NOT_IN_CHANGE: "named only things that are not in this change", 

100 REFUSED: "returned a refusal rather than a review", 

101 ADAPTER_FAILED: "reported an adapter failure", 

102} 

103 

104#: The run-metadata key each cause is published under. Kept here because three 

105#: modules read those keys and one of them is not spelled like its cause: 

106#: ``named_nothing`` ships as ``insubstantial``, the name it has had since #700's 

107#: second round, and renaming a published key to tidy a table would break every 

108#: consumer reading it. A renderer iterating this mapping cannot leave a cause 

109#: out, which is what a hand-written pair of ``if`` statements did. 

110PANEL_METADATA_KEYS = { 

111 SILENT: "silent", 

112 NAMED_NOTHING: "insubstantial", 

113 NOT_IN_CHANGE: "not_in_change", 

114 REFUSED: "refused", 

115 ADAPTER_FAILED: "adapter_failed", 

116} 

117 

118#: The key a ballot record carries its cause under. Read from the record, never 

119#: recomputed from it: whether a seat fell silent or refused is a fact about the 

120#: result that produced the ballot, and no predicate over the record alone can 

121#: tell those two apart — both are ``scope_substantive: false`` when the refusal 

122#: named nothing, which is why subtraction was reached for in the first place. 

123CAUSE_FIELD = "abstention_cause" 

124 

125 

126def responded(result) -> bool: 

127 """Did this round-1 seat return any output at all? 

128 

129 "Returned output", not "exited 0": adapters fail soft, so a nonzero exit can 

130 still carry a complete review on stdout. A seat with nothing at all still 

131 gets a ballot — an abstention naming it, so the report can say which seat 

132 fell silent (#700, round 2) — but this predicate is what lets that record 

133 state *which* kind of nothing came back. 

134 """ 

135 return bool((getattr(result, "output", "") or "").strip()) 

136 

137 

138def ballot_seats(reviews) -> list: 

139 """The round-1 seats that produce a ballot, in the stable panel order. 

140 

141 **Every** seat that ran. A seat that returned nothing used to be dropped 

142 here, which left the bundle unable to say which agent had gone silent: the 

143 report simply had one fewer entry than the bench, and a reader comparing the 

144 two had to guess which. It is recorded as an abstention instead, and 

145 :func:`is_review` is what keeps it out of the count. 

146 """ 

147 return list(reviews or []) 

148 

149 

150def is_review(ballot: Mapping[str, Any]) -> bool: 

151 """**The** definition: does this ballot record count as a review? (pure) 

152 

153 Three conditions, and each one is a defect this project has already shipped: 

154 

155 * ``role == "panelist"`` — the chair's synthesis record is the panel's 

156 consensus, not an *n+1*-th review (#699). 

157 * ``scope_substantive`` — the ballot names something a reader could go and 

158 check. A prose-only reply (*"Looks good to me, no concerns."*) names 

159 nothing, and the consumer refuses a verdict shaped like that (#700). 

160 * ``verdict != ABSTAIN`` — the seat actually cast a vote. An abstention is 

161 not an approval (#251), and it is not a review either. 

162 

163 Read from the record rather than recomputed from the agent result, because 

164 the record is what the consumer is handed: if the two could disagree, the 

165 number ai-jury announces would once again describe something other than the 

166 document it shipped. 

167 """ 

168 if not isinstance(ballot, Mapping): # pragma: no cover - defensive 

169 return False 

170 return ( 

171 (ballot.get("role") or "") == PANELIST_ROLE 

172 and bool(ballot.get("scope_substantive")) 

173 and (ballot.get("verdict") or "") != ABSTAIN 

174 ) 

175 

176 

177def review_count(ballots) -> int: 

178 """Reviews a consumer receives from these ballot records. 

179 

180 The one count. Every announcement, both halves of the ``min_reviews`` gate, 

181 the markdown report, the chair record's ``reviews_supplied`` and ``--doctor`` 

182 resolve to this, so the number ai-jury states and the number the consumer 

183 finds in the bundle are equal by construction. 

184 

185 Takes the **ballots**, not the raw round-1 results: whether a seat reviewed 

186 is a fact about the record it produced — its scope and its verdict — and no 

187 predicate over the raw result can see it. 

188 """ 

189 return sum(1 for b in ballots or [] if is_review(b)) 

190 

191 

192def panelist_ballots(ballots) -> list: 

193 """The ballot records that are *seats* — everything but the chair's synthesis. 

194 

195 The unit the buckets and the count are both taken over, so "how many seats 

196 balloted" and "how many of them reviewed" can never be measured against two 

197 different populations. 

198 """ 

199 return [ 

200 b 

201 for b in ballots or [] 

202 if isinstance(b, Mapping) and (b.get("role") or "") == PANELIST_ROLE 

203 ] 

204 

205 

206def abstention_cause(ballot: Mapping[str, Any]) -> str: 

207 """Why this ballot is not a review — one of :data:`ABSTENTION_CAUSES` (pure). 

208 

209 **Read** from :data:`CAUSE_FIELD` rather than recomputed. Which of the four 

210 happened is a fact about the result that produced the ballot: a seat that 

211 fell silent and one that answered with nothing checkable leave the same two 

212 structural fields behind, so a record that does not carry its cause has lost 

213 it. The ballot is what the consumer is handed, so the cause travels with it. 

214 

215 A record written by an older build — or by hand, as tests do — carries no 

216 cause, and is read from the fields every ballot has rather than refused: a 

217 failed adapter says so in ``round1_ok``, and an empty ``scope_substantive`` 

218 means nothing checkable was named. That reading cannot see silence, which is 

219 exactly why the field exists; it errs toward ``named_nothing``, which is true 

220 of a silent seat as well. 

221 """ 

222 if not isinstance(ballot, Mapping): # pragma: no cover - defensive 

223 return NAMED_NOTHING 

224 recorded = (ballot.get(CAUSE_FIELD) or "").strip() 

225 if recorded in ABSTENTION_CAUSES: 

226 return recorded 

227 if not ballot.get("round1_ok", True): 

228 return ADAPTER_FAILED 

229 if not ballot.get("scope_substantive"): 

230 return NAMED_NOTHING 

231 return REFUSED 

232 

233 

234def abstention_buckets(ballots) -> dict[str, int]: 

235 """The seats that balloted without reviewing, counted by cause (pure). 

236 

237 The five counts and :func:`review_count` partition the panelist ballots: 

238 every seat lands in exactly one, because each is classified by what happened 

239 to it. Subtraction is what this replaces, and what subtraction shipped: with 

240 ``insubstantial`` derived as ``ballots - silent - supplied``, every seat that 

241 was neither silent nor a counted review was billed as "named nothing 

242 checkable" — including, since a ballot could carry a substantive scope and 

243 still abstain, the seat that named a file and *then* refused, and the one 

244 whose adapter died holding a file name. The number was right and the cause 

245 printed beside it was contradicted by the ballot it described (#700, round 5). 

246 """ 

247 counts = dict.fromkeys(ABSTENTION_CAUSES, 0) 

248 for ballot in panelist_ballots(ballots): 

249 if not is_review(ballot): 

250 counts[abstention_cause(ballot)] += 1 

251 return counts 

252 

253 

254def bundle_records(ballots: int) -> int: 

255 """Total records in the bundle: the ballots, plus the chair's synthesis. 

256 

257 Deliberately distinct from the review count, and now distinct from it in two 

258 directions: the chair record is one more record than there are ballots, and 

259 a ballot is not necessarily a review. It exists for prose describing the 

260 bundle's *shape* ("this bundle carries N records") and must never answer 

261 "how many reviews", which is what :func:`review_count` is for. 

262 """ 

263 return int(ballots) + CHAIR_SYNTHESIS_RECORDS 

264 

265 

266def describe(reviews: int, *, available: int | None = None) -> str: 

267 """One line stating the number a consumer will actually receive. 

268 

269 Rendered before the panel runs, from the agents that are reachable — which 

270 can only be an **upper bound**, and now says so: every seat is a seat that 

271 *might* review, and one that returns nothing or names nothing checkable is 

272 recorded as an abstention rather than a review. Passing ``available`` 

273 selects that pre-run wording; without it the number is the count that 

274 actually materialised. 

275 

276 The chair record is named in the same breath, and named as *not* a review, 

277 because this line exists to forestall exactly the arithmetic that would add 

278 it. 

279 """ 

280 tail = f"plus {CHAIR_SYNTHESIS_RECORDS} chair synthesis record (not counted as a review)" 

281 if available is not None: 

282 return ( 

283 f"panel: {available} available agent(s) → at most {available} review(s) for a " 

284 f"downstream consumer (one per seat that names what it read and votes; a seat " 

285 f"that returns nothing, or names nothing checkable, abstains and is not one), " 

286 f"{tail}" 

287 ) 

288 return f"panel: {reviews} review(s) for a downstream consumer, {tail}" 

289 

290 

291def shortfall( 

292 supplied: int, 

293 required: int, 

294 *, 

295 stage: str, 

296 silent: int = 0, 

297 insubstantial: int = 0, 

298 not_in_change: int = 0, 

299 refused: int = 0, 

300 adapter_failed: int = 0, 

301) -> str | None: 

302 """Why this run cannot supply ``required`` reviews, or ``None`` if it can. 

303 

304 ``stage`` is the phrase naming when the check fired, so the three call sites 

305 ("before the panel runs" / "after the panel ran" / "on this machine") read as 

306 one message with the timing filled in rather than as unrelated errors. 

307 

308 The five counts are :func:`abstention_buckets` — the ways a seat that ran 

309 produces no review — and they are named separately because the remedies 

310 differ: a silent agent is usually a CLI that broke or a budget that ran out, 

311 a seat that answered and named nothing is a reviewer that did not review, a 

312 seat that named only things absent from the change reviewed something that 

313 is not this diff, a refusal is a model declining the task, and a failed 

314 adapter is an invocation to fix. ``insubstantial`` keeps its name for 

315 ``named_nothing`` and means exactly that. Pass them as the buckets report 

316 them: a caller that folds causes into one prints a sentence the ballots 

317 contradict. 

318 """ 

319 required = int(required or 0) 

320 if required <= 0: 

321 return None 

322 supplied = int(supplied) 

323 if supplied >= required: 

324 return None 

325 counted = { 

326 SILENT: int(silent), 

327 NAMED_NOTHING: int(insubstantial), 

328 NOT_IN_CHANGE: int(not_in_change), 

329 REFUSED: int(refused), 

330 ADAPTER_FAILED: int(adapter_failed), 

331 } 

332 reasons = [ 

333 f"{counted[cause]} {CAUSE_PHRASES[cause]}" for cause in ABSTENTION_CAUSES if counted[cause] 

334 ] 

335 missing = sum(counted.values()) 

336 clause = ( 

337 f" {missing} seat(s) ran without producing a review " 

338 f"({'; '.join(reasons)}), so they are recorded as abstentions " 

339 f"and are not counted." 

340 if missing 

341 else "" 

342 ) 

343 return ( 

344 f"panel too small {stage}: {supplied} review(s) available " 

345 f"(the chair's synthesis record is not one of them, and neither is an " 

346 f"abstaining ballot) but min_reviews requires {required}.{clause} " 

347 f"Enable another agent, or lower `[jury.ci] min_reviews` / `--min-reviews`." 

348 )