Coverage for src/keel/juryavail.py: 100%

126 statements  

« prev     ^ index     » next       coverage.py v7.16.2, created at 2026-10-02 20:26 +0000

1"""Can the cross-vendor panel actually be staffed here? (#1066) 

2 

3On a tier whose ``knobs.team.review.by_tier`` names ``jury``, s7 dispatches the panel and 

4its ballots *are* the review — #1014 round 3 deliberately made it so no operator flag can 

5take the panel back off. That is right while the panel can run. When it cannot — an agent 

6CLI is not installed, is unauthenticated, or the account is out of quota — the tier has no 

7way forward at all: the only review it has is one this machine cannot convene. 

8 

9This module is the pure half of the answer. The question it has to answer is narrower than 

10"are some agent CLIs installed": s7 does not convene a panel out of keel's delegate 

11inventory, it runs the **``jury`` binary** (``src/keel/adapters/commands/ship.md``), and 

12that binary holds its own configured panel. So the probe asks the runner first — 

13``jury --doctor --json``, ai-jury's own readiness document, which reports both that the 

14binary is there and which of *its* agents are usable — and only falls back to 

15:func:`keel.providerprobe.collect` (what ``keel doctor --providers`` prints) for a runner 

16whose document named no agents. A machine with ``claude`` and ``codex`` on ``PATH`` and no 

17``jury`` is **not** staffable, however healthy keel's own inventory looks: the panel s7 

18would dispatch cannot run. 

19 

20That document is also the binary's *identity*, and no document is no identity (#1068): 

21a ``jury`` on ``PATH`` that exits 0 without one has not established that it is ai-jury, so 

22it is unusable and keel's inventory cannot make it staffable. The proxy stands in for a 

23panel ai-jury declined to enumerate, never for a panel runner nobody established is there. 

24 

25Two answers to one question is what this ordering avoids. keel's delegate inventory is a 

26proxy for the panel, ai-jury's is the panel; when the panel can speak for itself it is the 

27authority, and the record says which inventory the verdict was read from. 

28 

29Three things the design holds to, all of them from the issue: 

30 

31* **Availability is measured, never asserted.** There is no flag that says "the panel is 

32 fine". What may not take the panel off is an operator's *preference*; availability is a 

33 fact about the world, and it is allowed to change the outcome precisely because it is 

34 recorded. 

35* **The policy is a configured allowance, not an automatic behaviour.** 

36 ``knobs.team.jury.on_unavailable`` is ``fallback`` (the sensible default for a solo 

37 project) or ``block`` (today's strictness, for a project whose product claim *is* 

38 cross-vendor review). 

39* **Never a silent downgrade.** ai-jury #682 exists because a panel that quietly collapsed 

40 to one vendor still reported success. So :meth:`Availability.as_dict` carries which 

41 seats were unavailable and why, all the way into the published assignment, the review 

42 contract, the run ledger and the closure comment. The fallback changes *who sat*, never 

43 *how many*: the seat count and the evidence requirement are the tier's, not the panel's. 

44 

45Pure and deterministic: the report goes in, the verdict comes out, and nothing here 

46touches PATH, a subprocess, or the clock. 

47""" 

48 

49from __future__ import annotations 

50 

51from collections.abc import Mapping, Sequence 

52from dataclasses import dataclass 

53from typing import Any 

54 

55from .team import ( 

56 DEFAULT_MIN_VENDORS, 

57 JURY_ON_UNAVAILABLE_DEFAULT, 

58 jury_on_unavailable, 

59) 

60from .team import JURY_RUNNER_COMMAND as JURY_RUNNER_COMMAND 

61from .team import JuryUnavailableError as JuryUnavailableError 

62from .team import refusal_message as refusal_message 

63 

64#: The module's public surface, in definition order (#1070). It is declared because the 

65#: ``X as X`` re-exports above are read from *other* modules — a use CodeQL's 

66#: ``py/unused-import`` cannot see, since it counts same-module uses only. A name listed 

67#: in ``__all__`` is used by definition, so the declaration answers the scanner with the 

68#: language's own statement of intent rather than with a dismissal. Being a real 

69#: declaration it has to be the *whole* surface, not the re-exports alone; 

70#: ``tests/test_reexport_surface.py`` holds it to that in both directions. 

71__all__ = [ 

72 "JURY_RUNNER_COMMAND", 

73 "JuryUnavailableError", 

74 "refusal_message", 

75 "JURY_RUNNER_VENDOR", 

76 "INVENTORY_RUNNER", 

77 "INVENTORY_PROVIDERS", 

78 "DECISION_AVAILABLE", 

79 "DECISION_FALLBACK", 

80 "DECISION_BLOCK", 

81 "Seat", 

82 "Runner", 

83 "RUNNER_UNPROBED", 

84 "Availability", 

85 "assess", 

86 "SOURCE_PROBE", 

87 "SOURCE_PULL_REQUEST", 

88 "SOURCE_RUN_LEDGER", 

89 "SOURCE_CLOSURE_COMMENT", 

90 "panel_sat", 

91 "recorded", 

92 "shipped", 

93 "is_ship_run_for_head", 

94 "states_panel", 

95 "pin", 

96 "is_pinnable_head", 

97] 

98 

99#: Re-exported from :mod:`keel.team`, which owns them because :func:`_review_seats` — the 

100#: one place a bench is resolved, and so the one place a blocked panel can refuse the work 

101#: it is actually about to review — cannot import this module: the import runs the other 

102#: way. They keep their names here because this is the module the feature is named for and 

103#: where a reader looks for them; ``keel.juryavail.JuryUnavailableError`` and 

104#: ``keel.team.JuryUnavailableError`` are one class, not two. 

105#: 

106#: ``JURY_RUNNER_COMMAND`` is the binary a jury-panel tier's s7 actually dispatches. Not a 

107#: delegate keel runs itself: keel does not depend on ai-jury, and every path through this 

108#: module stays total when it is absent — absent simply means the panel cannot sit here. 

109 

110#: The vendor the runner seat is attributed to, so a reader of ``unavailable`` can tell the 

111#: missing *panel* apart from a missing *panelist*. 

112JURY_RUNNER_VENDOR = "ai-jury" 

113 

114#: Where a verdict's vendor inventory was read from — recorded, because the two sources do 

115#: not have to agree and a reader must not have to guess which one spoke. 

116INVENTORY_RUNNER = f"{JURY_RUNNER_COMMAND} --doctor" 

117INVENTORY_PROVIDERS = "keel doctor --providers" 

118 

119#: The panel is staffable — nothing changes, the ballots are the review. 

120DECISION_AVAILABLE = "available" 

121#: The panel is not staffable and the policy allows a host bench in its place. 

122DECISION_FALLBACK = "fallback" 

123#: The panel is not staffable and the policy refuses the run. 

124DECISION_BLOCK = "block" 

125 

126 

127@dataclass(frozen=True) 

128class Seat: 

129 """One provider the panel could have used, and why it cannot.""" 

130 

131 provider: str 

132 vendor: str 

133 reason: str 

134 

135 def as_dict(self) -> dict[str, Any]: 

136 return {"provider": self.provider, "vendor": self.vendor, "reason": self.reason} 

137 

138 

139@dataclass(frozen=True) 

140class Runner: 

141 """The ``jury`` CLI itself: can s7 dispatch it here, and what panel does it hold? 

142 

143 Produced by :func:`keel.providerprobe.probe_jury_runner` — the thin-I/O half — and 

144 consumed here. ``doctor`` is ai-jury's own ``jury --doctor --json`` document, which is 

145 both the panel it holds and the way the binary identifies itself: without a readable 

146 one the probe reports ``usable=False`` (#1068). A document that identifies ai-jury but 

147 names no ``agents`` still means *usable*, and the verdict then reads keel's own 

148 delegate inventory for the vendor count instead. 

149 """ 

150 

151 usable: bool 

152 reason: str 

153 doctor: Mapping[str, Any] | None = None 

154 

155 @property 

156 def panel_rows(self) -> tuple[Any, ...] | None: 

157 """The agents ai-jury reports for its own panel, or ``None`` when it reported none.""" 

158 return _rows(self.doctor, "agents") 

159 

160 def as_dict(self) -> dict[str, Any]: 

161 return {"command": JURY_RUNNER_COMMAND, "usable": self.usable, "reason": self.reason} 

162 

163 

164#: What an unprobed runner reads as. Fail-closed on purpose: the whole point of #1066 

165#: round 2 is that a panel nobody established could run must not be reported staffable, so 

166#: "we did not ask" and "we asked and it is fine" cannot share an answer. 

167RUNNER_UNPROBED = Runner(False, f"the {JURY_RUNNER_COMMAND} runner was not probed") 

168 

169 

170@dataclass(frozen=True) 

171class Availability: 

172 """The probe's verdict on this tier's panel, ready to publish.""" 

173 

174 #: Distinct vendors the panel needs before it can be a *cross-vendor* panel. This is 

175 #: ``team.jury.min_vendors``, the same floor the verdict is later held to. 

176 required_vendors: int 

177 #: Vendors the probe found usable here, in the probe's own (deterministic) order. 

178 available_vendors: tuple[str, ...] 

179 #: Every seat the panel could not use, with the reason it reported. The ``jury`` runner 

180 #: itself is one of them when it is the thing that is missing. 

181 unavailable: tuple[Seat, ...] 

182 #: ``fallback`` | ``block`` — the configured allowance, already defaulted. 

183 policy: str = JURY_ON_UNAVAILABLE_DEFAULT 

184 #: The ``jury`` binary s7 dispatches. Unprobed reads as unusable, never as fine. 

185 runner: Runner = RUNNER_UNPROBED 

186 #: Which inventory the vendor counts came from — the runner's own, or keel's. 

187 inventory: str = INVENTORY_PROVIDERS 

188 

189 @property 

190 def staffable(self) -> bool: 

191 """Can this machine convene a panel spanning ``required_vendors`` vendors? 

192 

193 Both halves, because s7 needs both: the runner that dispatches the panel, and 

194 enough distinct vendors for it to *be* a cross-vendor panel. Agent CLIs on ``PATH`` 

195 with no ``jury`` to convene them is not a panel, it is an inventory. 

196 """ 

197 return self.runner.usable and len(self.available_vendors) >= self.required_vendors 

198 

199 @property 

200 def decision(self) -> str: 

201 """:data:`DECISION_AVAILABLE`, or the policy when the panel cannot be staffed.""" 

202 return DECISION_AVAILABLE if self.staffable else self.policy 

203 

204 @property 

205 def reason(self) -> str: 

206 """One sentence a reader can act on, naming the seats that were unavailable.""" 

207 counted = ( 

208 f"{len(self.available_vendors)} vendor(s) available " 

209 f"({', '.join(self.available_vendors) or 'none'}), {self.required_vendors} " 

210 f"required (per {self.inventory})" 

211 ) 

212 if self.staffable: 

213 return f"jury panel staffable: {counted}, dispatched by {self.runner.reason}" 

214 listed = ", ".join(f"{seat.provider} ({seat.reason})" for seat in self.unavailable) 

215 # Named apart from the vendor shortfall, because the numbers alone mislead: a 

216 # machine with two agent CLIs and no `jury` reads as "2 of 2 available" while the 

217 # panel s7 would dispatch cannot run at all. 

218 why = ( 

219 f"the {JURY_RUNNER_COMMAND} runner s7 dispatches is not usable here; {counted}" 

220 if not self.runner.usable 

221 else counted 

222 ) 

223 return f"jury panel not staffable: {why}; unavailable: {listed or 'none probed'}" 

224 

225 def as_dict(self) -> dict[str, Any]: 

226 """JSON-stable record — the shape the assignment and the contract publish.""" 

227 return { 

228 "probed": True, 

229 "staffable": self.staffable, 

230 "decision": self.decision, 

231 "on_unavailable": self.policy, 

232 "required_vendors": self.required_vendors, 

233 "available_vendors": list(self.available_vendors), 

234 "unavailable": [seat.as_dict() for seat in self.unavailable], 

235 "runner": self.runner.as_dict(), 

236 "inventory": self.inventory, 

237 # This record was measured *here*. A verification surface republishing what the 

238 # ship measured says so instead (:data:`SOURCE_PULL_REQUEST` / 

239 # :data:`SOURCE_RUN_LEDGER`), because "we checked" and "we were told" are not 

240 # the same claim. 

241 "source": SOURCE_PROBE, 

242 "reason": self.reason, 

243 } 

244 

245 

246def assess( 

247 report: Mapping[str, Any] | None, 

248 *, 

249 runner: Runner = RUNNER_UNPROBED, 

250 min_vendors: int = DEFAULT_MIN_VENDORS, 

251 policy: str | None = None, 

252) -> Availability: 

253 """Read the panel runner — and, failing that, keel's provider report — into a verdict. 

254 

255 ``runner`` is :func:`keel.providerprobe.probe_jury_runner`'s answer about the ``jury`` 

256 binary s7 actually dispatches. It gates the whole verdict: no runner, no panel, whatever 

257 keel's delegate inventory says. It defaults to :data:`RUNNER_UNPROBED` — unusable — so a 

258 caller that forgets to probe it gets the conservative answer rather than a staffable 

259 panel nobody checked. 

260 

261 The vendor inventory is the runner's own when ``jury --doctor --json`` reported one: 

262 ai-jury is the authority on the panel it would convene, and keel's delegate list is only 

263 a proxy for it. ``report`` — :func:`keel.providerprobe.build_report`'s document — is the 

264 fallback for a runner that identified itself and named no agents — never for one that 

265 produced no document at all, which is not established to be ai-jury and reaches here 

266 as ``usable=False``. Either way a panel spans *vendors*, so two rows 

267 that shell out to the same CLI are one opinion (the rule 

268 :func:`keel.providers.distinct_vendors` states), and a hosted API with its key set is as 

269 real a panelist as a CLI on ``PATH``. 

270 

271 Total by construction. A missing or malformed inventory yields *no* available vendors and 

272 *no* named seats, which reads as "not staffable" — the conservative answer, and the 

273 one that then goes through the project's own configured allowance rather than being 

274 quietly decided here. 

275 """ 

276 rows, inventory = _inventory(runner, report) 

277 available: list[str] = [] 

278 unavailable: list[Seat] = [] 

279 if not runner.usable: 

280 # First in the list, because it is the first thing to fix: a reader who sees 

281 # `codex not found on PATH` and installs codex has not made the panel runnable. 

282 unavailable.append(Seat(JURY_RUNNER_COMMAND, JURY_RUNNER_VENDOR, runner.reason)) 

283 for row in rows: 

284 if not isinstance(row, Mapping): 

285 continue 

286 name = _text(row.get("name")) or _text(row.get("vendor")) or "(unnamed provider)" 

287 vendor = _text(row.get("vendor")) or name 

288 if row.get("available"): 

289 if vendor not in available: 

290 available.append(vendor) 

291 else: 

292 unavailable.append(Seat(name, vendor, _text(row.get("reason")) or "no reason reported")) 

293 return Availability( 

294 required_vendors=max(1, min_vendors), 

295 available_vendors=tuple(available), 

296 unavailable=tuple(unavailable), 

297 policy=jury_on_unavailable(policy), 

298 runner=runner, 

299 inventory=inventory, 

300 ) 

301 

302 

303def _inventory(runner: Runner, report: Mapping[str, Any] | None) -> tuple[tuple[Any, ...], str]: 

304 """``(rows, where they came from)`` — the runner's own panel, or keel's providers.""" 

305 rows = runner.panel_rows 

306 if rows is not None: 

307 return rows, INVENTORY_RUNNER 

308 return _rows(report, "providers") or (), INVENTORY_PROVIDERS 

309 

310 

311def _rows(document: Mapping[str, Any] | None, key: str) -> tuple[Any, ...] | None: 

312 """``document[key]`` as a tuple of rows, or ``None`` when it is not a list of them.""" 

313 rows = document.get(key) if isinstance(document, Mapping) else None 

314 if not isinstance(rows, Sequence) or isinstance(rows, (str, bytes)): 

315 return None 

316 return tuple(rows) 

317 

318 

319#: Where a published availability record came from. A *probe* measured this machine; the 

320#: other two are a verification surface reading what the ship measured, which is not the 

321#: same claim and must not be published as if it were. 

322SOURCE_PROBE = "probe" 

323SOURCE_PULL_REQUEST = "pull-request" 

324SOURCE_RUN_LEDGER = "run-ledger" 

325#: The run's own record, read back off the closure comment it posted — the same statement 

326#: as :data:`SOURCE_RUN_LEDGER`, from the copy that travels with the pull request (#1068). 

327SOURCE_CLOSURE_COMMENT = "closure-comment" 

328 

329 

330def panel_sat( 

331 *, min_vendors: int = DEFAULT_MIN_VENDORS, policy: str | None = None 

332) -> dict[str, Any]: 

333 """The panel demonstrably sat: a head-pinned jury verdict is on the pull request (#1066). 

334 

335 A verification surface must not answer "was the panel available" by asking *its own* 

336 machine. ``keel evidence-verify`` and ``keel merge`` run wherever CI puts them, and a 

337 change juried on a workstation and checked on a bare runner would otherwise have its 

338 required evidence quietly rewritten — the panel item dropped, three host verdicts 

339 demanded that nobody was ever asked to post. The ballots are already on the pull 

340 request; that outranks anything *this host* can observe. 

341 

342 It is the **weakest** pin, and :func:`pin` — not this function — owns that order. It 

343 does not outrank the shipping run's own record of what it did, in either of the two 

344 places that record survives: the ``ship_run`` ledger (:func:`shipped`) and the closure 

345 comment rendered from it (:func:`recorded`). A posted verdict establishes that a panel 

346 sat for this head; it does not establish that this run's review *was* that panel. 

347 

348 ``probed: False`` says plainly that nothing was measured here. Everything else keeps the 

349 shape :meth:`Availability.as_dict` publishes, so every reader downstream is unchanged. 

350 """ 

351 return { 

352 "probed": False, 

353 "staffable": True, 

354 "decision": DECISION_AVAILABLE, 

355 "on_unavailable": jury_on_unavailable(policy), 

356 "required_vendors": max(1, min_vendors), 

357 "available_vendors": [], 

358 "unavailable": [], 

359 "runner": {"command": JURY_RUNNER_COMMAND, "usable": True, "reason": "the panel sat"}, 

360 "inventory": SOURCE_PULL_REQUEST, 

361 "source": SOURCE_PULL_REQUEST, 

362 "reason": ( 

363 "jury panel staffable: a head-pinned jury verdict is posted on the pull " 

364 "request, so the panel sat for this head; this surface did not re-probe" 

365 ), 

366 } 

367 

368 

369def recorded( 

370 decision: Any, *, min_vendors: int = DEFAULT_MIN_VENDORS, policy: str | None = None 

371) -> dict[str, Any] | None: 

372 """The panel decision this run published in its own closure comment (#1068 round 6). 

373 

374 The same statement :func:`shipped` reads, from the copy that travels with the pull 

375 request. It exists because the stronger copy does not travel: the ``ship_run`` ledger 

376 lives under the gitignored ``.keel/state/``, so on a hosted ``evidence-verify`` or 

377 ``merge`` — the CI check, or any machine other than the one that shipped — there is no 

378 same-head record and the ledger pin cannot fire at all. The precedence it establishes 

379 held on the workstation that shipped and nowhere else, while a leftover 

380 ``keel.jury-verdict.v1`` answered for that run everywhere else. 

381 

382 ``decision`` comes from :func:`keel.evidence.shipped_panel_decision`, which has already 

383 held it to a trusted author, an actual closure comment, and this exact head. Only 

384 ``available`` and ``fallback`` produce a record, exactly as in :func:`shipped`: 

385 ``block`` refused its run, so it is not a decision anything shipped under, and an 

386 unrecognised value is not a decision at all. Either reads as ``None``, and :func:`pin` 

387 returns that ``None`` as the answer rather than falling through — the same rule it 

388 holds a same-head ledger record to, for the same reason. 

389 

390 The record is thinner than the ledger's — the comment carries the decision and the 

391 seats' prose, not the structured inventory — so it publishes no vendors and no seats 

392 and says where it came from. Every consumer reads ``decision`` 

393 (:func:`keel.team._panel_falls_back`) and the ``reason`` sentence, both of which are 

394 here; nothing downstream needs the seat list to resolve a bench. 

395 

396 ``probed: False``, like every pin: a surface that read a comment measured nothing. 

397 """ 

398 if decision not in (DECISION_AVAILABLE, DECISION_FALLBACK): 

399 return None 

400 staffable = decision == DECISION_AVAILABLE 

401 outcome = ( 

402 "the panel sat" 

403 if staffable 

404 else "the panel could not be staffed there and a host bench reviewed instead" 

405 ) 

406 return { 

407 "probed": False, 

408 "staffable": staffable, 

409 "decision": decision, 

410 "on_unavailable": jury_on_unavailable(policy), 

411 "required_vendors": max(1, min_vendors), 

412 "available_vendors": [], 

413 "unavailable": [], 

414 "runner": { 

415 "command": JURY_RUNNER_COMMAND, 

416 "usable": staffable, 

417 "reason": outcome, 

418 }, 

419 "inventory": SOURCE_CLOSURE_COMMENT, 

420 "source": SOURCE_CLOSURE_COMMENT, 

421 "reason": ( 

422 f"jury panel {'staffable' if staffable else 'not staffable'}: the run that " 

423 f"produced this head recorded '{decision}' in the closure comment it posted " 

424 "on this pull request, so " + outcome + "; this surface did not re-probe" 

425 ), 

426 } 

427 

428 

429def shipped(record: Mapping[str, Any] | None, *, head_sha: str | None) -> dict[str, Any] | None: 

430 """The panel decision the ship that produced **this head** measured, or ``None``. 

431 

432 Read out of that run's ``ship_run`` ledger entry at ``run_context.jury_panel``, which 

433 :func:`keel.ledger.build_ship_run_record` writes for exactly this purpose. Total: a 

434 record from before the field existed, or one whose run resolved no panel, reads as 

435 ``None``. 

436 

437 What that ``None`` then means is :func:`pin`'s to decide and not this function's, and 

438 the two cases part there rather than here: a run that *was* asked and answered ``null`` 

439 silences the lower sources, while a record that never carried the key 

440 (:func:`states_panel`) is not consulted at all and the lower sources get their turn. 

441 

442 **Pinned to the exact head, the way the posted-verdict path is.** The record is selected 

443 by pull-request number (:func:`keel.ledger.latest_ship_run_for_pr`), and a pull request 

444 outlives its heads: a ship of an earlier head that fell back to a host bench would 

445 otherwise weaken the contract of the head being verified now, which is a stale run 

446 relaxing a live gate. So the record's ``git.head_sha`` must equal the head under 

447 verification, and anything else — an older head, a blank or absent head on either side, 

448 a malformed ``git`` block — reads as ``None``. 

449 

450 ``None`` **fails closed**, which is why it is safe to be strict here. It does not waive 

451 the panel; it drops the pin, and the caller then measures this machine. Both ways that 

452 can land are the refusing one: a fallback-shipped change verified where the panel *can* 

453 be staffed is held to a panel it did not run, and a panel-shipped change verified on a 

454 bare runner is held to host verdicts nobody posted. A run that genuinely convened the 

455 panel at this head is unaffected either way — this record then says ``available`` and 

456 its ballots are on the pull request, so both sources agree. 

457 

458 ``probed: False`` for the same reason :func:`panel_sat` sets it: this is a record being 

459 republished, not a measurement taken here. The ledger's own copy carries ``probed: True`` 

460 because the *ship* did probe; repeating that claim on a surface that only read a file 

461 would be the one thing this module refuses to let a record do — claim a provenance it 

462 does not have. 

463 """ 

464 context = record.get("run_context") if isinstance(record, Mapping) else None 

465 panel = context.get("jury_panel") if isinstance(context, Mapping) else None 

466 if not isinstance(panel, Mapping) or panel.get("decision") not in ( 

467 DECISION_AVAILABLE, 

468 DECISION_FALLBACK, 

469 ): 

470 return None 

471 if not _matches_head(record, head_sha): 

472 return None 

473 return {**dict(panel), "probed": False, "source": SOURCE_RUN_LEDGER} 

474 

475 

476def is_ship_run_for_head(record: Mapping[str, Any] | None, *, head_sha: Any) -> bool: 

477 """Did a ``ship_run`` for **this exact head** leave a record? (#1068) 

478 

479 Presence, not content: the run's ``run_context.jury_panel`` may say ``fallback``, 

480 ``block``, or ``None``, and this still answers ``True``. That separation is the 

481 whole point — :func:`pin` needs to know *whether the run left a record here* before it 

482 reads what the record says, because a record that says nothing about a panel is still 

483 a run that did not ship under one. 

484 """ 

485 return isinstance(record, Mapping) and _matches_head(record, head_sha) 

486 

487 

488def states_panel(record: Mapping[str, Any] | None) -> bool: 

489 """Does this record carry the ``jury_panel`` **key** at all? (#1068 round 7) 

490 

491 Not what it says — whether the run that wrote it had the word. The distinction is 

492 between a run that was *asked* about the panel and answered (even by answering 

493 ``None``: "this tier named no panel") and a record written before the field existed, 

494 which was never asked. 

495 

496 :func:`keel.ledger._run_context` always writes the key, ``None`` included, so for every 

497 record this feature produces the answer is ``True`` and rank 3 of :func:`pin` is 

498 unchanged: a same-head record whose ``jury_panel`` is ``null`` is the run saying it did 

499 not ship under a panel, and it silences the lower sources. A ledger row from before 

500 #1066 has no such key — missing vocabulary, not a statement — and silencing on its 

501 behalf would put words in a run's mouth: on a workstation carrying one, a change the 

502 panel really did jury would have had its posted ballots ignored and 

503 ``review-verdict-1..3`` demanded by a probe of the local machine. Those rows fall 

504 through to the closure comment and then to the ballots, which is exactly what they did 

505 before #1066 existed. 

506 

507 A record whose ``run_context`` is missing or unreadable has no key either, and reads the 

508 same way, for the same reason: absence of vocabulary, not a statement. 

509 """ 

510 context = record.get("run_context") if isinstance(record, Mapping) else None 

511 return isinstance(context, Mapping) and "jury_panel" in context 

512 

513 

514def pin( 

515 record: Mapping[str, Any] | None, 

516 *, 

517 head_sha: Any, 

518 panel_verdict_posted: bool, 

519 closure_panel_decision: Any = None, 

520) -> dict[str, Any] | None: 

521 """**The single authority on "what did this run ship under".** (#1066, #1068) 

522 

523 Every verification surface — ``keel evidence-verify``, ``keel merge``, and 

524 :func:`keel.cli._shipped_jury_availability`, which is only this function's thin-I/O 

525 wrapper — resolves that question here and nowhere else. It is one function rather than 

526 an order of ``if``-statements at a call site because the *precedence* between the two 

527 pins is itself a rule, and #1068 rounds 2–4 each found a rule written in one place and 

528 forgotten in its twin. There is one place now, and this docstring is it. 

529 

530 ``None`` means "nothing pins this head": the caller measures its own machine, exactly 

531 as it did before either pin existed. That is the fail-closed answer, never a waiver — 

532 a probe can only add the panel requirement back or demand the tier's host verdicts. 

533 

534 **Both run-record sources select the *latest* record for this head, and that direction 

535 is the rule** (#1068 round 7). The ledger source resolves through 

536 :func:`keel.ledger.latest_ship_run_for_pr`, which walks the chronologically appended 

537 records and keeps the **last** match; the closure source resolves through 

538 :func:`keel.evidence.shipped_panel_decision`, which walks ``pr_comments`` in GitHub's 

539 oldest-first order and keeps the **last** match. They are two copies of one statement — 

540 ranks 2 and 4 are the same sentence read off two artifacts — so if they disagreed about 

541 which run they were quoting, the precedence between them would be meaningless: the 

542 machine with the ledger would answer for the newest ship and the machine without it for 

543 the oldest. Round 6 had exactly that, and worse: the closure source was first-match 

544 *and* a run whose panel sat rendered no marker, so a commit shipped once under the 

545 fallback and then again on a machine that could staff the panel left one marker on the 

546 pull request saying ``fallback``, and CI — where there is no ledger — pinned the 

547 host-bench contract onto a panel-reviewed change and never asked for the panel's own 

548 verdict. :func:`keel.closure._jury_panel` now emits ``decision=available`` too, so every 

549 ship speaks and "latest wins" is well defined on both sides. 

550 ``tests/test_juryavail.py::TestBothRunRecordSourcesSelectTheLatest`` holds them to it. 

551 

552 The order, and why it is this way round: 

553 

554 1. **No head, no pin.** :func:`is_pinnable_head`. A pin removes requirements — it takes 

555 ``review-verdict-1..3`` off the required set outright — so it may only ever be taken 

556 against an exact commit. ``panel_verdict_posted`` is ignored here even when ``True``, 

557 because :func:`keel.evidence.panel_verdict_posted` reads a blank head as *no head 

558 filter*: right for counting evidence, wrong for a pin. 

559 2. **The run's own ledger record for this head wins** (:func:`is_ship_run_for_head`, 

560 then :func:`shipped`). The ledger records what *this run actually did*; a posted 

561 verdict records what somebody put on the pull request. At the same head the two can 

562 disagree, and then the ledger is the one that is evidence of the run: a ship that 

563 measured ``fallback`` seated three host reviewers and owes ``review-verdict-1..3``, 

564 and a leftover jury verdict at that head — from an earlier ship of the same commit, 

565 from a force-push back onto it, or from a collaborator who ran ``jury`` by hand — 

566 is not that run's review. Letting the verdict win dropped three required items on 

567 the strength of a comment nobody's run had promised. 

568 3. **A same-head record that says nothing about a panel still speaks — if it had the 

569 word.** :func:`shipped` returns ``None`` for a record whose ``run_context.jury_panel`` 

570 is ``null``, malformed, or ``block``, and that ``None`` is returned as-is rather than 

571 falling through to a comment: this run left a record here and it does not say the run 

572 shipped under a panel, so nobody may say otherwise on its behalf. 

573 

574 The one record that does *not* speak is one whose ``run_context`` never carried the 

575 key (:func:`states_panel`) — a ledger row written before #1066, or one whose 

576 ``run_context`` is unreadable. That is missing vocabulary, not a statement, and 

577 silencing the lower sources on its behalf would answer a question the run was never 

578 asked: a change the panel really did jury, verified on the workstation that still 

579 has that row, would have had its posted ballots ignored and ``review-verdict-1..3`` 

580 demanded by a probe of *this* machine. Such a record falls through to rank 4 and then 

581 rank 5, which is what it did before #1066 existed. Every record this feature writes 

582 carries the key — :func:`keel.ledger._run_context` always writes it, ``None`` 

583 included — so the fall-through applies to legacy rows and to nothing else. 

584 4. **Failing that, the run's own closure comment** (:func:`recorded`), whose decision 

585 :func:`keel.evidence.shipped_panel_decision` has already held to a trusted author, 

586 to an actual closure comment, and to this head. This rank is what makes rank 2 mean 

587 anything off the shipping workstation (#1068 round 6): the ledger lives under the 

588 gitignored ``.keel/state/``, so a hosted ``evidence-verify`` or ``merge`` has no 

589 same-head record at all and fell straight through to the verdict — the run's 

590 fallback was outranked by a leftover comment on every machine except the one that 

591 had no need of the rule. The closure comment is the *same statement* as the ledger 

592 record it was rendered from, in the one place that travels with the pull request, 

593 so it ranks with the ledger and above the verdict. Since round 7 it is silent only 

594 for a head no keel closure comment names — a run that shipped under a staffable panel 

595 records ``decision=available`` and pins the panel here rather than leaving rank 5 to 

596 infer it from ballots. 

597 

598 Silent and refusing are different answers, and rank 4 gives both: no marker for this 

599 head is silence and rank 5 gets its turn, while a marker that says ``block`` or 

600 something unrecognised is a record that does not say the run shipped under a panel, 

601 and :func:`recorded` returns ``None`` as the answer for the same reason rank 3 does. 

602 5. **Only with no record of the run's own does a posted verdict pin** 

603 (:func:`panel_sat`). Head-pinned ballots prove a panel *sat* for this head; they do 

604 not prove this run's review **was** that panel, which is why they rank last. 

605 

606 Every pin therefore runs in the same direction: the strongest available statement about 

607 *this run at this head*, falling back to measuring rather than to guessing. 

608 """ 

609 if not is_pinnable_head(head_sha): 

610 return None 

611 if is_ship_run_for_head(record, head_sha=head_sha) and states_panel(record): 

612 return shipped(record, head_sha=head_sha) 

613 if closure_panel_decision is not None: 

614 return recorded(closure_panel_decision) 

615 return panel_sat() if panel_verdict_posted else None 

616 

617 

618def is_pinnable_head(head_sha: Any) -> bool: 

619 """Is ``head_sha`` a head a pin may be taken against? (#1068) 

620 

621 **The one blank-head rule both panel pins read.** A pin republishes an earlier run's 

622 panel decision in place of measuring this machine, so it may only ever be taken against 

623 an exact commit: an unknown head must not be authorized by a record — or a comment — 

624 from some other one. :func:`keel.ledger.gates_pass_for_head` already holds the merge 

625 gate to this, and :func:`shipped` to the ledger pin; round 3 hardened those and left 

626 the posted-verdict pin reading :func:`keel.evidence._matches_head`, which treats a 

627 blank head as *unfiltered* and so counted any trusted jury marker on the pull request 

628 as this head's. Written twice, hardened once. It is written here now, and :func:`pin` 

629 — the one place the two sources are ranked — asks it before either is consulted. 

630 

631 :func:`keel.evidence._matches_head` is deliberately the other rule and stays that way: 

632 it filters *evidence items* inside a gate that, with no head resolved, runs 

633 head-agnostic throughout — every review verdict counts too. Nothing reached through it 

634 removes a requirement. A pin does: it takes ``review-verdict-1..3`` off the required 

635 set entirely. 

636 

637 **The test is what the answer can do, not which module asks.** So this predicate has a 

638 third caller that is not a pin: :func:`keel.evidence.jury_participating_vendors`, whose 

639 count downgrades a gating jury to advisory below ``jury.min_vendors`` and thereby drops 

640 ``jury-verdict`` from the required evidence (#1069). It removes a requirement, so it 

641 takes this rule; its two panel-shaped siblings there cannot, so they keep the 

642 permissive one. 

643 """ 

644 return isinstance(head_sha, str) and bool(head_sha.strip()) 

645 

646 

647def _matches_head(record: Mapping[str, Any], head_sha: str | None) -> bool: 

648 """Was this ledger record written for exactly ``head_sha``? A blank head never matches.""" 

649 if not is_pinnable_head(head_sha): 

650 return False 

651 git = record.get("git") 

652 recorded = git.get("head_sha") if isinstance(git, Mapping) else None 

653 return isinstance(recorded, str) and recorded == head_sha 

654 

655 

656def _text(value: Any) -> str | None: 

657 """A non-blank string, or ``None`` — so a blank field reads as unset.""" 

658 return value.strip() if isinstance(value, str) and value.strip() else None