Coverage for src/keel/juryavail.py: 100%
126 statements
« prev ^ index » next coverage.py v7.16.2, created at 2026-10-02 20:26 +0000
« prev ^ index » next coverage.py v7.16.2, created at 2026-10-02 20:26 +0000
1"""Can the cross-vendor panel actually be staffed here? (#1066)
3On a tier whose ``knobs.team.review.by_tier`` names ``jury``, s7 dispatches the panel and
4its ballots *are* the review — #1014 round 3 deliberately made it so no operator flag can
5take the panel back off. That is right while the panel can run. When it cannot — an agent
6CLI is not installed, is unauthenticated, or the account is out of quota — the tier has no
7way forward at all: the only review it has is one this machine cannot convene.
9This module is the pure half of the answer. The question it has to answer is narrower than
10"are some agent CLIs installed": s7 does not convene a panel out of keel's delegate
11inventory, it runs the **``jury`` binary** (``src/keel/adapters/commands/ship.md``), and
12that binary holds its own configured panel. So the probe asks the runner first —
13``jury --doctor --json``, ai-jury's own readiness document, which reports both that the
14binary is there and which of *its* agents are usable — and only falls back to
15:func:`keel.providerprobe.collect` (what ``keel doctor --providers`` prints) for a runner
16whose document named no agents. A machine with ``claude`` and ``codex`` on ``PATH`` and no
17``jury`` is **not** staffable, however healthy keel's own inventory looks: the panel s7
18would dispatch cannot run.
20That document is also the binary's *identity*, and no document is no identity (#1068):
21a ``jury`` on ``PATH`` that exits 0 without one has not established that it is ai-jury, so
22it is unusable and keel's inventory cannot make it staffable. The proxy stands in for a
23panel ai-jury declined to enumerate, never for a panel runner nobody established is there.
25Two answers to one question is what this ordering avoids. keel's delegate inventory is a
26proxy for the panel, ai-jury's is the panel; when the panel can speak for itself it is the
27authority, and the record says which inventory the verdict was read from.
29Three things the design holds to, all of them from the issue:
31* **Availability is measured, never asserted.** There is no flag that says "the panel is
32 fine". What may not take the panel off is an operator's *preference*; availability is a
33 fact about the world, and it is allowed to change the outcome precisely because it is
34 recorded.
35* **The policy is a configured allowance, not an automatic behaviour.**
36 ``knobs.team.jury.on_unavailable`` is ``fallback`` (the sensible default for a solo
37 project) or ``block`` (today's strictness, for a project whose product claim *is*
38 cross-vendor review).
39* **Never a silent downgrade.** ai-jury #682 exists because a panel that quietly collapsed
40 to one vendor still reported success. So :meth:`Availability.as_dict` carries which
41 seats were unavailable and why, all the way into the published assignment, the review
42 contract, the run ledger and the closure comment. The fallback changes *who sat*, never
43 *how many*: the seat count and the evidence requirement are the tier's, not the panel's.
45Pure and deterministic: the report goes in, the verdict comes out, and nothing here
46touches PATH, a subprocess, or the clock.
47"""
49from __future__ import annotations
51from collections.abc import Mapping, Sequence
52from dataclasses import dataclass
53from typing import Any
55from .team import (
56 DEFAULT_MIN_VENDORS,
57 JURY_ON_UNAVAILABLE_DEFAULT,
58 jury_on_unavailable,
59)
60from .team import JURY_RUNNER_COMMAND as JURY_RUNNER_COMMAND
61from .team import JuryUnavailableError as JuryUnavailableError
62from .team import refusal_message as refusal_message
64#: The module's public surface, in definition order (#1070). It is declared because the
65#: ``X as X`` re-exports above are read from *other* modules — a use CodeQL's
66#: ``py/unused-import`` cannot see, since it counts same-module uses only. A name listed
67#: in ``__all__`` is used by definition, so the declaration answers the scanner with the
68#: language's own statement of intent rather than with a dismissal. Being a real
69#: declaration it has to be the *whole* surface, not the re-exports alone;
70#: ``tests/test_reexport_surface.py`` holds it to that in both directions.
71__all__ = [
72 "JURY_RUNNER_COMMAND",
73 "JuryUnavailableError",
74 "refusal_message",
75 "JURY_RUNNER_VENDOR",
76 "INVENTORY_RUNNER",
77 "INVENTORY_PROVIDERS",
78 "DECISION_AVAILABLE",
79 "DECISION_FALLBACK",
80 "DECISION_BLOCK",
81 "Seat",
82 "Runner",
83 "RUNNER_UNPROBED",
84 "Availability",
85 "assess",
86 "SOURCE_PROBE",
87 "SOURCE_PULL_REQUEST",
88 "SOURCE_RUN_LEDGER",
89 "SOURCE_CLOSURE_COMMENT",
90 "panel_sat",
91 "recorded",
92 "shipped",
93 "is_ship_run_for_head",
94 "states_panel",
95 "pin",
96 "is_pinnable_head",
97]
99#: Re-exported from :mod:`keel.team`, which owns them because :func:`_review_seats` — the
100#: one place a bench is resolved, and so the one place a blocked panel can refuse the work
101#: it is actually about to review — cannot import this module: the import runs the other
102#: way. They keep their names here because this is the module the feature is named for and
103#: where a reader looks for them; ``keel.juryavail.JuryUnavailableError`` and
104#: ``keel.team.JuryUnavailableError`` are one class, not two.
105#:
106#: ``JURY_RUNNER_COMMAND`` is the binary a jury-panel tier's s7 actually dispatches. Not a
107#: delegate keel runs itself: keel does not depend on ai-jury, and every path through this
108#: module stays total when it is absent — absent simply means the panel cannot sit here.
110#: The vendor the runner seat is attributed to, so a reader of ``unavailable`` can tell the
111#: missing *panel* apart from a missing *panelist*.
112JURY_RUNNER_VENDOR = "ai-jury"
114#: Where a verdict's vendor inventory was read from — recorded, because the two sources do
115#: not have to agree and a reader must not have to guess which one spoke.
116INVENTORY_RUNNER = f"{JURY_RUNNER_COMMAND} --doctor"
117INVENTORY_PROVIDERS = "keel doctor --providers"
119#: The panel is staffable — nothing changes, the ballots are the review.
120DECISION_AVAILABLE = "available"
121#: The panel is not staffable and the policy allows a host bench in its place.
122DECISION_FALLBACK = "fallback"
123#: The panel is not staffable and the policy refuses the run.
124DECISION_BLOCK = "block"
127@dataclass(frozen=True)
128class Seat:
129 """One provider the panel could have used, and why it cannot."""
131 provider: str
132 vendor: str
133 reason: str
135 def as_dict(self) -> dict[str, Any]:
136 return {"provider": self.provider, "vendor": self.vendor, "reason": self.reason}
139@dataclass(frozen=True)
140class Runner:
141 """The ``jury`` CLI itself: can s7 dispatch it here, and what panel does it hold?
143 Produced by :func:`keel.providerprobe.probe_jury_runner` — the thin-I/O half — and
144 consumed here. ``doctor`` is ai-jury's own ``jury --doctor --json`` document, which is
145 both the panel it holds and the way the binary identifies itself: without a readable
146 one the probe reports ``usable=False`` (#1068). A document that identifies ai-jury but
147 names no ``agents`` still means *usable*, and the verdict then reads keel's own
148 delegate inventory for the vendor count instead.
149 """
151 usable: bool
152 reason: str
153 doctor: Mapping[str, Any] | None = None
155 @property
156 def panel_rows(self) -> tuple[Any, ...] | None:
157 """The agents ai-jury reports for its own panel, or ``None`` when it reported none."""
158 return _rows(self.doctor, "agents")
160 def as_dict(self) -> dict[str, Any]:
161 return {"command": JURY_RUNNER_COMMAND, "usable": self.usable, "reason": self.reason}
164#: What an unprobed runner reads as. Fail-closed on purpose: the whole point of #1066
165#: round 2 is that a panel nobody established could run must not be reported staffable, so
166#: "we did not ask" and "we asked and it is fine" cannot share an answer.
167RUNNER_UNPROBED = Runner(False, f"the {JURY_RUNNER_COMMAND} runner was not probed")
170@dataclass(frozen=True)
171class Availability:
172 """The probe's verdict on this tier's panel, ready to publish."""
174 #: Distinct vendors the panel needs before it can be a *cross-vendor* panel. This is
175 #: ``team.jury.min_vendors``, the same floor the verdict is later held to.
176 required_vendors: int
177 #: Vendors the probe found usable here, in the probe's own (deterministic) order.
178 available_vendors: tuple[str, ...]
179 #: Every seat the panel could not use, with the reason it reported. The ``jury`` runner
180 #: itself is one of them when it is the thing that is missing.
181 unavailable: tuple[Seat, ...]
182 #: ``fallback`` | ``block`` — the configured allowance, already defaulted.
183 policy: str = JURY_ON_UNAVAILABLE_DEFAULT
184 #: The ``jury`` binary s7 dispatches. Unprobed reads as unusable, never as fine.
185 runner: Runner = RUNNER_UNPROBED
186 #: Which inventory the vendor counts came from — the runner's own, or keel's.
187 inventory: str = INVENTORY_PROVIDERS
189 @property
190 def staffable(self) -> bool:
191 """Can this machine convene a panel spanning ``required_vendors`` vendors?
193 Both halves, because s7 needs both: the runner that dispatches the panel, and
194 enough distinct vendors for it to *be* a cross-vendor panel. Agent CLIs on ``PATH``
195 with no ``jury`` to convene them is not a panel, it is an inventory.
196 """
197 return self.runner.usable and len(self.available_vendors) >= self.required_vendors
199 @property
200 def decision(self) -> str:
201 """:data:`DECISION_AVAILABLE`, or the policy when the panel cannot be staffed."""
202 return DECISION_AVAILABLE if self.staffable else self.policy
204 @property
205 def reason(self) -> str:
206 """One sentence a reader can act on, naming the seats that were unavailable."""
207 counted = (
208 f"{len(self.available_vendors)} vendor(s) available "
209 f"({', '.join(self.available_vendors) or 'none'}), {self.required_vendors} "
210 f"required (per {self.inventory})"
211 )
212 if self.staffable:
213 return f"jury panel staffable: {counted}, dispatched by {self.runner.reason}"
214 listed = ", ".join(f"{seat.provider} ({seat.reason})" for seat in self.unavailable)
215 # Named apart from the vendor shortfall, because the numbers alone mislead: a
216 # machine with two agent CLIs and no `jury` reads as "2 of 2 available" while the
217 # panel s7 would dispatch cannot run at all.
218 why = (
219 f"the {JURY_RUNNER_COMMAND} runner s7 dispatches is not usable here; {counted}"
220 if not self.runner.usable
221 else counted
222 )
223 return f"jury panel not staffable: {why}; unavailable: {listed or 'none probed'}"
225 def as_dict(self) -> dict[str, Any]:
226 """JSON-stable record — the shape the assignment and the contract publish."""
227 return {
228 "probed": True,
229 "staffable": self.staffable,
230 "decision": self.decision,
231 "on_unavailable": self.policy,
232 "required_vendors": self.required_vendors,
233 "available_vendors": list(self.available_vendors),
234 "unavailable": [seat.as_dict() for seat in self.unavailable],
235 "runner": self.runner.as_dict(),
236 "inventory": self.inventory,
237 # This record was measured *here*. A verification surface republishing what the
238 # ship measured says so instead (:data:`SOURCE_PULL_REQUEST` /
239 # :data:`SOURCE_RUN_LEDGER`), because "we checked" and "we were told" are not
240 # the same claim.
241 "source": SOURCE_PROBE,
242 "reason": self.reason,
243 }
246def assess(
247 report: Mapping[str, Any] | None,
248 *,
249 runner: Runner = RUNNER_UNPROBED,
250 min_vendors: int = DEFAULT_MIN_VENDORS,
251 policy: str | None = None,
252) -> Availability:
253 """Read the panel runner — and, failing that, keel's provider report — into a verdict.
255 ``runner`` is :func:`keel.providerprobe.probe_jury_runner`'s answer about the ``jury``
256 binary s7 actually dispatches. It gates the whole verdict: no runner, no panel, whatever
257 keel's delegate inventory says. It defaults to :data:`RUNNER_UNPROBED` — unusable — so a
258 caller that forgets to probe it gets the conservative answer rather than a staffable
259 panel nobody checked.
261 The vendor inventory is the runner's own when ``jury --doctor --json`` reported one:
262 ai-jury is the authority on the panel it would convene, and keel's delegate list is only
263 a proxy for it. ``report`` — :func:`keel.providerprobe.build_report`'s document — is the
264 fallback for a runner that identified itself and named no agents — never for one that
265 produced no document at all, which is not established to be ai-jury and reaches here
266 as ``usable=False``. Either way a panel spans *vendors*, so two rows
267 that shell out to the same CLI are one opinion (the rule
268 :func:`keel.providers.distinct_vendors` states), and a hosted API with its key set is as
269 real a panelist as a CLI on ``PATH``.
271 Total by construction. A missing or malformed inventory yields *no* available vendors and
272 *no* named seats, which reads as "not staffable" — the conservative answer, and the
273 one that then goes through the project's own configured allowance rather than being
274 quietly decided here.
275 """
276 rows, inventory = _inventory(runner, report)
277 available: list[str] = []
278 unavailable: list[Seat] = []
279 if not runner.usable:
280 # First in the list, because it is the first thing to fix: a reader who sees
281 # `codex not found on PATH` and installs codex has not made the panel runnable.
282 unavailable.append(Seat(JURY_RUNNER_COMMAND, JURY_RUNNER_VENDOR, runner.reason))
283 for row in rows:
284 if not isinstance(row, Mapping):
285 continue
286 name = _text(row.get("name")) or _text(row.get("vendor")) or "(unnamed provider)"
287 vendor = _text(row.get("vendor")) or name
288 if row.get("available"):
289 if vendor not in available:
290 available.append(vendor)
291 else:
292 unavailable.append(Seat(name, vendor, _text(row.get("reason")) or "no reason reported"))
293 return Availability(
294 required_vendors=max(1, min_vendors),
295 available_vendors=tuple(available),
296 unavailable=tuple(unavailable),
297 policy=jury_on_unavailable(policy),
298 runner=runner,
299 inventory=inventory,
300 )
303def _inventory(runner: Runner, report: Mapping[str, Any] | None) -> tuple[tuple[Any, ...], str]:
304 """``(rows, where they came from)`` — the runner's own panel, or keel's providers."""
305 rows = runner.panel_rows
306 if rows is not None:
307 return rows, INVENTORY_RUNNER
308 return _rows(report, "providers") or (), INVENTORY_PROVIDERS
311def _rows(document: Mapping[str, Any] | None, key: str) -> tuple[Any, ...] | None:
312 """``document[key]`` as a tuple of rows, or ``None`` when it is not a list of them."""
313 rows = document.get(key) if isinstance(document, Mapping) else None
314 if not isinstance(rows, Sequence) or isinstance(rows, (str, bytes)):
315 return None
316 return tuple(rows)
319#: Where a published availability record came from. A *probe* measured this machine; the
320#: other two are a verification surface reading what the ship measured, which is not the
321#: same claim and must not be published as if it were.
322SOURCE_PROBE = "probe"
323SOURCE_PULL_REQUEST = "pull-request"
324SOURCE_RUN_LEDGER = "run-ledger"
325#: The run's own record, read back off the closure comment it posted — the same statement
326#: as :data:`SOURCE_RUN_LEDGER`, from the copy that travels with the pull request (#1068).
327SOURCE_CLOSURE_COMMENT = "closure-comment"
330def panel_sat(
331 *, min_vendors: int = DEFAULT_MIN_VENDORS, policy: str | None = None
332) -> dict[str, Any]:
333 """The panel demonstrably sat: a head-pinned jury verdict is on the pull request (#1066).
335 A verification surface must not answer "was the panel available" by asking *its own*
336 machine. ``keel evidence-verify`` and ``keel merge`` run wherever CI puts them, and a
337 change juried on a workstation and checked on a bare runner would otherwise have its
338 required evidence quietly rewritten — the panel item dropped, three host verdicts
339 demanded that nobody was ever asked to post. The ballots are already on the pull
340 request; that outranks anything *this host* can observe.
342 It is the **weakest** pin, and :func:`pin` — not this function — owns that order. It
343 does not outrank the shipping run's own record of what it did, in either of the two
344 places that record survives: the ``ship_run`` ledger (:func:`shipped`) and the closure
345 comment rendered from it (:func:`recorded`). A posted verdict establishes that a panel
346 sat for this head; it does not establish that this run's review *was* that panel.
348 ``probed: False`` says plainly that nothing was measured here. Everything else keeps the
349 shape :meth:`Availability.as_dict` publishes, so every reader downstream is unchanged.
350 """
351 return {
352 "probed": False,
353 "staffable": True,
354 "decision": DECISION_AVAILABLE,
355 "on_unavailable": jury_on_unavailable(policy),
356 "required_vendors": max(1, min_vendors),
357 "available_vendors": [],
358 "unavailable": [],
359 "runner": {"command": JURY_RUNNER_COMMAND, "usable": True, "reason": "the panel sat"},
360 "inventory": SOURCE_PULL_REQUEST,
361 "source": SOURCE_PULL_REQUEST,
362 "reason": (
363 "jury panel staffable: a head-pinned jury verdict is posted on the pull "
364 "request, so the panel sat for this head; this surface did not re-probe"
365 ),
366 }
369def recorded(
370 decision: Any, *, min_vendors: int = DEFAULT_MIN_VENDORS, policy: str | None = None
371) -> dict[str, Any] | None:
372 """The panel decision this run published in its own closure comment (#1068 round 6).
374 The same statement :func:`shipped` reads, from the copy that travels with the pull
375 request. It exists because the stronger copy does not travel: the ``ship_run`` ledger
376 lives under the gitignored ``.keel/state/``, so on a hosted ``evidence-verify`` or
377 ``merge`` — the CI check, or any machine other than the one that shipped — there is no
378 same-head record and the ledger pin cannot fire at all. The precedence it establishes
379 held on the workstation that shipped and nowhere else, while a leftover
380 ``keel.jury-verdict.v1`` answered for that run everywhere else.
382 ``decision`` comes from :func:`keel.evidence.shipped_panel_decision`, which has already
383 held it to a trusted author, an actual closure comment, and this exact head. Only
384 ``available`` and ``fallback`` produce a record, exactly as in :func:`shipped`:
385 ``block`` refused its run, so it is not a decision anything shipped under, and an
386 unrecognised value is not a decision at all. Either reads as ``None``, and :func:`pin`
387 returns that ``None`` as the answer rather than falling through — the same rule it
388 holds a same-head ledger record to, for the same reason.
390 The record is thinner than the ledger's — the comment carries the decision and the
391 seats' prose, not the structured inventory — so it publishes no vendors and no seats
392 and says where it came from. Every consumer reads ``decision``
393 (:func:`keel.team._panel_falls_back`) and the ``reason`` sentence, both of which are
394 here; nothing downstream needs the seat list to resolve a bench.
396 ``probed: False``, like every pin: a surface that read a comment measured nothing.
397 """
398 if decision not in (DECISION_AVAILABLE, DECISION_FALLBACK):
399 return None
400 staffable = decision == DECISION_AVAILABLE
401 outcome = (
402 "the panel sat"
403 if staffable
404 else "the panel could not be staffed there and a host bench reviewed instead"
405 )
406 return {
407 "probed": False,
408 "staffable": staffable,
409 "decision": decision,
410 "on_unavailable": jury_on_unavailable(policy),
411 "required_vendors": max(1, min_vendors),
412 "available_vendors": [],
413 "unavailable": [],
414 "runner": {
415 "command": JURY_RUNNER_COMMAND,
416 "usable": staffable,
417 "reason": outcome,
418 },
419 "inventory": SOURCE_CLOSURE_COMMENT,
420 "source": SOURCE_CLOSURE_COMMENT,
421 "reason": (
422 f"jury panel {'staffable' if staffable else 'not staffable'}: the run that "
423 f"produced this head recorded '{decision}' in the closure comment it posted "
424 "on this pull request, so " + outcome + "; this surface did not re-probe"
425 ),
426 }
429def shipped(record: Mapping[str, Any] | None, *, head_sha: str | None) -> dict[str, Any] | None:
430 """The panel decision the ship that produced **this head** measured, or ``None``.
432 Read out of that run's ``ship_run`` ledger entry at ``run_context.jury_panel``, which
433 :func:`keel.ledger.build_ship_run_record` writes for exactly this purpose. Total: a
434 record from before the field existed, or one whose run resolved no panel, reads as
435 ``None``.
437 What that ``None`` then means is :func:`pin`'s to decide and not this function's, and
438 the two cases part there rather than here: a run that *was* asked and answered ``null``
439 silences the lower sources, while a record that never carried the key
440 (:func:`states_panel`) is not consulted at all and the lower sources get their turn.
442 **Pinned to the exact head, the way the posted-verdict path is.** The record is selected
443 by pull-request number (:func:`keel.ledger.latest_ship_run_for_pr`), and a pull request
444 outlives its heads: a ship of an earlier head that fell back to a host bench would
445 otherwise weaken the contract of the head being verified now, which is a stale run
446 relaxing a live gate. So the record's ``git.head_sha`` must equal the head under
447 verification, and anything else — an older head, a blank or absent head on either side,
448 a malformed ``git`` block — reads as ``None``.
450 ``None`` **fails closed**, which is why it is safe to be strict here. It does not waive
451 the panel; it drops the pin, and the caller then measures this machine. Both ways that
452 can land are the refusing one: a fallback-shipped change verified where the panel *can*
453 be staffed is held to a panel it did not run, and a panel-shipped change verified on a
454 bare runner is held to host verdicts nobody posted. A run that genuinely convened the
455 panel at this head is unaffected either way — this record then says ``available`` and
456 its ballots are on the pull request, so both sources agree.
458 ``probed: False`` for the same reason :func:`panel_sat` sets it: this is a record being
459 republished, not a measurement taken here. The ledger's own copy carries ``probed: True``
460 because the *ship* did probe; repeating that claim on a surface that only read a file
461 would be the one thing this module refuses to let a record do — claim a provenance it
462 does not have.
463 """
464 context = record.get("run_context") if isinstance(record, Mapping) else None
465 panel = context.get("jury_panel") if isinstance(context, Mapping) else None
466 if not isinstance(panel, Mapping) or panel.get("decision") not in (
467 DECISION_AVAILABLE,
468 DECISION_FALLBACK,
469 ):
470 return None
471 if not _matches_head(record, head_sha):
472 return None
473 return {**dict(panel), "probed": False, "source": SOURCE_RUN_LEDGER}
476def is_ship_run_for_head(record: Mapping[str, Any] | None, *, head_sha: Any) -> bool:
477 """Did a ``ship_run`` for **this exact head** leave a record? (#1068)
479 Presence, not content: the run's ``run_context.jury_panel`` may say ``fallback``,
480 ``block``, or ``None``, and this still answers ``True``. That separation is the
481 whole point — :func:`pin` needs to know *whether the run left a record here* before it
482 reads what the record says, because a record that says nothing about a panel is still
483 a run that did not ship under one.
484 """
485 return isinstance(record, Mapping) and _matches_head(record, head_sha)
488def states_panel(record: Mapping[str, Any] | None) -> bool:
489 """Does this record carry the ``jury_panel`` **key** at all? (#1068 round 7)
491 Not what it says — whether the run that wrote it had the word. The distinction is
492 between a run that was *asked* about the panel and answered (even by answering
493 ``None``: "this tier named no panel") and a record written before the field existed,
494 which was never asked.
496 :func:`keel.ledger._run_context` always writes the key, ``None`` included, so for every
497 record this feature produces the answer is ``True`` and rank 3 of :func:`pin` is
498 unchanged: a same-head record whose ``jury_panel`` is ``null`` is the run saying it did
499 not ship under a panel, and it silences the lower sources. A ledger row from before
500 #1066 has no such key — missing vocabulary, not a statement — and silencing on its
501 behalf would put words in a run's mouth: on a workstation carrying one, a change the
502 panel really did jury would have had its posted ballots ignored and
503 ``review-verdict-1..3`` demanded by a probe of the local machine. Those rows fall
504 through to the closure comment and then to the ballots, which is exactly what they did
505 before #1066 existed.
507 A record whose ``run_context`` is missing or unreadable has no key either, and reads the
508 same way, for the same reason: absence of vocabulary, not a statement.
509 """
510 context = record.get("run_context") if isinstance(record, Mapping) else None
511 return isinstance(context, Mapping) and "jury_panel" in context
514def pin(
515 record: Mapping[str, Any] | None,
516 *,
517 head_sha: Any,
518 panel_verdict_posted: bool,
519 closure_panel_decision: Any = None,
520) -> dict[str, Any] | None:
521 """**The single authority on "what did this run ship under".** (#1066, #1068)
523 Every verification surface — ``keel evidence-verify``, ``keel merge``, and
524 :func:`keel.cli._shipped_jury_availability`, which is only this function's thin-I/O
525 wrapper — resolves that question here and nowhere else. It is one function rather than
526 an order of ``if``-statements at a call site because the *precedence* between the two
527 pins is itself a rule, and #1068 rounds 2–4 each found a rule written in one place and
528 forgotten in its twin. There is one place now, and this docstring is it.
530 ``None`` means "nothing pins this head": the caller measures its own machine, exactly
531 as it did before either pin existed. That is the fail-closed answer, never a waiver —
532 a probe can only add the panel requirement back or demand the tier's host verdicts.
534 **Both run-record sources select the *latest* record for this head, and that direction
535 is the rule** (#1068 round 7). The ledger source resolves through
536 :func:`keel.ledger.latest_ship_run_for_pr`, which walks the chronologically appended
537 records and keeps the **last** match; the closure source resolves through
538 :func:`keel.evidence.shipped_panel_decision`, which walks ``pr_comments`` in GitHub's
539 oldest-first order and keeps the **last** match. They are two copies of one statement —
540 ranks 2 and 4 are the same sentence read off two artifacts — so if they disagreed about
541 which run they were quoting, the precedence between them would be meaningless: the
542 machine with the ledger would answer for the newest ship and the machine without it for
543 the oldest. Round 6 had exactly that, and worse: the closure source was first-match
544 *and* a run whose panel sat rendered no marker, so a commit shipped once under the
545 fallback and then again on a machine that could staff the panel left one marker on the
546 pull request saying ``fallback``, and CI — where there is no ledger — pinned the
547 host-bench contract onto a panel-reviewed change and never asked for the panel's own
548 verdict. :func:`keel.closure._jury_panel` now emits ``decision=available`` too, so every
549 ship speaks and "latest wins" is well defined on both sides.
550 ``tests/test_juryavail.py::TestBothRunRecordSourcesSelectTheLatest`` holds them to it.
552 The order, and why it is this way round:
554 1. **No head, no pin.** :func:`is_pinnable_head`. A pin removes requirements — it takes
555 ``review-verdict-1..3`` off the required set outright — so it may only ever be taken
556 against an exact commit. ``panel_verdict_posted`` is ignored here even when ``True``,
557 because :func:`keel.evidence.panel_verdict_posted` reads a blank head as *no head
558 filter*: right for counting evidence, wrong for a pin.
559 2. **The run's own ledger record for this head wins** (:func:`is_ship_run_for_head`,
560 then :func:`shipped`). The ledger records what *this run actually did*; a posted
561 verdict records what somebody put on the pull request. At the same head the two can
562 disagree, and then the ledger is the one that is evidence of the run: a ship that
563 measured ``fallback`` seated three host reviewers and owes ``review-verdict-1..3``,
564 and a leftover jury verdict at that head — from an earlier ship of the same commit,
565 from a force-push back onto it, or from a collaborator who ran ``jury`` by hand —
566 is not that run's review. Letting the verdict win dropped three required items on
567 the strength of a comment nobody's run had promised.
568 3. **A same-head record that says nothing about a panel still speaks — if it had the
569 word.** :func:`shipped` returns ``None`` for a record whose ``run_context.jury_panel``
570 is ``null``, malformed, or ``block``, and that ``None`` is returned as-is rather than
571 falling through to a comment: this run left a record here and it does not say the run
572 shipped under a panel, so nobody may say otherwise on its behalf.
574 The one record that does *not* speak is one whose ``run_context`` never carried the
575 key (:func:`states_panel`) — a ledger row written before #1066, or one whose
576 ``run_context`` is unreadable. That is missing vocabulary, not a statement, and
577 silencing the lower sources on its behalf would answer a question the run was never
578 asked: a change the panel really did jury, verified on the workstation that still
579 has that row, would have had its posted ballots ignored and ``review-verdict-1..3``
580 demanded by a probe of *this* machine. Such a record falls through to rank 4 and then
581 rank 5, which is what it did before #1066 existed. Every record this feature writes
582 carries the key — :func:`keel.ledger._run_context` always writes it, ``None``
583 included — so the fall-through applies to legacy rows and to nothing else.
584 4. **Failing that, the run's own closure comment** (:func:`recorded`), whose decision
585 :func:`keel.evidence.shipped_panel_decision` has already held to a trusted author,
586 to an actual closure comment, and to this head. This rank is what makes rank 2 mean
587 anything off the shipping workstation (#1068 round 6): the ledger lives under the
588 gitignored ``.keel/state/``, so a hosted ``evidence-verify`` or ``merge`` has no
589 same-head record at all and fell straight through to the verdict — the run's
590 fallback was outranked by a leftover comment on every machine except the one that
591 had no need of the rule. The closure comment is the *same statement* as the ledger
592 record it was rendered from, in the one place that travels with the pull request,
593 so it ranks with the ledger and above the verdict. Since round 7 it is silent only
594 for a head no keel closure comment names — a run that shipped under a staffable panel
595 records ``decision=available`` and pins the panel here rather than leaving rank 5 to
596 infer it from ballots.
598 Silent and refusing are different answers, and rank 4 gives both: no marker for this
599 head is silence and rank 5 gets its turn, while a marker that says ``block`` or
600 something unrecognised is a record that does not say the run shipped under a panel,
601 and :func:`recorded` returns ``None`` as the answer for the same reason rank 3 does.
602 5. **Only with no record of the run's own does a posted verdict pin**
603 (:func:`panel_sat`). Head-pinned ballots prove a panel *sat* for this head; they do
604 not prove this run's review **was** that panel, which is why they rank last.
606 Every pin therefore runs in the same direction: the strongest available statement about
607 *this run at this head*, falling back to measuring rather than to guessing.
608 """
609 if not is_pinnable_head(head_sha):
610 return None
611 if is_ship_run_for_head(record, head_sha=head_sha) and states_panel(record):
612 return shipped(record, head_sha=head_sha)
613 if closure_panel_decision is not None:
614 return recorded(closure_panel_decision)
615 return panel_sat() if panel_verdict_posted else None
618def is_pinnable_head(head_sha: Any) -> bool:
619 """Is ``head_sha`` a head a pin may be taken against? (#1068)
621 **The one blank-head rule both panel pins read.** A pin republishes an earlier run's
622 panel decision in place of measuring this machine, so it may only ever be taken against
623 an exact commit: an unknown head must not be authorized by a record — or a comment —
624 from some other one. :func:`keel.ledger.gates_pass_for_head` already holds the merge
625 gate to this, and :func:`shipped` to the ledger pin; round 3 hardened those and left
626 the posted-verdict pin reading :func:`keel.evidence._matches_head`, which treats a
627 blank head as *unfiltered* and so counted any trusted jury marker on the pull request
628 as this head's. Written twice, hardened once. It is written here now, and :func:`pin`
629 — the one place the two sources are ranked — asks it before either is consulted.
631 :func:`keel.evidence._matches_head` is deliberately the other rule and stays that way:
632 it filters *evidence items* inside a gate that, with no head resolved, runs
633 head-agnostic throughout — every review verdict counts too. Nothing reached through it
634 removes a requirement. A pin does: it takes ``review-verdict-1..3`` off the required
635 set entirely.
637 **The test is what the answer can do, not which module asks.** So this predicate has a
638 third caller that is not a pin: :func:`keel.evidence.jury_participating_vendors`, whose
639 count downgrades a gating jury to advisory below ``jury.min_vendors`` and thereby drops
640 ``jury-verdict`` from the required evidence (#1069). It removes a requirement, so it
641 takes this rule; its two panel-shaped siblings there cannot, so they keep the
642 permissive one.
643 """
644 return isinstance(head_sha, str) and bool(head_sha.strip())
647def _matches_head(record: Mapping[str, Any], head_sha: str | None) -> bool:
648 """Was this ledger record written for exactly ``head_sha``? A blank head never matches."""
649 if not is_pinnable_head(head_sha):
650 return False
651 git = record.get("git")
652 recorded = git.get("head_sha") if isinstance(git, Mapping) else None
653 return isinstance(recorded, str) and recorded == head_sha
656def _text(value: Any) -> str | None:
657 """A non-blank string, or ``None`` — so a blank field reads as unset."""
658 return value.strip() if isinstance(value, str) and value.strip() else None