4836c6cb48
CI / Lint (ruff) (pull_request) Successful in 9s
CI / Tests (py3.10 / ubuntu-latest) (pull_request) Successful in 11s
CI / Tests (py3.12 / ubuntu-latest) (pull_request) Successful in 10s
CI / Tests (py3.12 / windows-latest) (pull_request) Successful in 35s
CI / Tests (py3.13 / ubuntu-latest) (pull_request) Successful in 10s
CI / Catalog signature (pull_request) Failing after 6s
Wires the maintainer's freshly-rotated signing keys into BCC (issue #68 finding 5: the old key was shared between catalog+release and had been exposed to CI), adds a key-rotation re-attestation mode to the Catalog Console so the Console can actually re-sign under the new key, and hardens the show-seed-b64 CLI prompt that led to a private key being pasted into a chat. ## New keys (#68 finding 5) - bcc_core.CATALOG_PUBKEYS -> new CATALOG pubkey (0s24PmkZcTT5yxNDdyTPHl5fyxArrHNPJKBjnXoQd8k=). Signs data/catalog.json, Console-only, offline. - .github/workflows/ci.yml EXPECTED_CATALOG_PUBKEY_B64 -> same new catalog pubkey (the CI trust anchor added for #68 finding 4). - scripts/sign_checksums.py RELEASE_PUBKEYS -> new RELEASE pubkey (6BnPgJEHJFyVltFoLTCNadIsehjy00iiW8IRlC1TfhA=). Signs SHA256SUMS only, CI-resident. - README.md 'Verifying your download' / 'Signing keys' sections filled in with both pubkeys, explicit about which key is which. The old key 082NOwVB7uURkvfyS3+knJ+40Fk6C9unsF47+2uPKo4= is retired (it was shared and exposed to CI) and is deliberately NOT retained in either trust list -- keeping a burned key in CATALOG_PUBKEYS would defeat the point of rotating it. ## Key-rotation re-attestation mode (Catalog Console) Problem: after the key swap, the existing data/catalog.json.sig (signed with the OLD key) no longer verifies under the NEW CATALOG_PUBKEYS, but catalog content is unchanged, so diff_catalogs(last_signed, current) is empty -- and can_sign()'s empty-diff guard (load-bearing, #68 finding 1) correctly refuses to sign an empty changeset. Without a rotation-aware path, the Console could never re-sign and CI would stay red forever. Fix: treat rotation as a full re-attestation, not a diff. - catalog_review.py: ReviewSession/start_review gain reattest: bool = False. When set, changes is built via diff_catalogs(None, new_catalog) -- every entry presented as if newly added, requiring a fresh acknowledgement -- instead of diffing against old_catalog. can_sign() is UNCHANGED: it still refuses a genuinely-empty changeset and still enforces the TOCTOU blob-SHA pin and the blocking-risk check, because reattest sessions simply never produce an empty changeset (unless the catalog itself is empty). - catalog_console.py: adds catalog_signature_valid_at(repo_dir, commit, raw), which calls bcc_core.verify_catalog_signature directly (never reimplemented) to detect whether the committed .sig verifies under the CURRENT CATALOG_PUBKEYS. ReviewWindow._on_load's source=main path uses this to decide reattest=True/False, and shows a loud, explicit red banner ("KEY ROTATION IN PROGRESS...") whenever reattest mode is entered -- never silent. Status text and Sign-button gating flow through the same can_sign()/all_entries_acknowledged() path as normal review. - tests/test_catalog_review.py: 6 new tests covering re-attest mode (one change-entry per server, gating until all acknowledged, then permits), confirming normal mode still refuses an empty diff (rotation path is not a general bypass), and confirming reattest mode still enforces the blocking-risk check and the TOCTOU pin. ## show-seed-b64 hardening (Task 3) The maintainer ran show-seed-b64 --release, saw an ambiguous prompt, and pasted the printed PRIVATE seed into a chat believing it was public. - cmd_show_seed_b64: passphrase prompt is now explicit ("Passphrase for the release signing key (the one YOU chose when generating it)"). A loud three-line warning banner prints to STDERR immediately before the seed ("!!! PRIVATE KEY BELOW..."); the seed itself stays alone on STDOUT so piping into pbcopy or a CI secret field still works cleanly. - cmd_keygen: labels for both key kinds now say PUBLIC/PRIVATE explicitly and state safety properties inline (safe to commit vs. never commit), so the printed output can't be mistaken for the other key's. ## Verification - ruff check . / ruff format --check .: clean. - pytest: full suite green (360 passed, 1 skipped). - Delete-the-check-and-watch-it-fail: each of can_sign()'s four gates (empty-diff, TOCTOU pin, blocking-risk, acknowledge-all) and the reattest branch itself were individually removed and confirmed to turn a test red, then restored -- see PR description for the exact failures. Not touched: data/catalog.json (content-signed, maintainer's job via the Console). No private key generated, requested, or committed. bcc.py untouched (open PR #67).
854 lines
32 KiB
Python
854 lines
32 KiB
Python
"""
|
|
catalog_review.py -- pure, GUI-free review/diff/risk/crypto logic for the
|
|
Catalog Console (issue #62).
|
|
|
|
This module is deliberately Qt-free and network-free so every function in it
|
|
is unit-testable offline, exactly like bcc_core.py. catalog_console.py (the
|
|
PySide6 GUI) is a thin shell over these functions -- it owns Qt widgets,
|
|
subprocess/git calls, and HTTP registry lookups; this module owns judgment.
|
|
|
|
Nothing here is reimplemented from bcc_core: the command allowlist and the
|
|
signature domain-separation prefix are imported, not retyped, so the two
|
|
modules cannot silently drift apart (see bcc_core.validate_catalog /
|
|
bcc_core.verify_catalog_signature and the project's "the check drifted on a
|
|
new surface" recurring-bug lesson).
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
import re
|
|
from collections.abc import Callable
|
|
from dataclasses import dataclass, field
|
|
from urllib.parse import urlsplit
|
|
|
|
from bcc_core import _CATALOG_SIG_DOMAIN as CATALOG_SIG_DOMAIN
|
|
from bcc_core import CATALOG_ALLOWED_COMMANDS
|
|
from bcc_core import verify_catalog_signature as _verify_catalog_signature
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# Semantic diff
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
# Top-level scalar/simple fields compared directly (not drilled into).
|
|
_DIFF_FIELDS = (
|
|
"display",
|
|
"description",
|
|
"category",
|
|
"official",
|
|
"setup",
|
|
"homepage",
|
|
"docs_url",
|
|
"source",
|
|
"notes",
|
|
"stars",
|
|
"last_release",
|
|
"env_required",
|
|
)
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class FieldChange:
|
|
"""One field that differs between the old and new version of an entry."""
|
|
|
|
field: str
|
|
old: object
|
|
new: object
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class EntryChange:
|
|
"""One catalog entry's change: added, removed, or changed.
|
|
|
|
`old`/`new` are the raw entry dicts (or None for added/removed) so risk
|
|
predicates and the GUI can inspect anything not captured by
|
|
`field_changes` (which only lists fields that actually differ).
|
|
"""
|
|
|
|
entry_id: str
|
|
status: str # "added" | "removed" | "changed"
|
|
old: dict | None
|
|
new: dict | None
|
|
field_changes: tuple[FieldChange, ...] = ()
|
|
|
|
|
|
def _config_field_changes(old_cfg: dict | None, new_cfg: dict | None) -> list[FieldChange]:
|
|
old_cfg = old_cfg or {}
|
|
new_cfg = new_cfg or {}
|
|
changes: list[FieldChange] = []
|
|
for f in ("command", "args", "env"):
|
|
ov, nv = old_cfg.get(f), new_cfg.get(f)
|
|
if ov != nv:
|
|
changes.append(FieldChange(f"config.{f}", ov, nv))
|
|
return changes
|
|
|
|
|
|
def _entry_field_changes(old_entry: dict, new_entry: dict) -> tuple[FieldChange, ...]:
|
|
changes: list[FieldChange] = []
|
|
for f in _DIFF_FIELDS:
|
|
ov, nv = old_entry.get(f), new_entry.get(f)
|
|
if ov != nv:
|
|
changes.append(FieldChange(f, ov, nv))
|
|
changes.extend(_config_field_changes(old_entry.get("config"), new_entry.get("config")))
|
|
return tuple(changes)
|
|
|
|
|
|
def diff_catalogs(old: dict | None, new: dict | None) -> list[EntryChange]:
|
|
"""Semantic (per-entry) diff between two parsed catalog dicts.
|
|
|
|
NOT a text diff: entries are matched by `id`, and each changed entry
|
|
reports exactly which fields differ (with before/after values), which is
|
|
what lets the Console render "command changed from X to Y" instead of a
|
|
JSON line diff a reviewer has to mentally reconstruct.
|
|
|
|
Entries missing/malformed `id` are ignored here -- that is a
|
|
validate_catalog() rejection, not a diffing concern, and diffing must not
|
|
silently invent a match for two differently-broken entries.
|
|
"""
|
|
old_servers = {
|
|
e["id"]: e
|
|
for e in (old or {}).get("servers", []) or []
|
|
if isinstance(e, dict) and isinstance(e.get("id"), str) and e.get("id")
|
|
}
|
|
new_servers = {
|
|
e["id"]: e
|
|
for e in (new or {}).get("servers", []) or []
|
|
if isinstance(e, dict) and isinstance(e.get("id"), str) and e.get("id")
|
|
}
|
|
|
|
changes: list[EntryChange] = []
|
|
for entry_id in sorted(set(old_servers) | set(new_servers)):
|
|
old_e = old_servers.get(entry_id)
|
|
new_e = new_servers.get(entry_id)
|
|
if old_e is None:
|
|
changes.append(EntryChange(entry_id, "added", None, new_e, ()))
|
|
elif new_e is None:
|
|
changes.append(EntryChange(entry_id, "removed", old_e, None, ()))
|
|
elif old_e != new_e:
|
|
fc = _entry_field_changes(old_e, new_e)
|
|
if fc:
|
|
changes.append(EntryChange(entry_id, "changed", old_e, new_e, fc))
|
|
return changes
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# Risk annotations -- each predicate is pure and independently unit-tested.
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class RiskFinding:
|
|
severity: str # "blocking" | "warning" | "info"
|
|
code: str
|
|
message: str
|
|
|
|
|
|
def _escape_non_ascii(s: str) -> str:
|
|
"""Render a string with any non-ASCII code point shown as an escape
|
|
sequence, so a homoglyph/RTL-override character can't visually pass as
|
|
the real thing in the review UI."""
|
|
return s.encode("unicode_escape").decode("ascii")
|
|
|
|
|
|
def risk_env_required(change: EntryChange) -> list[RiskFinding]:
|
|
"""A non-empty env_required value is blocking: catalog entries must ship
|
|
only the *names* of env vars the user fills in, never values."""
|
|
entry = change.new or {}
|
|
env_required = entry.get("env_required")
|
|
findings: list[RiskFinding] = []
|
|
if isinstance(env_required, dict):
|
|
for k, v in env_required.items():
|
|
if v not in (None, ""):
|
|
findings.append(
|
|
RiskFinding(
|
|
"blocking",
|
|
"env_required_value",
|
|
f"env_required[{k!r}] carries a non-empty value -- catalog "
|
|
"entries must never ship secret values, only placeholder names.",
|
|
)
|
|
)
|
|
return findings
|
|
|
|
|
|
def risk_command_allowlist(change: EntryChange) -> list[RiskFinding]:
|
|
"""A command outside bcc_core.CATALOG_ALLOWED_COMMANDS is blocking.
|
|
Imports the allowlist rather than redefining it."""
|
|
entry = change.new or {}
|
|
config = entry.get("config") or {}
|
|
command = config.get("command")
|
|
if isinstance(command, str) and command and command not in CATALOG_ALLOWED_COMMANDS:
|
|
return [
|
|
RiskFinding(
|
|
"blocking",
|
|
"command_not_allowed",
|
|
f"command {command!r} is not on the catalog allowlist "
|
|
f"({', '.join(sorted(CATALOG_ALLOWED_COMMANDS))}).",
|
|
)
|
|
]
|
|
return []
|
|
|
|
|
|
def risk_non_ascii(change: EntryChange) -> list[RiskFinding]:
|
|
"""Non-ASCII code points in id/command/args are blocking -- homoglyph /
|
|
RTL-override typosquatting can make a malicious package name visually
|
|
identical to a legitimate one in a naive diff view."""
|
|
entry = change.new or {}
|
|
findings: list[RiskFinding] = []
|
|
|
|
entry_id = entry.get("id")
|
|
if isinstance(entry_id, str) and not entry_id.isascii():
|
|
findings.append(
|
|
RiskFinding(
|
|
"blocking",
|
|
"non_ascii_id",
|
|
f"id contains non-ASCII code points: {_escape_non_ascii(entry_id)!r}",
|
|
)
|
|
)
|
|
|
|
config = entry.get("config") or {}
|
|
command = config.get("command")
|
|
if isinstance(command, str) and not command.isascii():
|
|
findings.append(
|
|
RiskFinding(
|
|
"blocking",
|
|
"non_ascii_command",
|
|
f"command contains non-ASCII code points: {_escape_non_ascii(command)!r}",
|
|
)
|
|
)
|
|
|
|
for a in config.get("args") or []:
|
|
if isinstance(a, str) and not a.isascii():
|
|
findings.append(
|
|
RiskFinding(
|
|
"blocking",
|
|
"non_ascii_arg",
|
|
f"arg contains non-ASCII code points: {_escape_non_ascii(a)!r}",
|
|
)
|
|
)
|
|
|
|
return findings
|
|
|
|
|
|
def _npm_candidate_args(command: str | None, args: list[str]) -> list[str]:
|
|
if command != "npx":
|
|
return []
|
|
return [a for a in args if isinstance(a, str) and a and not a.startswith("-")]
|
|
|
|
|
|
def split_npm_spec(spec: str) -> tuple[str, str | None]:
|
|
"""Split an npm package spec into (name, version). version is None if
|
|
unpinned. Handles scoped (@scope/name@version) and unscoped
|
|
(name@version) specs."""
|
|
if spec.startswith("@"):
|
|
rest = spec[1:]
|
|
if "/" not in rest:
|
|
return spec, None # malformed scope, can't tell -- treat unpinned
|
|
scope, _, remainder = rest.partition("/")
|
|
if "@" in remainder:
|
|
pkg_name, _, version = remainder.partition("@")
|
|
return f"@{scope}/{pkg_name}", (version or None)
|
|
return f"@{scope}/{remainder}", None
|
|
if "@" in spec:
|
|
name, _, version = spec.partition("@")
|
|
return name, (version or None)
|
|
return spec, None
|
|
|
|
|
|
def is_pinned_npm_spec(spec: str) -> bool:
|
|
_name, version = split_npm_spec(spec)
|
|
return bool(version)
|
|
|
|
|
|
_DOCKER_VALUE_FLAGS = {
|
|
"-e",
|
|
"--env",
|
|
"-v",
|
|
"--volume",
|
|
"-p",
|
|
"--publish",
|
|
"-w",
|
|
"--workdir",
|
|
"-u",
|
|
"--user",
|
|
"--name",
|
|
"--network",
|
|
"--entrypoint",
|
|
}
|
|
|
|
|
|
def docker_image_candidates(args: list[str]) -> list[str]:
|
|
"""Best-effort extraction of the image reference from a `docker run
|
|
[OPTIONS] IMAGE [CMD...]` args list: the first positional token after
|
|
any leading `run` and flag(+value) pairs."""
|
|
candidates: list[str] = []
|
|
i = 0
|
|
while i < len(args):
|
|
a = args[i]
|
|
if a == "run":
|
|
i += 1
|
|
continue
|
|
if isinstance(a, str) and a.startswith("-"):
|
|
if "=" not in a and a in _DOCKER_VALUE_FLAGS:
|
|
i += 2
|
|
continue
|
|
i += 1
|
|
continue
|
|
if isinstance(a, str):
|
|
candidates.append(a)
|
|
break # first positional token after `run` is the image ref
|
|
return candidates
|
|
|
|
|
|
def is_pinned_docker_image(image: str) -> bool:
|
|
if "@sha256:" in image:
|
|
return True
|
|
tag_part = image.rsplit("/", 1)[-1]
|
|
if ":" not in tag_part:
|
|
return False # no tag => implicit :latest
|
|
tag = tag_part.rsplit(":", 1)[-1]
|
|
return bool(tag) and tag != "latest"
|
|
|
|
|
|
def risk_unpinned_package(change: EntryChange) -> list[RiskFinding]:
|
|
"""Every entry must pin an exact version: an `@scope/pkg` npm arg with
|
|
no `@version`, or a docker image with no tag / `:latest`, is blocking.
|
|
A later-compromised package must not be able to auto-upgrade into every
|
|
user just because the catalog entry never pinned a version."""
|
|
entry = change.new or {}
|
|
config = entry.get("config") or {}
|
|
command = config.get("command")
|
|
args = config.get("args") or []
|
|
findings: list[RiskFinding] = []
|
|
|
|
if command == "npx":
|
|
for a in _npm_candidate_args(command, args):
|
|
if not is_pinned_npm_spec(a):
|
|
findings.append(
|
|
RiskFinding(
|
|
"blocking",
|
|
"unpinned_npm_package",
|
|
f"npm package arg {a!r} has no pinned @version.",
|
|
)
|
|
)
|
|
elif command == "docker":
|
|
for img in docker_image_candidates(args):
|
|
if not is_pinned_docker_image(img):
|
|
findings.append(
|
|
RiskFinding(
|
|
"blocking",
|
|
"unpinned_docker_image",
|
|
f"docker image {img!r} is not pinned to an exact tag "
|
|
"(uses :latest or no tag).",
|
|
)
|
|
)
|
|
|
|
return findings
|
|
|
|
|
|
_URL_FIELDS = ("homepage", "docs_url", "source")
|
|
|
|
|
|
def _domain(url: str) -> str:
|
|
try:
|
|
return urlsplit(url).netloc.lower()
|
|
except ValueError:
|
|
return ""
|
|
|
|
|
|
def risk_url_domain_change(change: EntryChange) -> list[RiskFinding]:
|
|
"""Non-https URLs and, more importantly, a *domain change* on any URL
|
|
field are surfaced loudly with old-vs-new domains broken out -- the
|
|
lookalike-domain-swap defence."""
|
|
findings: list[RiskFinding] = []
|
|
old_entry = change.old or {}
|
|
new_entry = change.new or {}
|
|
|
|
for f in _URL_FIELDS:
|
|
new_url = new_entry.get(f)
|
|
if not isinstance(new_url, str) or not new_url:
|
|
continue
|
|
if not new_url.startswith("https://"):
|
|
findings.append(
|
|
RiskFinding("warning", "non_https_url", f"{f} is not https://: {new_url!r}")
|
|
)
|
|
old_url = old_entry.get(f)
|
|
if isinstance(old_url, str) and old_url:
|
|
old_domain, new_domain = _domain(old_url), _domain(new_url)
|
|
if old_domain and new_domain and old_domain != new_domain:
|
|
findings.append(
|
|
RiskFinding(
|
|
"warning",
|
|
"domain_changed",
|
|
f"{f} domain changed from {old_domain!r} to {new_domain!r} -- "
|
|
"verify this isn't a lookalike-domain swap.",
|
|
)
|
|
)
|
|
return findings
|
|
|
|
|
|
def risk_new_entry(change: EntryChange) -> list[RiskFinding]:
|
|
"""A brand-new entry is flagged for extra scrutiny -- not blocking on its
|
|
own, but it's the category of change the registry lookup exists for."""
|
|
if change.status == "added":
|
|
return [
|
|
RiskFinding(
|
|
"info",
|
|
"new_entry",
|
|
"Brand-new catalog entry -- extra scrutiny: check publisher identity "
|
|
"via the registry lookup before signing.",
|
|
)
|
|
]
|
|
return []
|
|
|
|
|
|
_RISK_PREDICATES: tuple[Callable[[EntryChange], list[RiskFinding]], ...] = (
|
|
risk_env_required,
|
|
risk_command_allowlist,
|
|
risk_non_ascii,
|
|
risk_unpinned_package,
|
|
risk_url_domain_change,
|
|
risk_new_entry,
|
|
)
|
|
|
|
|
|
def entry_risk_findings(change: EntryChange) -> list[RiskFinding]:
|
|
"""Run every risk predicate against one entry change and return the
|
|
combined findings (order matches _RISK_PREDICATES)."""
|
|
findings: list[RiskFinding] = []
|
|
for predicate in _RISK_PREDICATES:
|
|
findings.extend(predicate(change))
|
|
return findings
|
|
|
|
|
|
def has_blocking_risk(change: EntryChange) -> bool:
|
|
return any(f.severity == "blocking" for f in entry_risk_findings(change))
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# Review session: acknowledge-gating + TOCTOU blob-SHA pinning
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
@dataclass
|
|
class ReviewSession:
|
|
"""State for one review pass. `pinned_blob_sha` is the git blob SHA of
|
|
data/catalog.json as it existed the moment review began -- see
|
|
can_sign()/sign_precondition().
|
|
|
|
`loaded_ref` is the exact ref this review was loaded from ("main", or a
|
|
PR's `refs/pull/<n>/head`) -- see issue #68 finding 1. It exists so the
|
|
Sign path can re-resolve the TOCTOU blob SHA from *the ref that was
|
|
actually reviewed*, instead of a hardcoded "main" that silently diverges
|
|
from the reviewed ref on every PR review (the bug that made the PR path
|
|
unable to sign at all, and forced everyone onto the vacuous
|
|
main-vs-itself path instead).
|
|
|
|
`reattest` marks a KEY-ROTATION re-attestation pass (issue #68 finding 5
|
|
follow-up): the currently-trusted `bcc_core.CATALOG_PUBKEYS` key changed
|
|
and the existing `data/catalog.json.sig` no longer verifies under it.
|
|
Content-wise nothing may have changed -- `diff_catalogs(old, new)` can be
|
|
genuinely empty -- but the NEW key has never vouched for any of this
|
|
catalog before, so every entry needs a first-time attestation under the
|
|
new key, not a diff against the old one. When `reattest` is set,
|
|
`changes` is built as "every entry in `new_catalog`, presented as if
|
|
newly added" (via `diff_catalogs(None, new_catalog)`) instead of a
|
|
diff against `old_catalog`, so the acknowledge-gate in `can_sign()`
|
|
requires re-reviewing everything the new key will sign -- which is the
|
|
intended cost of a key rotation, not a bypass of the empty-diff guard.
|
|
"""
|
|
|
|
pinned_blob_sha: str
|
|
old_catalog: dict
|
|
new_catalog: dict
|
|
loaded_ref: str = "main"
|
|
reattest: bool = False
|
|
changes: list[EntryChange] = field(default_factory=list)
|
|
acknowledged: set[str] = field(default_factory=set)
|
|
|
|
def __post_init__(self) -> None:
|
|
if not self.changes:
|
|
if self.reattest:
|
|
# Every entry in the new catalog is treated as though it
|
|
# were newly added -- because, under the NEW signing key,
|
|
# it is: nothing this key signs was ever attested by it
|
|
# before. Reusing diff_catalogs(None, new) (rather than a
|
|
# bespoke code path) means the same "added" risk predicate
|
|
# (risk_new_entry) and the same EntryChange shape the rest
|
|
# of this module and the GUI already know how to render
|
|
# apply here unmodified.
|
|
self.changes = diff_catalogs(None, self.new_catalog)
|
|
else:
|
|
self.changes = diff_catalogs(self.old_catalog, self.new_catalog)
|
|
|
|
|
|
def start_review(
|
|
pinned_blob_sha: str,
|
|
old_catalog: dict,
|
|
new_catalog: dict,
|
|
loaded_ref: str = "main",
|
|
reattest: bool = False,
|
|
) -> ReviewSession:
|
|
return ReviewSession(
|
|
pinned_blob_sha=pinned_blob_sha,
|
|
old_catalog=old_catalog,
|
|
new_catalog=new_catalog,
|
|
loaded_ref=loaded_ref,
|
|
reattest=reattest,
|
|
)
|
|
|
|
|
|
def find_last_signed_catalog_raw(
|
|
candidates: list[bytes], sig: bytes, pubkeys: list[bytes]
|
|
) -> bytes | None:
|
|
"""Given `candidates` (candidate raw catalog.json byte-strings -- e.g.
|
|
successive historical versions from git log, most-recent-first),
|
|
return the first one whose signature verifies against `sig`/`pubkeys`,
|
|
or None if none do.
|
|
|
|
This is how source="main" review diffs against "the last catalog a
|
|
maintainer actually signed" instead of against itself (issue #68
|
|
finding 1): `catalog_console.last_signed_catalog_raw` walks
|
|
data/catalog.json's git history on main and hands the candidates here.
|
|
"""
|
|
for raw in candidates:
|
|
if _verify_catalog_signature(raw, sig, pubkeys):
|
|
return raw
|
|
return None
|
|
|
|
|
|
def acknowledge_entry(session: ReviewSession, entry_id: str) -> None:
|
|
ids = {c.entry_id for c in session.changes}
|
|
if entry_id not in ids:
|
|
raise ValueError(f"{entry_id!r} is not part of this review session's diff.")
|
|
session.acknowledged.add(entry_id)
|
|
|
|
|
|
def unacknowledge_entry(session: ReviewSession, entry_id: str) -> None:
|
|
session.acknowledged.discard(entry_id)
|
|
|
|
|
|
def all_entries_acknowledged(session: ReviewSession) -> bool:
|
|
return {c.entry_id for c in session.changes} <= session.acknowledged
|
|
|
|
|
|
# NOTE for future editors: do NOT add an "acknowledge all" shortcut here, now
|
|
# or ever. The friction of individually acknowledging every changed entry is
|
|
# the entire point of this tool (issue #62) -- a shortcut would let a tired
|
|
# reviewer rubber-stamp a diff exactly like the "merge PR, run script, push"
|
|
# reflex this Console exists to replace. If this comment is the only thing
|
|
# stopping you, that is the point: it is stopping you on purpose.
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class SignDecision:
|
|
ok: bool
|
|
reason: str | None = None
|
|
|
|
|
|
def can_sign(session: ReviewSession, current_blob_sha: str) -> SignDecision:
|
|
"""Whether the Sign button may fire right now.
|
|
|
|
Four independent gates, all required, checked in this order:
|
|
|
|
0. The diff must be non-empty. An empty diff historically meant "Sign
|
|
unlocks instantly" (`set() <= set()` is vacuously True), which is
|
|
exactly backwards: a vacuously-satisfied gate is worse than no gate
|
|
at all, because it *manufactures confidence* -- the signature looks
|
|
identical to one produced by a real review. "Nothing changed" must
|
|
mean "nothing to sign", never "sign unlocked". (Issue #68 finding 1;
|
|
this is what let commit b08cf21 sign all 19 entries with zero of them
|
|
ever reviewed.)
|
|
|
|
This gate is unaffected by `session.reattest`: a key-rotation
|
|
re-attestation session's `changes` is built from
|
|
`diff_catalogs(None, new_catalog)` (see ReviewSession), which is
|
|
empty ONLY if the catalog itself has zero entries -- a genuinely
|
|
empty catalog either way. Rotation never manufactures a non-empty
|
|
changeset out of an empty one; it just changes *what* "non-empty"
|
|
is computed against.
|
|
1. TOCTOU: `current_blob_sha` (fetched fresh, immediately before signing,
|
|
from the ref that was actually reviewed -- see sign_precondition())
|
|
must match the blob SHA pinned when review began. If the bytes on the
|
|
remote changed since -- a new commit pushed to the same PR, a
|
|
force-push, another PR merged in between -- signing is refused and a
|
|
re-review is forced. This is what makes "signing is the approval act"
|
|
true rather than aspirational: the signature is bound to the exact
|
|
reviewed bytes, not to "whatever the file happens to be now".
|
|
2. No blocking risk finding may be outstanding on ANY changed entry, full
|
|
stop -- checked here, not just in the GUI. The GUI additionally
|
|
disables the acknowledge checkbox for a blocking entry, but that is a
|
|
UI nicety, not the enforcement point: if this pure gate didn't also
|
|
check it, a blocking risk would only be stopped by the GUI happening
|
|
to have wired the checkbox correctly, and nothing would catch a
|
|
regression in that wiring. The GUI must not be the only thing
|
|
standing between a blocking risk and a signature.
|
|
3. Every changed entry in the diff must be individually acknowledged.
|
|
"""
|
|
if not session.changes:
|
|
return SignDecision(
|
|
False,
|
|
"Nothing to sign: this review's diff is empty. If you expected "
|
|
"changes here, you may be diffing the wrong source/ref.",
|
|
)
|
|
if current_blob_sha != session.pinned_blob_sha:
|
|
return SignDecision(
|
|
False,
|
|
"The reviewed bytes changed since this review began (blob SHA "
|
|
"mismatch) -- re-review required before signing.",
|
|
)
|
|
blocking_ids = sorted({c.entry_id for c in session.changes if has_blocking_risk(c)})
|
|
if blocking_ids:
|
|
return SignDecision(
|
|
False,
|
|
"Blocking risk finding(s) outstanding on: "
|
|
f"{', '.join(blocking_ids)} -- fix the underlying change, do not sign around it.",
|
|
)
|
|
if not all_entries_acknowledged(session):
|
|
pending = sorted({c.entry_id for c in session.changes} - session.acknowledged)
|
|
return SignDecision(
|
|
False, f"Not every changed entry has been acknowledged yet: {', '.join(pending)}"
|
|
)
|
|
return SignDecision(True, None)
|
|
|
|
|
|
def sign_precondition(
|
|
session: ReviewSession, resolve_blob_sha: Callable[[str], str]
|
|
) -> SignDecision:
|
|
"""The real Sign-button gate: resolves the current TOCTOU blob SHA from
|
|
*the ref this session was actually loaded from* (`session.loaded_ref`),
|
|
never a hardcoded "main", then delegates to can_sign().
|
|
|
|
`resolve_blob_sha` is injected so this stays testable without git/Qt --
|
|
catalog_console.ReviewWindow._on_sign passes a real resolver
|
|
(fetch_ref + blob_sha_at against self.repo_dir); tests pass a fake
|
|
dict-backed lookup. This is the fix for issue #68 finding 1's first bug:
|
|
`_on_sign` used to hardcode `fetch_ref(self.repo_dir, "main")` as the
|
|
comparison ref, so for any PR review (where `loaded_ref` is the PR's
|
|
head, not main) the SHAs differed by definition and Sign could never
|
|
fire -- and the retry path re-called the same hardcoded resolver, so it
|
|
re-pinned the same wrong value and looped forever instead of forcing a
|
|
genuine re-review.
|
|
"""
|
|
current_blob_sha = resolve_blob_sha(session.loaded_ref)
|
|
return can_sign(session, current_blob_sha)
|
|
|
|
|
|
def catalog_signing_message(raw_bytes: bytes) -> bytes:
|
|
"""The exact bytes that get signed: bcc_core's domain-separation prefix
|
|
(imported, never retyped) + the raw catalog bytes. Using this function
|
|
guarantees the Console's signature and bcc_core.verify_catalog_signature
|
|
can never drift apart on the prefix."""
|
|
return CATALOG_SIG_DOMAIN + raw_bytes
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# Key management: passphrase-encrypted-at-rest Ed25519 seed
|
|
#
|
|
# The private key is NEVER stored plaintext, never an env var, never
|
|
# committed. encrypt_private_key/decrypt_private_key are pure and offline
|
|
# (scrypt KDF + AES-256-GCM via `cryptography`, already a project
|
|
# dependency); catalog_console.py decides WHERE the resulting blob lives
|
|
# (OS keychain if available, else a file outside the repo).
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
_KDF_SALT_LEN = 16
|
|
_KDF_N = 2**15 # scrypt cost parameter, tuned for a one-off interactive unlock
|
|
_KDF_R = 8
|
|
_KDF_P = 1
|
|
_NONCE_LEN = 12
|
|
_AAD = b"bcc-catalog-console-key-v1"
|
|
|
|
|
|
def _derive_key(passphrase: str, salt: bytes) -> bytes:
|
|
from cryptography.hazmat.primitives.kdf.scrypt import Scrypt
|
|
|
|
kdf = Scrypt(salt=salt, length=32, n=_KDF_N, r=_KDF_R, p=_KDF_P)
|
|
return kdf.derive(passphrase.encode("utf-8"))
|
|
|
|
|
|
def generate_keypair() -> tuple[bytes, bytes]:
|
|
"""Generate a new Ed25519 keypair. Returns (seed_32_bytes, pubkey_32_bytes)."""
|
|
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey
|
|
from cryptography.hazmat.primitives.serialization import Encoding, PublicFormat
|
|
|
|
private_key = Ed25519PrivateKey.generate()
|
|
seed = private_key.private_bytes_raw()
|
|
pubkey = private_key.public_key().public_bytes(Encoding.Raw, PublicFormat.Raw)
|
|
return seed, pubkey
|
|
|
|
|
|
def encrypt_private_key(seed: bytes, passphrase: str) -> bytes:
|
|
"""Encrypt a 32-byte Ed25519 seed at rest with a passphrase. Returns a
|
|
self-contained blob: salt || nonce || ciphertext+tag."""
|
|
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
|
|
|
|
if len(seed) != 32:
|
|
raise ValueError(f"expected a 32-byte raw Ed25519 seed, got {len(seed)} bytes")
|
|
if not passphrase:
|
|
raise ValueError("a non-empty passphrase is required")
|
|
salt = os.urandom(_KDF_SALT_LEN)
|
|
key = _derive_key(passphrase, salt)
|
|
nonce = os.urandom(_NONCE_LEN)
|
|
ciphertext = AESGCM(key).encrypt(nonce, seed, _AAD)
|
|
return salt + nonce + ciphertext
|
|
|
|
|
|
def decrypt_private_key(blob: bytes, passphrase: str) -> bytes:
|
|
"""Decrypt a blob produced by encrypt_private_key. Raises ValueError on a
|
|
wrong passphrase or corrupt blob -- never silently returns garbage."""
|
|
from cryptography.exceptions import InvalidTag
|
|
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
|
|
|
|
if len(blob) < _KDF_SALT_LEN + _NONCE_LEN:
|
|
raise ValueError("key blob is too short to be valid")
|
|
salt = blob[:_KDF_SALT_LEN]
|
|
nonce = blob[_KDF_SALT_LEN : _KDF_SALT_LEN + _NONCE_LEN]
|
|
ciphertext = blob[_KDF_SALT_LEN + _NONCE_LEN :]
|
|
key = _derive_key(passphrase, salt)
|
|
try:
|
|
return AESGCM(key).decrypt(nonce, ciphertext, _AAD)
|
|
except InvalidTag as e:
|
|
raise ValueError("wrong passphrase or corrupted key file") from e
|
|
|
|
|
|
def sign_catalog_bytes(raw: bytes, seed: bytes) -> bytes:
|
|
"""Sign `raw` catalog bytes with a 32-byte Ed25519 seed, using the exact
|
|
domain-separated message bcc_core.verify_catalog_signature expects."""
|
|
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey
|
|
|
|
if len(seed) != 32:
|
|
raise ValueError(f"expected a 32-byte raw Ed25519 seed, got {len(seed)} bytes")
|
|
private_key = Ed25519PrivateKey.from_private_bytes(seed)
|
|
return private_key.sign(catalog_signing_message(raw))
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# Registry lookup -- the check a human genuinely can't do.
|
|
#
|
|
# The network call itself is injected (a `Fetcher` callable) so this stays
|
|
# testable offline; catalog_console.py supplies the real npm/PyPI HTTP
|
|
# fetcher. Fails soft everywhere: a fetcher returning None/raising just
|
|
# yields RegistryInfo(available=False), never an exception into the caller
|
|
# and never a block on review.
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class PackageRef:
|
|
entry_id: str
|
|
ecosystem: str # "npm" | "pypi"
|
|
name: str
|
|
version: str | None
|
|
|
|
|
|
def split_pypi_spec(spec: str) -> tuple[str, str | None]:
|
|
for sep in ("==", "@"):
|
|
if sep in spec:
|
|
name, _, version = spec.partition(sep)
|
|
return name, (version or None)
|
|
return spec, None
|
|
|
|
|
|
def extract_package_refs(entry: dict) -> list[PackageRef]:
|
|
"""Pull out the package(s) a basic-tier entry's args reference, for the
|
|
registry lookup. Returns [] for link-only entries or entries whose
|
|
command isn't npx/uvx (docker images aren't registry-lookup candidates
|
|
in the npm/PyPI sense used here)."""
|
|
config = entry.get("config") or {}
|
|
command = config.get("command")
|
|
args = config.get("args") or []
|
|
entry_id = entry.get("id", "") if isinstance(entry.get("id"), str) else ""
|
|
refs: list[PackageRef] = []
|
|
|
|
if command == "npx":
|
|
for a in _npm_candidate_args(command, args):
|
|
name, version = split_npm_spec(a)
|
|
if name:
|
|
refs.append(PackageRef(entry_id, "npm", name, version))
|
|
elif command == "uvx":
|
|
for a in args:
|
|
if isinstance(a, str) and a and not a.startswith("-"):
|
|
name, version = split_pypi_spec(a)
|
|
if name:
|
|
refs.append(PackageRef(entry_id, "pypi", name, version))
|
|
break # `uvx <pkg>` -- first positional token is the package
|
|
|
|
return refs
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class RegistryInfo:
|
|
ref: PackageRef
|
|
available: bool
|
|
publisher: str | None = None
|
|
age_days: int | None = None
|
|
last_release: str | None = None
|
|
downloads: int | None = None
|
|
near_neighbor_ids: tuple[str, ...] = ()
|
|
|
|
|
|
Fetcher = Callable[[PackageRef], dict | None]
|
|
|
|
|
|
def edit_distance(a: str, b: str) -> int:
|
|
"""Levenshtein distance, iterative DP (no recursion depth concerns)."""
|
|
if a == b:
|
|
return 0
|
|
la, lb = len(a), len(b)
|
|
if la == 0:
|
|
return lb
|
|
if lb == 0:
|
|
return la
|
|
prev = list(range(lb + 1))
|
|
for i, ca in enumerate(a, 1):
|
|
cur = [i] + [0] * lb
|
|
for j, cb in enumerate(b, 1):
|
|
cost = 0 if ca == cb else 1
|
|
cur[j] = min(prev[j] + 1, cur[j - 1] + 1, prev[j - 1] + cost)
|
|
prev = cur
|
|
return prev[lb]
|
|
|
|
|
|
def near_neighbor_ids(name: str, other_ids: list[str], max_distance: int = 2) -> list[str]:
|
|
"""Catalog ids within `max_distance` edits of `name` (case-insensitive),
|
|
excluding an exact match -- the dependency-confusion / typosquat
|
|
near-neighbour warning."""
|
|
lname = name.lower()
|
|
return [
|
|
oid
|
|
for oid in other_ids
|
|
if oid != name and edit_distance(lname, oid.lower()) <= max_distance
|
|
]
|
|
|
|
|
|
def lookup_registry_info(
|
|
ref: PackageRef, fetcher: Fetcher, all_entry_ids: list[str]
|
|
) -> RegistryInfo:
|
|
"""Resolve one package against the live registry via the injected
|
|
fetcher. Never raises: any fetcher exception or falsy return means
|
|
`available=False` ("unavailable"), which the GUI renders plainly rather
|
|
than blocking or erroring the review."""
|
|
neighbors = tuple(near_neighbor_ids(ref.name, all_entry_ids))
|
|
try:
|
|
raw = fetcher(ref)
|
|
except Exception:
|
|
raw = None
|
|
if not raw:
|
|
return RegistryInfo(ref=ref, available=False, near_neighbor_ids=neighbors)
|
|
return RegistryInfo(
|
|
ref=ref,
|
|
available=True,
|
|
publisher=raw.get("publisher"),
|
|
age_days=raw.get("age_days"),
|
|
last_release=raw.get("last_release"),
|
|
downloads=raw.get("downloads"),
|
|
near_neighbor_ids=neighbors,
|
|
)
|
|
|
|
|
|
_NON_ASCII_RE = re.compile(r"[^\x00-\x7f]")
|
|
|
|
|
|
def contains_non_ascii(s: str) -> bool:
|
|
return bool(_NON_ASCII_RE.search(s))
|