Files
better-claude-config/catalog_review.py
BCC Fix Agent 82483e693d
CI / Lint (ruff) (pull_request) Successful in 9s
CI / Tests (py3.10 / ubuntu-latest) (pull_request) Successful in 11s
CI / Tests (py3.12 / windows-latest) (pull_request) Successful in 34s
CI / Tests (py3.12 / ubuntu-latest) (pull_request) Successful in 9s
CI / Tests (py3.13 / ubuntu-latest) (pull_request) Successful in 10s
CI / Catalog signature (pull_request) Successful in 6s
fix(catalog-console): close the vacuous review gate; split catalog/release signing keys
Fixes #68 findings 1 and 5.

Finding 1 -- the review gate signed without reviewing anything:
- ReviewWindow._on_sign hardcoded "main" as the TOCTOU comparison ref, so
  any PR review (where _on_load pins the PR head's blob SHA) could never
  sign; the only working path was main-vs-itself, whose empty diff made
  can_sign() vacuously True (set() <= set()). Commit b08cf21 signed 19
  entries through exactly that path with zero of them reviewed.
- can_sign() now refuses an empty changeset outright, and itself checks
  has_blocking_risk() across every changed entry rather than trusting the
  GUI to have disabled a checkbox.
- ReviewSession now carries loaded_ref (the exact ref reviewed); a new pure
  sign_precondition(session, resolve_blob_sha) resolves the TOCTOU SHA from
  that ref, never a hardcoded "main". _on_load's retry path re-diffs
  instead of re-pinning the same stale SHA, so a blob-mismatch refusal
  can't loop forever.
- source="main" now diffs against the last catalog a maintainer actually
  SIGNED (walking catalog.json's git history until a version verifies
  against the current .sig), not against itself.
- Replaced the theatre-only test_no_acknowledge_all_function_exists (only
  asserted no function was *named* acknowledge_all) with a test that also
  exercises the real gate. Added can_sign/sign_precondition coverage for
  the empty-diff, blocking-risk, and ref-resolution seams -- each verified
  to fail when its guard is removed.

Finding 5 -- the catalog key and release key were the same CI-resident key:
- scripts/sign_checksums.py gets its own RELEASE_PUBKEYS (separate from
  bcc_core.CATALOG_PUBKEYS) and a verify_checksums_against_any() helper.
- release.yml's signing-smoke-test now verifies RELEASE_SIGNING_KEY against
  RELEASE_PUBKEYS only -- it no longer imports bcc_core/CATALOG_PUBKEYS at
  all, so this workflow can never compare a CI secret against the
  catalog's root of trust.
- catalog_console.py: keygen/show-seed-b64 gain --release, with separate
  keychain/file storage per key kind. show-seed-b64 refuses to run without
  --release, so the catalog seed can't be exported to a CI secret by habit.
- README documents both keys' trust properties and the asymmetry: a CI
  compromise burns the release key, never the catalog key.

The maintainer must rotate the catalog key (it was CI-resident, so treat it
as burned for catalog use) and generate a fresh release key -- see the PR
description for the exact steps. No key is generated or committed here.
2026-07-12 21:25:34 -04:00

819 lines
30 KiB
Python

"""
catalog_review.py -- pure, GUI-free review/diff/risk/crypto logic for the
Catalog Console (issue #62).
This module is deliberately Qt-free and network-free so every function in it
is unit-testable offline, exactly like bcc_core.py. catalog_console.py (the
PySide6 GUI) is a thin shell over these functions -- it owns Qt widgets,
subprocess/git calls, and HTTP registry lookups; this module owns judgment.
Nothing here is reimplemented from bcc_core: the command allowlist and the
signature domain-separation prefix are imported, not retyped, so the two
modules cannot silently drift apart (see bcc_core.validate_catalog /
bcc_core.verify_catalog_signature and the project's "the check drifted on a
new surface" recurring-bug lesson).
"""
from __future__ import annotations
import os
import re
from collections.abc import Callable
from dataclasses import dataclass, field
from urllib.parse import urlsplit
from bcc_core import _CATALOG_SIG_DOMAIN as CATALOG_SIG_DOMAIN
from bcc_core import CATALOG_ALLOWED_COMMANDS
from bcc_core import verify_catalog_signature as _verify_catalog_signature
# --------------------------------------------------------------------------- #
# Semantic diff
# --------------------------------------------------------------------------- #
# Top-level scalar/simple fields compared directly (not drilled into).
_DIFF_FIELDS = (
"display",
"description",
"category",
"official",
"setup",
"homepage",
"docs_url",
"source",
"notes",
"stars",
"last_release",
"env_required",
)
@dataclass(frozen=True)
class FieldChange:
"""One field that differs between the old and new version of an entry."""
field: str
old: object
new: object
@dataclass(frozen=True)
class EntryChange:
"""One catalog entry's change: added, removed, or changed.
`old`/`new` are the raw entry dicts (or None for added/removed) so risk
predicates and the GUI can inspect anything not captured by
`field_changes` (which only lists fields that actually differ).
"""
entry_id: str
status: str # "added" | "removed" | "changed"
old: dict | None
new: dict | None
field_changes: tuple[FieldChange, ...] = ()
def _config_field_changes(old_cfg: dict | None, new_cfg: dict | None) -> list[FieldChange]:
old_cfg = old_cfg or {}
new_cfg = new_cfg or {}
changes: list[FieldChange] = []
for f in ("command", "args", "env"):
ov, nv = old_cfg.get(f), new_cfg.get(f)
if ov != nv:
changes.append(FieldChange(f"config.{f}", ov, nv))
return changes
def _entry_field_changes(old_entry: dict, new_entry: dict) -> tuple[FieldChange, ...]:
changes: list[FieldChange] = []
for f in _DIFF_FIELDS:
ov, nv = old_entry.get(f), new_entry.get(f)
if ov != nv:
changes.append(FieldChange(f, ov, nv))
changes.extend(_config_field_changes(old_entry.get("config"), new_entry.get("config")))
return tuple(changes)
def diff_catalogs(old: dict | None, new: dict | None) -> list[EntryChange]:
"""Semantic (per-entry) diff between two parsed catalog dicts.
NOT a text diff: entries are matched by `id`, and each changed entry
reports exactly which fields differ (with before/after values), which is
what lets the Console render "command changed from X to Y" instead of a
JSON line diff a reviewer has to mentally reconstruct.
Entries missing/malformed `id` are ignored here -- that is a
validate_catalog() rejection, not a diffing concern, and diffing must not
silently invent a match for two differently-broken entries.
"""
old_servers = {
e["id"]: e
for e in (old or {}).get("servers", []) or []
if isinstance(e, dict) and isinstance(e.get("id"), str) and e.get("id")
}
new_servers = {
e["id"]: e
for e in (new or {}).get("servers", []) or []
if isinstance(e, dict) and isinstance(e.get("id"), str) and e.get("id")
}
changes: list[EntryChange] = []
for entry_id in sorted(set(old_servers) | set(new_servers)):
old_e = old_servers.get(entry_id)
new_e = new_servers.get(entry_id)
if old_e is None:
changes.append(EntryChange(entry_id, "added", None, new_e, ()))
elif new_e is None:
changes.append(EntryChange(entry_id, "removed", old_e, None, ()))
elif old_e != new_e:
fc = _entry_field_changes(old_e, new_e)
if fc:
changes.append(EntryChange(entry_id, "changed", old_e, new_e, fc))
return changes
# --------------------------------------------------------------------------- #
# Risk annotations -- each predicate is pure and independently unit-tested.
# --------------------------------------------------------------------------- #
@dataclass(frozen=True)
class RiskFinding:
severity: str # "blocking" | "warning" | "info"
code: str
message: str
def _escape_non_ascii(s: str) -> str:
"""Render a string with any non-ASCII code point shown as an escape
sequence, so a homoglyph/RTL-override character can't visually pass as
the real thing in the review UI."""
return s.encode("unicode_escape").decode("ascii")
def risk_env_required(change: EntryChange) -> list[RiskFinding]:
"""A non-empty env_required value is blocking: catalog entries must ship
only the *names* of env vars the user fills in, never values."""
entry = change.new or {}
env_required = entry.get("env_required")
findings: list[RiskFinding] = []
if isinstance(env_required, dict):
for k, v in env_required.items():
if v not in (None, ""):
findings.append(
RiskFinding(
"blocking",
"env_required_value",
f"env_required[{k!r}] carries a non-empty value -- catalog "
"entries must never ship secret values, only placeholder names.",
)
)
return findings
def risk_command_allowlist(change: EntryChange) -> list[RiskFinding]:
"""A command outside bcc_core.CATALOG_ALLOWED_COMMANDS is blocking.
Imports the allowlist rather than redefining it."""
entry = change.new or {}
config = entry.get("config") or {}
command = config.get("command")
if isinstance(command, str) and command and command not in CATALOG_ALLOWED_COMMANDS:
return [
RiskFinding(
"blocking",
"command_not_allowed",
f"command {command!r} is not on the catalog allowlist "
f"({', '.join(sorted(CATALOG_ALLOWED_COMMANDS))}).",
)
]
return []
def risk_non_ascii(change: EntryChange) -> list[RiskFinding]:
"""Non-ASCII code points in id/command/args are blocking -- homoglyph /
RTL-override typosquatting can make a malicious package name visually
identical to a legitimate one in a naive diff view."""
entry = change.new or {}
findings: list[RiskFinding] = []
entry_id = entry.get("id")
if isinstance(entry_id, str) and not entry_id.isascii():
findings.append(
RiskFinding(
"blocking",
"non_ascii_id",
f"id contains non-ASCII code points: {_escape_non_ascii(entry_id)!r}",
)
)
config = entry.get("config") or {}
command = config.get("command")
if isinstance(command, str) and not command.isascii():
findings.append(
RiskFinding(
"blocking",
"non_ascii_command",
f"command contains non-ASCII code points: {_escape_non_ascii(command)!r}",
)
)
for a in config.get("args") or []:
if isinstance(a, str) and not a.isascii():
findings.append(
RiskFinding(
"blocking",
"non_ascii_arg",
f"arg contains non-ASCII code points: {_escape_non_ascii(a)!r}",
)
)
return findings
def _npm_candidate_args(command: str | None, args: list[str]) -> list[str]:
if command != "npx":
return []
return [a for a in args if isinstance(a, str) and a and not a.startswith("-")]
def split_npm_spec(spec: str) -> tuple[str, str | None]:
"""Split an npm package spec into (name, version). version is None if
unpinned. Handles scoped (@scope/name@version) and unscoped
(name@version) specs."""
if spec.startswith("@"):
rest = spec[1:]
if "/" not in rest:
return spec, None # malformed scope, can't tell -- treat unpinned
scope, _, remainder = rest.partition("/")
if "@" in remainder:
pkg_name, _, version = remainder.partition("@")
return f"@{scope}/{pkg_name}", (version or None)
return f"@{scope}/{remainder}", None
if "@" in spec:
name, _, version = spec.partition("@")
return name, (version or None)
return spec, None
def is_pinned_npm_spec(spec: str) -> bool:
_name, version = split_npm_spec(spec)
return bool(version)
_DOCKER_VALUE_FLAGS = {
"-e",
"--env",
"-v",
"--volume",
"-p",
"--publish",
"-w",
"--workdir",
"-u",
"--user",
"--name",
"--network",
"--entrypoint",
}
def docker_image_candidates(args: list[str]) -> list[str]:
"""Best-effort extraction of the image reference from a `docker run
[OPTIONS] IMAGE [CMD...]` args list: the first positional token after
any leading `run` and flag(+value) pairs."""
candidates: list[str] = []
i = 0
while i < len(args):
a = args[i]
if a == "run":
i += 1
continue
if isinstance(a, str) and a.startswith("-"):
if "=" not in a and a in _DOCKER_VALUE_FLAGS:
i += 2
continue
i += 1
continue
if isinstance(a, str):
candidates.append(a)
break # first positional token after `run` is the image ref
return candidates
def is_pinned_docker_image(image: str) -> bool:
if "@sha256:" in image:
return True
tag_part = image.rsplit("/", 1)[-1]
if ":" not in tag_part:
return False # no tag => implicit :latest
tag = tag_part.rsplit(":", 1)[-1]
return bool(tag) and tag != "latest"
def risk_unpinned_package(change: EntryChange) -> list[RiskFinding]:
"""Every entry must pin an exact version: an `@scope/pkg` npm arg with
no `@version`, or a docker image with no tag / `:latest`, is blocking.
A later-compromised package must not be able to auto-upgrade into every
user just because the catalog entry never pinned a version."""
entry = change.new or {}
config = entry.get("config") or {}
command = config.get("command")
args = config.get("args") or []
findings: list[RiskFinding] = []
if command == "npx":
for a in _npm_candidate_args(command, args):
if not is_pinned_npm_spec(a):
findings.append(
RiskFinding(
"blocking",
"unpinned_npm_package",
f"npm package arg {a!r} has no pinned @version.",
)
)
elif command == "docker":
for img in docker_image_candidates(args):
if not is_pinned_docker_image(img):
findings.append(
RiskFinding(
"blocking",
"unpinned_docker_image",
f"docker image {img!r} is not pinned to an exact tag "
"(uses :latest or no tag).",
)
)
return findings
_URL_FIELDS = ("homepage", "docs_url", "source")
def _domain(url: str) -> str:
try:
return urlsplit(url).netloc.lower()
except ValueError:
return ""
def risk_url_domain_change(change: EntryChange) -> list[RiskFinding]:
"""Non-https URLs and, more importantly, a *domain change* on any URL
field are surfaced loudly with old-vs-new domains broken out -- the
lookalike-domain-swap defence."""
findings: list[RiskFinding] = []
old_entry = change.old or {}
new_entry = change.new or {}
for f in _URL_FIELDS:
new_url = new_entry.get(f)
if not isinstance(new_url, str) or not new_url:
continue
if not new_url.startswith("https://"):
findings.append(
RiskFinding("warning", "non_https_url", f"{f} is not https://: {new_url!r}")
)
old_url = old_entry.get(f)
if isinstance(old_url, str) and old_url:
old_domain, new_domain = _domain(old_url), _domain(new_url)
if old_domain and new_domain and old_domain != new_domain:
findings.append(
RiskFinding(
"warning",
"domain_changed",
f"{f} domain changed from {old_domain!r} to {new_domain!r} -- "
"verify this isn't a lookalike-domain swap.",
)
)
return findings
def risk_new_entry(change: EntryChange) -> list[RiskFinding]:
"""A brand-new entry is flagged for extra scrutiny -- not blocking on its
own, but it's the category of change the registry lookup exists for."""
if change.status == "added":
return [
RiskFinding(
"info",
"new_entry",
"Brand-new catalog entry -- extra scrutiny: check publisher identity "
"via the registry lookup before signing.",
)
]
return []
_RISK_PREDICATES: tuple[Callable[[EntryChange], list[RiskFinding]], ...] = (
risk_env_required,
risk_command_allowlist,
risk_non_ascii,
risk_unpinned_package,
risk_url_domain_change,
risk_new_entry,
)
def entry_risk_findings(change: EntryChange) -> list[RiskFinding]:
"""Run every risk predicate against one entry change and return the
combined findings (order matches _RISK_PREDICATES)."""
findings: list[RiskFinding] = []
for predicate in _RISK_PREDICATES:
findings.extend(predicate(change))
return findings
def has_blocking_risk(change: EntryChange) -> bool:
return any(f.severity == "blocking" for f in entry_risk_findings(change))
# --------------------------------------------------------------------------- #
# Review session: acknowledge-gating + TOCTOU blob-SHA pinning
# --------------------------------------------------------------------------- #
@dataclass
class ReviewSession:
"""State for one review pass. `pinned_blob_sha` is the git blob SHA of
data/catalog.json as it existed the moment review began -- see
can_sign()/sign_precondition().
`loaded_ref` is the exact ref this review was loaded from ("main", or a
PR's `refs/pull/<n>/head`) -- see issue #68 finding 1. It exists so the
Sign path can re-resolve the TOCTOU blob SHA from *the ref that was
actually reviewed*, instead of a hardcoded "main" that silently diverges
from the reviewed ref on every PR review (the bug that made the PR path
unable to sign at all, and forced everyone onto the vacuous
main-vs-itself path instead).
"""
pinned_blob_sha: str
old_catalog: dict
new_catalog: dict
loaded_ref: str = "main"
changes: list[EntryChange] = field(default_factory=list)
acknowledged: set[str] = field(default_factory=set)
def __post_init__(self) -> None:
if not self.changes:
self.changes = diff_catalogs(self.old_catalog, self.new_catalog)
def start_review(
pinned_blob_sha: str,
old_catalog: dict,
new_catalog: dict,
loaded_ref: str = "main",
) -> ReviewSession:
return ReviewSession(
pinned_blob_sha=pinned_blob_sha,
old_catalog=old_catalog,
new_catalog=new_catalog,
loaded_ref=loaded_ref,
)
def find_last_signed_catalog_raw(
candidates: list[bytes], sig: bytes, pubkeys: list[bytes]
) -> bytes | None:
"""Given `candidates` (candidate raw catalog.json byte-strings -- e.g.
successive historical versions from git log, most-recent-first),
return the first one whose signature verifies against `sig`/`pubkeys`,
or None if none do.
This is how source="main" review diffs against "the last catalog a
maintainer actually signed" instead of against itself (issue #68
finding 1): `catalog_console.last_signed_catalog_raw` walks
data/catalog.json's git history on main and hands the candidates here.
"""
for raw in candidates:
if _verify_catalog_signature(raw, sig, pubkeys):
return raw
return None
def acknowledge_entry(session: ReviewSession, entry_id: str) -> None:
ids = {c.entry_id for c in session.changes}
if entry_id not in ids:
raise ValueError(f"{entry_id!r} is not part of this review session's diff.")
session.acknowledged.add(entry_id)
def unacknowledge_entry(session: ReviewSession, entry_id: str) -> None:
session.acknowledged.discard(entry_id)
def all_entries_acknowledged(session: ReviewSession) -> bool:
return {c.entry_id for c in session.changes} <= session.acknowledged
# NOTE for future editors: do NOT add an "acknowledge all" shortcut here, now
# or ever. The friction of individually acknowledging every changed entry is
# the entire point of this tool (issue #62) -- a shortcut would let a tired
# reviewer rubber-stamp a diff exactly like the "merge PR, run script, push"
# reflex this Console exists to replace. If this comment is the only thing
# stopping you, that is the point: it is stopping you on purpose.
@dataclass(frozen=True)
class SignDecision:
ok: bool
reason: str | None = None
def can_sign(session: ReviewSession, current_blob_sha: str) -> SignDecision:
"""Whether the Sign button may fire right now.
Four independent gates, all required, checked in this order:
0. The diff must be non-empty. An empty diff historically meant "Sign
unlocks instantly" (`set() <= set()` is vacuously True), which is
exactly backwards: a vacuously-satisfied gate is worse than no gate
at all, because it *manufactures confidence* -- the signature looks
identical to one produced by a real review. "Nothing changed" must
mean "nothing to sign", never "sign unlocked". (Issue #68 finding 1;
this is what let commit b08cf21 sign all 19 entries with zero of them
ever reviewed.)
1. TOCTOU: `current_blob_sha` (fetched fresh, immediately before signing,
from the ref that was actually reviewed -- see sign_precondition())
must match the blob SHA pinned when review began. If the bytes on the
remote changed since -- a new commit pushed to the same PR, a
force-push, another PR merged in between -- signing is refused and a
re-review is forced. This is what makes "signing is the approval act"
true rather than aspirational: the signature is bound to the exact
reviewed bytes, not to "whatever the file happens to be now".
2. No blocking risk finding may be outstanding on ANY changed entry, full
stop -- checked here, not just in the GUI. The GUI additionally
disables the acknowledge checkbox for a blocking entry, but that is a
UI nicety, not the enforcement point: if this pure gate didn't also
check it, a blocking risk would only be stopped by the GUI happening
to have wired the checkbox correctly, and nothing would catch a
regression in that wiring. The GUI must not be the only thing
standing between a blocking risk and a signature.
3. Every changed entry in the diff must be individually acknowledged.
"""
if not session.changes:
return SignDecision(
False,
"Nothing to sign: this review's diff is empty. If you expected "
"changes here, you may be diffing the wrong source/ref.",
)
if current_blob_sha != session.pinned_blob_sha:
return SignDecision(
False,
"The reviewed bytes changed since this review began (blob SHA "
"mismatch) -- re-review required before signing.",
)
blocking_ids = sorted({c.entry_id for c in session.changes if has_blocking_risk(c)})
if blocking_ids:
return SignDecision(
False,
"Blocking risk finding(s) outstanding on: "
f"{', '.join(blocking_ids)} -- fix the underlying change, do not sign around it.",
)
if not all_entries_acknowledged(session):
pending = sorted({c.entry_id for c in session.changes} - session.acknowledged)
return SignDecision(
False, f"Not every changed entry has been acknowledged yet: {', '.join(pending)}"
)
return SignDecision(True, None)
def sign_precondition(
session: ReviewSession, resolve_blob_sha: Callable[[str], str]
) -> SignDecision:
"""The real Sign-button gate: resolves the current TOCTOU blob SHA from
*the ref this session was actually loaded from* (`session.loaded_ref`),
never a hardcoded "main", then delegates to can_sign().
`resolve_blob_sha` is injected so this stays testable without git/Qt --
catalog_console.ReviewWindow._on_sign passes a real resolver
(fetch_ref + blob_sha_at against self.repo_dir); tests pass a fake
dict-backed lookup. This is the fix for issue #68 finding 1's first bug:
`_on_sign` used to hardcode `fetch_ref(self.repo_dir, "main")` as the
comparison ref, so for any PR review (where `loaded_ref` is the PR's
head, not main) the SHAs differed by definition and Sign could never
fire -- and the retry path re-called the same hardcoded resolver, so it
re-pinned the same wrong value and looped forever instead of forcing a
genuine re-review.
"""
current_blob_sha = resolve_blob_sha(session.loaded_ref)
return can_sign(session, current_blob_sha)
def catalog_signing_message(raw_bytes: bytes) -> bytes:
"""The exact bytes that get signed: bcc_core's domain-separation prefix
(imported, never retyped) + the raw catalog bytes. Using this function
guarantees the Console's signature and bcc_core.verify_catalog_signature
can never drift apart on the prefix."""
return CATALOG_SIG_DOMAIN + raw_bytes
# --------------------------------------------------------------------------- #
# Key management: passphrase-encrypted-at-rest Ed25519 seed
#
# The private key is NEVER stored plaintext, never an env var, never
# committed. encrypt_private_key/decrypt_private_key are pure and offline
# (scrypt KDF + AES-256-GCM via `cryptography`, already a project
# dependency); catalog_console.py decides WHERE the resulting blob lives
# (OS keychain if available, else a file outside the repo).
# --------------------------------------------------------------------------- #
_KDF_SALT_LEN = 16
_KDF_N = 2**15 # scrypt cost parameter, tuned for a one-off interactive unlock
_KDF_R = 8
_KDF_P = 1
_NONCE_LEN = 12
_AAD = b"bcc-catalog-console-key-v1"
def _derive_key(passphrase: str, salt: bytes) -> bytes:
from cryptography.hazmat.primitives.kdf.scrypt import Scrypt
kdf = Scrypt(salt=salt, length=32, n=_KDF_N, r=_KDF_R, p=_KDF_P)
return kdf.derive(passphrase.encode("utf-8"))
def generate_keypair() -> tuple[bytes, bytes]:
"""Generate a new Ed25519 keypair. Returns (seed_32_bytes, pubkey_32_bytes)."""
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey
from cryptography.hazmat.primitives.serialization import Encoding, PublicFormat
private_key = Ed25519PrivateKey.generate()
seed = private_key.private_bytes_raw()
pubkey = private_key.public_key().public_bytes(Encoding.Raw, PublicFormat.Raw)
return seed, pubkey
def encrypt_private_key(seed: bytes, passphrase: str) -> bytes:
"""Encrypt a 32-byte Ed25519 seed at rest with a passphrase. Returns a
self-contained blob: salt || nonce || ciphertext+tag."""
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
if len(seed) != 32:
raise ValueError(f"expected a 32-byte raw Ed25519 seed, got {len(seed)} bytes")
if not passphrase:
raise ValueError("a non-empty passphrase is required")
salt = os.urandom(_KDF_SALT_LEN)
key = _derive_key(passphrase, salt)
nonce = os.urandom(_NONCE_LEN)
ciphertext = AESGCM(key).encrypt(nonce, seed, _AAD)
return salt + nonce + ciphertext
def decrypt_private_key(blob: bytes, passphrase: str) -> bytes:
"""Decrypt a blob produced by encrypt_private_key. Raises ValueError on a
wrong passphrase or corrupt blob -- never silently returns garbage."""
from cryptography.exceptions import InvalidTag
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
if len(blob) < _KDF_SALT_LEN + _NONCE_LEN:
raise ValueError("key blob is too short to be valid")
salt = blob[:_KDF_SALT_LEN]
nonce = blob[_KDF_SALT_LEN : _KDF_SALT_LEN + _NONCE_LEN]
ciphertext = blob[_KDF_SALT_LEN + _NONCE_LEN :]
key = _derive_key(passphrase, salt)
try:
return AESGCM(key).decrypt(nonce, ciphertext, _AAD)
except InvalidTag as e:
raise ValueError("wrong passphrase or corrupted key file") from e
def sign_catalog_bytes(raw: bytes, seed: bytes) -> bytes:
"""Sign `raw` catalog bytes with a 32-byte Ed25519 seed, using the exact
domain-separated message bcc_core.verify_catalog_signature expects."""
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey
if len(seed) != 32:
raise ValueError(f"expected a 32-byte raw Ed25519 seed, got {len(seed)} bytes")
private_key = Ed25519PrivateKey.from_private_bytes(seed)
return private_key.sign(catalog_signing_message(raw))
# --------------------------------------------------------------------------- #
# Registry lookup -- the check a human genuinely can't do.
#
# The network call itself is injected (a `Fetcher` callable) so this stays
# testable offline; catalog_console.py supplies the real npm/PyPI HTTP
# fetcher. Fails soft everywhere: a fetcher returning None/raising just
# yields RegistryInfo(available=False), never an exception into the caller
# and never a block on review.
# --------------------------------------------------------------------------- #
@dataclass(frozen=True)
class PackageRef:
entry_id: str
ecosystem: str # "npm" | "pypi"
name: str
version: str | None
def split_pypi_spec(spec: str) -> tuple[str, str | None]:
for sep in ("==", "@"):
if sep in spec:
name, _, version = spec.partition(sep)
return name, (version or None)
return spec, None
def extract_package_refs(entry: dict) -> list[PackageRef]:
"""Pull out the package(s) a basic-tier entry's args reference, for the
registry lookup. Returns [] for link-only entries or entries whose
command isn't npx/uvx (docker images aren't registry-lookup candidates
in the npm/PyPI sense used here)."""
config = entry.get("config") or {}
command = config.get("command")
args = config.get("args") or []
entry_id = entry.get("id", "") if isinstance(entry.get("id"), str) else ""
refs: list[PackageRef] = []
if command == "npx":
for a in _npm_candidate_args(command, args):
name, version = split_npm_spec(a)
if name:
refs.append(PackageRef(entry_id, "npm", name, version))
elif command == "uvx":
for a in args:
if isinstance(a, str) and a and not a.startswith("-"):
name, version = split_pypi_spec(a)
if name:
refs.append(PackageRef(entry_id, "pypi", name, version))
break # `uvx <pkg>` -- first positional token is the package
return refs
@dataclass(frozen=True)
class RegistryInfo:
ref: PackageRef
available: bool
publisher: str | None = None
age_days: int | None = None
last_release: str | None = None
downloads: int | None = None
near_neighbor_ids: tuple[str, ...] = ()
Fetcher = Callable[[PackageRef], dict | None]
def edit_distance(a: str, b: str) -> int:
"""Levenshtein distance, iterative DP (no recursion depth concerns)."""
if a == b:
return 0
la, lb = len(a), len(b)
if la == 0:
return lb
if lb == 0:
return la
prev = list(range(lb + 1))
for i, ca in enumerate(a, 1):
cur = [i] + [0] * lb
for j, cb in enumerate(b, 1):
cost = 0 if ca == cb else 1
cur[j] = min(prev[j] + 1, cur[j - 1] + 1, prev[j - 1] + cost)
prev = cur
return prev[lb]
def near_neighbor_ids(name: str, other_ids: list[str], max_distance: int = 2) -> list[str]:
"""Catalog ids within `max_distance` edits of `name` (case-insensitive),
excluding an exact match -- the dependency-confusion / typosquat
near-neighbour warning."""
lname = name.lower()
return [
oid
for oid in other_ids
if oid != name and edit_distance(lname, oid.lower()) <= max_distance
]
def lookup_registry_info(
ref: PackageRef, fetcher: Fetcher, all_entry_ids: list[str]
) -> RegistryInfo:
"""Resolve one package against the live registry via the injected
fetcher. Never raises: any fetcher exception or falsy return means
`available=False` ("unavailable"), which the GUI renders plainly rather
than blocking or erroring the review."""
neighbors = tuple(near_neighbor_ids(ref.name, all_entry_ids))
try:
raw = fetcher(ref)
except Exception:
raw = None
if not raw:
return RegistryInfo(ref=ref, available=False, near_neighbor_ids=neighbors)
return RegistryInfo(
ref=ref,
available=True,
publisher=raw.get("publisher"),
age_days=raw.get("age_days"),
last_release=raw.get("last_release"),
downloads=raw.get("downloads"),
near_neighbor_ids=neighbors,
)
_NON_ASCII_RE = re.compile(r"[^\x00-\x7f]")
def contains_non_ascii(s: str) -> bool:
return bool(_NON_ASCII_RE.search(s))