dnf DryRun: --assumeno cancels the transaction, so DNF exits non-zero
even on a dry run that resolved cleanly — the old check failed those
(curl-style single-package upgrades with no extra deps showed FAILED).
But non-empty stdout is not success either: 'No match for argument',
'Nothing to do', and 'Error:' all print output and exit non-zero, and
treating them as success would mint a capability token for a transaction
that never installs (fail-open). Gate on an actually-resolved
transaction (a 'Transaction Summary' block, which never coexists with
'Nothing to do') instead.
web logout: clear the zustand persist key (auth-storage) alongside
auth_token and user, so a JWT from a prior server reinstall (JWT_SECRET
rotation) does not survive a logout + re-login cycle. Drop the redundant
localStorage removal in Layout — the store owns logout cleanup.
- Replace unbounded goroutine fan-out with bounded pool (8 concurrent)
so 300-package dnf scans no longer timeout every request against api.osv.dev
- Drop in-memory osvDedup sync.Map; gate on persisted supply_chain_checked_at
so the dedup survives restart and failed checks retry naturally
- On query failure, record the error without checked_at so the package stays
a candidate for the next cycle (ETHOS: errors are history, assume failure)
- Shared RunOSVChecks in services/ used by both scan path and startup backfill
- Add FreshSupplyChainPackages query for persist-driven freshness lookup
- Bump version to 0.2.3.0
Add server/internal/orchestrator — a stateless, DB-driven service that advances
the package update lifecycle instead of leaving every transition to an operator
button press. It holds no state of its own: each 60s sweep re-reads from the DB,
so a restart reconciles on the next tick, and every advance rides the
LIFECYCLE-001 guarded transition (idempotent by construction).
- Auto-approval is policy-gated by policy.auto_approve_max_severity (default
off). Eligible pending packages are approved and their dry-run enqueued
(-> checking_dependencies). It never advances to installing on its own, never
runs when allow_dry_runs=false, never approves packages carrying
supply_chain_vulns, and an unrecognized policy value fails safe to disabled.
- Stuck-state recovery: checking_dependencies past 30m re-enqueues the dry-run
up to twice then fails the package; installing past 60m fails it;
pending_dependencies past 24h logs a stale count without changing state.
- Synchronous fast-path: ReportUpdates fires OnPackagesDiscovered() so
auto-approve runs at scan time, serialized with the timer via TryLock.
Supporting changes: exported TransitionByID/TransitionByPackage and
GetPackagesInStatus/BumpRetryCounter on UpdateQueries; GetPolicyString on the
settings service; EnqueueDryRun extracted from the manual dry-run endpoint and
shared with the orchestrator. Unit tests cover policy gating, requeue-then-fail,
and the install timeout.
Route every current_package_state status change through one transitionStatus
path: read the observed status, validate against PackageStatusTransitions,
run a status-guarded UPDATE, record terminal history. Replaces ten raw-SQL
transition functions whose WHERE guards validated nothing and silently
no-op'd on an illegal state. ApproveUpdate, the Reject/Install/Set* family,
BulkApprove and UpdatePackageStatus now share the core; illegal moves return
a named from->to error instead of a silent miss, and concurrent callers are
caught by the guarded row count.
Migration 047 renames the terminal success state updated -> installed in
current_package_state and update_version_history, realigning both CHECK
constraints with the Go PackageStatus/HistoryStatus constants.
UpdateStats updated_updates -> installed_updates to match.
UpdateCurrentStateInTx documents its reconcile CASE as the SQL twin of
models.ReconcileFromScan so the two stay in lockstep.
Dashboard: vulnerable-package count surfaced in AttentionPanel, plus a
Vulnerable quick-filter on the Updates view.
When ErrRefreshTokenInvalid or ErrMachineMismatch fires, the agent now
waits 10 minutes between poll attempts instead of exponential backoff.
These are permanent states — no amount of retrying fixes a dead
refresh token or a machine_id mismatch. The agent stays alive and
visible, waiting for operator intervention through the dashboard.
Old key stays is_active=true after SetPrimaryKey with no TTL or auto-deprecate.
README now says what actually exists: operator promotes new key, then explicitly
deprecates the old one. No sliding-window automation yet.
RAF/SUPPLY_CHAIN_GATE_PLAN.md unblocked — the architectural thesis for the
capability-token model. Updated to reflect agent-self upgrade path, OSV
expansion, and current verification state.
README Status section rewritten: "implemented and locally exercised, not
production-proven" replaces the misleading "working in production" header.
Honest gaps listed (GATE-002, CRITICAL-004).
Trust model: refresh-token rotation paragraph now states the 90-day expiry.
OPERATIONS.md (operator runbook) whitelisted in .gitignore — useful for
anyone deploying RedFlag.
SEC-001: Registration tokens stored as SHA-256 hashes. Migration 046 adds
token_hash column, backfills from plaintext, drops token column. All queries
use hash. Token plaintext shown once at creation (reveal panel in UI), never
retrievable again. Follows the refresh-token pattern.
SEC-005: README "no sanitization" claims corrected — code correctly sanitizes
against log injection (ANSI stripping, control char replacement, truncation).
Wording updated to match reality.
SEC-008: Command creation with idempotency_key uses ON CONFLICT DO NOTHING
instead of blind insert. Prevents duplicate command execution.
Trust model: Ed25519 key rotation documented — signing_keys table supports
multiple concurrent active keys with a sliding window for zero-downtime
rotation. OSV.dev ecosystem coverage updated (apt, dnf added).
A retried command carries the same action/result as its original, so the
history read as a fresh attempt. The lineage already lived in
agent_commands.retried_from_id — it just was not projected.
GetAllUnifiedHistory now selects is_retry + retried_from_id (both UNION
halves; logs are always false/null); UnifiedHistoryItem carries them; the
handler prefixes the narrative with "Retry — ". ChatTimeline composes its
own command sentences (narrative is only a log fallback), so it gets the
same prefix guarded by entry.is_retry, matching the is_retry/
retried_from_id convention LiveOperations already consumes.
Also drop a leftover heartbeat console.log debug block in Agents.tsx.
Make the capability gate runnable as installed. The helper is now a
first-class signed artifact distributed through the same pipeline as the
agent binary: built in the server image, signed at startup, listed in the
signed release manifest, served over GET /api/v1/helper/:arch with
X-Content-Signature, and Ed25519-verified + provisioned at install time
(binary root:root 0755, keyring, replay-guard dir, agent_id).
dnf discovery runs unprivileged: SandboxOpts redirect log/cache to an
agent-writable temp dir, so the agent holds zero dnf sudo (only the helper
invocation line). Removed the dead dnf discovery sudoers grants.
dnf artifact resolution: dnf5 pulls the matching .src.rpm from COPR-style
repos alongside the binary, which made singleRPMInDir refuse as ambiguous,
dropping the closure to empty and failing the mint closed ("no resolved
closure stored"). Filter source rpms before the ambiguity check so it pins
the one install artifact.
Drop the Fedora "updates" repo from the dnf security-severity heuristic;
it is not security-specific. Delete orphaned installer/sudoers.go (no
callers; emitted a contradictory unit). Migration 036: remove embedded
BEGIN/COMMIT that closed the runner's own transaction early.
- NeedsSupplyChainCheck: unchanged, OSV.dev eligibility (npm/pypi only)
- CanServerFetchArtifact: server can download from public registries (npm/pypi)
- NeedsCapabilityGate: ecosystems that route through capability tokens (dnf, apt, npm, pypi)
- computeAndStorePackageHash now uses CanServerFetchArtifact
- usesCapabilityExecution now uses NeedsCapabilityGate
Previously, ApproveUpdate logged mint_skipped for dnf/apt because
computeAndStorePackageHash returns empty (server cannot download those
artifacts). But the agent already resolved and reported the full closure
with per-artifact hashes via ReportDependencies → pinReportedClosure.
Now when artifactHash is empty, the handler calls mintResolvedClosure
which reads the stored closure from the dry-run phase. Server-fetched
ecosystems (npm/pypi) use the existing artifactHash path unchanged.
Remove Install/InstallMultiple/Upgrade/UpdatePackage from the Installer
interface. Mutation for gated ecosystems (dnf/apt) is refused with a
[SECURITY] error directing to the capability-token path. Non-gated
ecosystems (winget, windows_update, docker_image) type-assert to the
concrete type for UpdatePackage/Upgrade/InstallMultiple.
HandleInstallUpdates: gate dnf/apt; type-assert switch for non-gated.
HandleConfirmDependencies: gate dnf/apt; type-assert switch with
InstallMultiple for dependency batches.
Replace raw exec.Command calls in the scanner and the SecureCommandExecutor
calls in DryRun with DiscoveryRunner, matching the pattern already applied
to DNF (5a27f7b0). Mutation methods (Install, InstallMultiple, Upgrade,
UpdatePackage) are unchanged — Task 4 will delete them.
Replace raw exec.Command (scanner) and SecureCommandExecutor (installer dry-run)
with the unified DiscoveryRunner chokepoint. DryRun no longer manages its own
temp dir — DiscoveryRunner handles sandbox compatibility per ecosystem config.
Bind /renew to the registered machine so a stolen refresh token can't mint
tokens from another host. Rotate the refresh token on every renewal; replaying
a consumed token whose successor is also consumed revokes the family. Accept-
previous-once grace covers agent crash-before-save. Typed auth errors so the
polling loop renews on 401 and treats refresh/machine failures as terminal.
No unsigned binary path: build refuses when signing is disabled, downloads
return 404 when no signed package resolves. Signed release manifest endpoint,
installer verifies manifest signature and pins binary hash, Windows token
mandatory, Rust helper verify-binary.
Extract the agent polling loop from main.go into internal/agent/loop.go
so the Windows service and the CLI agent share one code path. The loop
now reads jitter cap and backoff curve from PollingConfig (struct with
merge + file/env defaults) instead of hardcoding 30s/10s/300s. Machine
ID resolution uses the canonical system.GetMachineID() in both the
registration and runtime paths, removing the inline 'unknown-' fallback.
Stuck command retries are parameterized (maxRetries arg) rather than
hardcoded to < 5.
Restructure settings into a uniform hub-of-cards pattern: extract inline
Account Settings into /settings/general, un-orphan SecuritySettings with
working /settings/security/:tab routes. Add fleet-wide polling resilience
tuning (jitter_max_seconds, backoff_base_seconds, backoff_max_seconds)
as operational settings — stored in security_settings, delivered over
GET /api/v1/agents/:id/config, merged into the agent's local config at
runtime with a 15-minute refresh cadence. Frontend includes AgentPolling
page, hook, hub card, and route.
Lands the long-dropped in-flight work plus two slices of the pinning-mirror direction.
Registry-gap closure (in-flight, was repeatedly dropped):
- Agent resolves canonical artifact hashes from its own signed repo metadata
(dnf download + rpm header; apt-cache policy+show) — server no longer serves a
placeholder dnf URL and says so honestly.
- Server pins the agent-reported closure and mints the capability token at the
dependency-confirmation boundary; receipt updates package status.
Slice 1 — package detail pane:
- GET /updates/:id/fleet (cross-agent view). Detail pane gains Supply Chain card
(pinned sha256, published/age, age-gate verdict, resolved closure) and Affected
Agents card (per-host version delta + status, click-to-pivot).
Package-centric Updates list:
- ListAggregatedPackages rollup (GET /packages): one row per package across the
fleet — agent/version counts, max severity, vuln + hash-pin rollups, status
breakdown. List view rewritten to package rows that drill into the fleet view.
Slice 2 — version timeline catalog:
- migration 043 package_versions; idempotent upsert populated at scan, enriched at
approval (OSV posture, publish date, hash) and at closure pin (per-artifact hash).
- GET /updates/:id/versions + Version Timeline card.
UI: description overflow fix, shared table density px-6->px-4, status label cleanup.
Version: 0.2.0.7 across versions.go, docker-compose, Makefile (Makefile was stale at
0.2.0.3/0.2.0).
Server:
- ApproveUpdate() now calls computeAndStorePackageHash() to download artifact,
compute SHA256, and store in DB
- GET /dashboard/updates/verify-hash endpoint for agents to fetch hashes
Database:
- Migration 040: added expected_sha256 VARCHAR(64) to current_package_state table
Agent:
- HandleInstallUpdates() fetches expected hash from server before install
- DNFInstaller.VerifyHash() downloads and verifies package hash
- APT/Docker/Winget/WindowsUpdate: hash verification stubs (fail-open)
- LRU cache (100 entries) to reduce server load
Security:
- Hash verification happens BEFORE package manager install
- Mismatch blocks installation with error logged
- Fail-open: hash fetch failure doesn't block, but verification failure does
- AgentUpdatesEnhanced: ['active-commands'] → ['activeCommands'] (hyphenated key
never matched the camelCase query key, so invalidation was silently dead)
- useUpdates (install + approve): invalidate ['dashboard-stats'] and
['activeCommands'] on success so Dashboard and Live Operations react without
waiting for their independent poll cycles
- Updates.tsx handleConfirmDependencies: replace window.location.reload() with
targeted queryClient.invalidateQueries calls (BUG-017 pattern)
- LiveOperations: updateId = cmd.params?.update_id || cmd.id so "View Update
Details" navigates to the correct update package, not the command record
- useUpdates query: add refetchInterval: 30000 / staleTime: 15000 so agent-side
completions surface without window-focus or manual refresh
- updates.go ReportLog: emit system_event (agent_update/failed) when an
update_agent or verify_command command returns result=failed, using
RenderUpdateLog for operator-facing narrative
The agent's validateNonce was hardcoded to 5 min, equal to the default
check-in interval. When queue-to-fetch timing drifts past expiry (e.g.
agent restarted, check-in cycle reset) the nonce expires before the agent
can fetch the command. Age+1s-late → nonce_expired → command rejected.
Changed validateNonce to accept maxAge, computed in HandleUpdateAgent as
2 × cfg.CheckInInterval. The nonce still binds the update to this agent
within a bounded window — same security model, no auto-retry magic.