Watch
1
0
Fork
You've already forked SouveraineOS
0
SouveraineOS/docs/substrate/NET_LAYER_PLAN.md
Fimeg e480809c70 docs: rescue the agent-substrate tree out of a gitignored directory
219 files, 2.0 MB, untracked in souveraine/docs and existing nowhere else.
The volume is at 100% with no snapshots.
2026-07-26 12:11:50 -04:00

382 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# SouveraineOS Networking Layer
Status: design, under discussion. Not yet building.
Target: the product, not a v1. May not compile on first run; that's fine.
## The product
A machine — laptop or phone — where every consumer of the network (the
substrate's own federation/sync/voice sockets; system services; a Waydroid
Android instance and its apps; an arbitrary shelled-out process) is governed
by an explicit per-consumer network policy, enforced at the right kernel/
userspace layer for what it is. Default baseline is "behave like a normal
Linux host"; policy is an override. The user and the agent both express it;
the substrate enforces it.
Mobile Linux has not solved this. We are.
## What's live on the device (verified 2026-07-17)
Two data links, NM-tracked, correct metrics:
- `wlan0` wifi — IPv4, default route metric 600
- `qrtr0` gsm "Fido LTE" — IPv6-only, global v6 on `qmapmux0.0`, no native v4
- `clat` tun — 464XLAT IPv4-over-IPv6, MTU 1260
NetworkManager + ModemManager active and correct. Kernel prefers WiFi via
metric — we READ priority, never implement it.
### Kernel facts that drove the design (not assumptions)
- **cgroup v2 only** (`cgroup2fs`). Mounted controllers: cpuset cpu io
memory hugetlb pids. **No `net_cls`, no `net_prio`** — those are v1-only,
deliberately absent in v2.
- `nft`, `iptables`, `ip` present. No libcgroup userspace (not needed on v2).
- **Conclusion:** the textbook "net_cls.classid + nftables match" recipe is
dead on this kernel. Per-process classification must be **cgroup-bpf**
(`CGROUP_SKB` / `CGROUP_SOCK_ADDR`) stamping a socket mark, with `ip rule`
+ per-link routing tables (and/or nft mark-match) doing the enforcement.
- Waydroid not yet installed — it's future alongside-software, governed as a
classified consumer class, not integrated via binder.
## Architecture — two regimes, one model
```
Consumer (who)
│ has a LinkPolicy (which links, with what constraints)
│ enforced via an Enforcement (how)
┌───────────────────────────────┬───────────────────────────────┐
│ Substrate-socket regime │ Kernel-classification regime │
│ (userspace, no privilege) │ (privileged: cgroup-bpf+nft) │
│ │ │
│ federation WS │ Waydroid instance + its apps │
│ memfs git sync │ system services (systemd units)│
│ voice STT/TTS HTTP │ arbitrary shelled-out procs │
│ any socket WE open │ any socket we DIDN'T open │
│ │ │
│ SO_BINDTODEVICE / bind() │ cgroup dir + pid write + │
│ before connect() │ BPF mark + ip rule per link │
│ + re-resolve on handoff │ + nft allow/deny/mark-match │
└───────────────────────────────┴───────────────────────────────┘
LinkState (authoritative, from NM/MM D-Bus)
owned by machined::net, same as before
```
Why two regimes, not a choice: the substrate's own connections NEED
userspace control — only the substrate can re-resolve peer DNS on egress
change and reconnect its WS with the right semantics; a kernel policy layer
can't express that. Arbitrary processes NEED kernel classification — only
the kernel sees sockets the substrate never opened. These aren't competing
approaches; they're the correct tools for two disjoint categories of traffic,
unified by one Consumer/LinkPolicy model.
## The shared model (the spine)
```rust
/// A physical or virtual link, as the kernel + NM know it.
pub struct Link {
pub device: String, // "wlan0", "qrtr0", "clat"
pub kind: LinkKind, // Wifi | Cellular | Wired | Tun
pub metered: bool,
pub v4: bool,
pub v6: bool,
pub default: bool, // holds the lowest-metric default route
pub source_v4: Option<Ipv4Addr>, // bind address for SO_BINDTODEVICE fallback
pub source_v6: Option<Ipv6Addr>,
pub signal: Option<u8>,
pub roaming: bool, // MM AccessTechnologies → guess
}
/// Which links a consumer may use, and how. The primary type.
pub struct LinkPolicy {
pub allow: LinkSet, // explicit allow-list of link constraints
pub deny: LinkSet, // explicit deny-list (wins over allow)
pub max_metered: Option<DataBudget>, // cap cellular bytes per period
}
pub enum LinkSet {
/// Match by capability, not by device name — survives interface rename.
AnyMatching { kind: Option<LinkKind>, metered: Option<bool>,
v4: Option<bool>, v6: Option<bool>, roaming: Option<bool> },
/// Pin to a named device (rare; for diagnostics).
Device(String),
/// The kernel default route — the "no policy" baseline.
DefaultRoute,
}
/// Who the policy applies to.
pub struct Consumer {
pub id: ConsumerId,
pub policy: LinkPolicy,
pub enforcement: Enforcement,
}
pub enum ConsumerId {
/// A named substrate component we author.
Component(&'static str), // "federation", "memfs-sync", "voice-stt"
/// A cgroup path — matches any process placed there.
Cgroup(String), // "/souveraine/waydroid", "/souveraine/web"
/// A systemd unit we'll enroll.
Unit(String),
}
pub enum Enforcement {
/// Substrate opens the socket itself and binds it. No privilege.
SubstrateSocket,
/// Kernel enforces via cgroup-bpf mark + ip rule + nft. Privileged.
KernelClassify { cgroup: String, fwmark: u32 },
}
```
`LinkClass` (Any/Validated/Unmetered/Bulk) survives as a **convenience
derived from LinkPolicy + live LinkState** — it's how a call-site asks "does
my policy permit a Bulk operation right now" without re-stating constraints.
It is NOT the primary type; the primary type is the explicit LinkPolicy.
## Tier 1 — machined::net (link truth)
Unchanged from the earlier sketch: NM/MM D-Bus reads via the existing zbus
dep, builds `LinkState`, serves it on the guarded ok/reason socket as
`NetState`. std-only, synchronous, audited, SO_PEERCRED-logged — same
discipline as the rest of machined. LinkState is the input to everything.
Add: a `LinkPolicy` store (where per-consumer policies live) also served via
machined, since policy is machine-level truth, not agent-level.
## Tier 2 — core/net (the policy engine)
In-process. Holds the current LinkState (from machined, or direct NM D-Bus
on desktop). The engine:
1. Resolves a Consumer's LinkPolicy against live LinkState → the set of
currently-usable Links (or: "deny, no matching link" / "deferred, only
metered available and budget exhausted").
2. Exposes `tokio::sync::watch` of (LinkState, resolved-policies) so
consumers react to handoff.
3. Provides the substrate-socket bind helper:
`bind_socket_to(link: &Link)` — SO_BINDTODEVICE, falling back to
bind(source_addr) — used by federation/sync/voice before connect().
4. Provides `enroll(consumer)` for KernelClassify consumers: create the
cgroup dir, attach the BPF program, install the ip rules + nft rules.
## Tier 3 — the kernel-classification regime (the hard, real part)
This is what makes it a product instead of a library.
### Mechanism (cgroup v2 + BPF, since net_cls is gone)
1. **cgroup layout** under `/sys/fs/cgroup/souveraine.net/<consumer>/`.
Each KernelClassify consumer gets a dir; processes are enrolled by
writing their pid to `cgroup.procs`. Waydroid's container init goes into
`/souveraine.net/waydroid/`; its apps inherit.
2. **BPF classifier**: a `BPF_PROG_TYPE_CGROUP_SKB` (egress) and/or
`CGROUP_SOCK_ADDR` (connect) program attached to each consumer cgroup.
It stamps the socket's skb / socket with a fwmark identifying the
consumer + its current policy verdict. Verdict recomputed on LinkState
change (userspace pushes a new map; BPF reads the map).
3. **Routing**: per-link routing tables (wifi in table 100, lte in table 200,
clat in table 300) + `ip rule` entries keyed on fwmark → table. A socket
marked "wifi-only" routes through table 100 regardless of the default.
A consumer with no policy gets the default route (baseline behavior).
4. **Allow/deny**: nftables matches fwmark + egress interface → accept or
drop. A wifi-only consumer attempting to egress on qrtr0 is dropped.
5. **Metered budget**: a BPF map counter per consumer; once the period
budget is hit, the verdict flips to deny-cellular. Reset by userspace
on the period boundary.
### Privilege + trust model (inherits RedFlag's binary-trust pattern)
This is the load-bearing security design, lifted directly from RedFlag's
`redflag-helper` (capability-token executor, verified in helper/src/main.rs).
**The standing-daemon question is resolved: netd is NOT standing.** The BPF
programs are long-lived *kernel* state, pinned to `/sys/fs/bpf/souveraine/`
(bpffs is already mounted, mode 700). They survive netd exiting. So the
privileged *process* and the privileged *kernel state* are decoupled — a
standing daemon is not needed to hold the programs, and RedFlag's "no standing
elevated daemon" doctrine holds.
**`souveraine-netd` — transient, per-operation, capability-token executor.**
Modeled 1:1 on `redflag-helper`:
- Invoked via one `systemd-run --wait ... souveraine-netd --token-file ...`
line (the agent's/machined's *only* sudo for net ops). Scoped by polkit to
transient units, exactly as RedFlag does.
- Reads one Ed25519-signed capability token on stdin, performs exactly one
mutation, exits. No shell, no inherited env, fail-closed on every path.
- Exit codes double as deny taxonomy (copy RedFlag's 1026 scheme).
- Token operations (fixed vocabulary — this is the "models won't rewrite
code" guarantee):
`attach-program`, `set-policy <consumer>`, `enroll-cgroup <pid>`,
`set-ip-rule`, `set-metered-budget`, `detach-program`.
**Signing authority (decided 2026-07-17): machined signs netd tokens.**
machined already holds the machine Ed25519 seed and the
`souveraine-machined:v1` signing context. netd tokens are signed by machined
under a new domain `netd-policy`. One root of trust per machine, already
provisioned and audited (SO_PEERCRED, guarded ok/reason). No second key, no
off-host authority. The agent never signs; it step-up-requests, machined
signs on policy authority, netd verifies against the machine pubkey.
**The agent step-up path** (the product surface you want):
```
agent decides "consumer X should be wifi-only"
→ step-up IPC to machined: {operation: set-policy, consumer: X,
policy: {allow: wifi-only}, reason: ...}
→ machined verifies agent's step-up auth (ambient/personal/stepUp tiers,
same machinery as the session capability work — gated, acked, verified)
→ machined signs netd-policy capability token:
closure_hash = sha256(canonical policy mutation)
signed = machined.sign("souveraine-machined:v1:netd-policy:{token}")
→ machined invokes netd via the one systemd-run line, token on stdin
→ netd: verify token vs machine pubkey → validate trust paths →
hash-check → perform the one set-policy → write verdict map → exit
→ result back to machined → back to agent
```
The verbs (`set-policy`, `enroll-cgroup`, ...) and the policy schema are
fixed. A model expresses intent over this stable vocabulary; it can never
ask netd to do something outside the enumerated operations, because netd
refuses unknown ops (EXIT_UNSUPPORTED_OP). Extending the engine adds verbs;
it never requires the model to emit different logic for existing ones.
**Inherited RedFlag mechanisms (all verified in helper/src/main.rs):**
1. **Pinned keyring, verify-keys-not-servers** — netd trusts the machine
pubkey by fingerprint. `/etc/souveraine/trusted-keys` (or derive from
machined's seed pub). No URL trust.
2. **Trust-path self-validation (SEC-021)** — every file netd relies on
(keyring, the BPF ELF being attached, policy files) must be root-owned,
not a symlink, not group/other-writable. Violation = hard denial
(EXIT_TRUST_PATH). Defends its own inputs.
3. **Hash-pinned BPF objects** — the BPF program ELF is part of the signed
closure; netd re-computes sha256 constant-time before attach. A swapped
`.o` is denied. `stage_and_verify` into root-only staging before verify,
so the caller can't swap bytes mid-flight.
4. **Atomic replace** — programs/maps pinned atomically; failed attach leaves
no half-state (rename, never write-in-place).
5. **Replay guard** — consumed tokens recorded in netd state file.
6. **Self-upgrade via the same gate** — netd replaces itself using a
`netd-self` token verifying the new binary's hash. Trusted path never
widens for updates.
**Reactivity without a standing daemon (the key trick):**
Link-change reactions (flip the verdict map when wifi drops) are handled by
a **small unprivileged watcher** — NM-D-Bus / netlink listener with write
access only to the bpffs map file, *not* to netd's privilege and *not* to
CAP_NET_ADMIN. bpffs map permissions gate who may write the verdict. So the
privileged surface stays transient (netd runs per-op, exits), and the
reactive surface is unprivileged (map writes only). This is how you get live
handoff behavior without a standing elevated process.
**netd systemd unit** (inherits RedFlag's hardening template):
- `AmbientCapabilities=CAP_NET_ADMIN` ONLY. No `CapabilityBoundingSet`
(RedFlag rule: strips setuid caps from sudo inside the unit, kills the
invocation path).
- `ProtectSystem=strict`, `ProtectHome=true`, `PrivateTmp=true`,
`ProtectKernelTunables=true`, `RestrictSUIDSGID=true`.
- `ProtectControlGroups` left default — netd needs cgroup dir creation.
- Documents the kernel-version floor (see NET_KERNEL_GAPS.md) and refuses
to attach if the floor isn't met (BTF absent, mark target missing) —
fail-closed, never silently unenforced. RedFlag's floor doctrine.
### Waydroid as a consumer class
Not a binder integration. Waydroid's LXC container init is enrolled into
`/souveraine.net/waydroid/` (or a subtree with per-app subgroups if we want
per-Android-app granularity). Its sockets are classified and routed like
any other process. The Android `ConnectivityManager` inside the container
keeps working against whatever egress the kernel gives it; we govern that
egress from outside. binder stays out of scope.
## The handoff (the thing you feel leaving the house)
Two halves, both required:
1. **Substrate connections re-resolve + reconnect on link change.**
`peer_outbound_task` today pins its resolved peer across backoff. Fix:
hold endpoint as URL, re-resolve via `tokio::net::lookup_host` on every
reconnect, subscribe to core/net's watch and force-reconnect on `primary`
change. Cellular-v6-only + CLAT case: log the CLAT ride, don't pointlessly
retry an IPv4 peer.
2. **Kernel-classified consumers follow the link set their policy allows.**
When wifi drops, a wifi-only consumer's ip-rule + nft deny drops its
egress; when wifi returns, it flows again. No userspace action needed —
the BPF verdict map flips with LinkState.
## Surfaces
- health panel: active link, metered, cellular signal, connectivity, AND the
resolved verdict per enrolled consumer ("waydroid: wifi-only (denied on
lte)"). The consumer view is the new part.
- (later) a policy UI: assign consumers to policies. Read-only first.
## Dependencies & privilege
- zbus (present) for NM/MM D-Bus.
- New: **aya** (pure-Rust BPF loader). Chosen over libbpf-rs to keep C out
of netd's privileged TCB. Requires kernel BTF (the rebuild adds it).
- New: nftables — shell `nft` first (invoked by netd per-op); bind nftnl-rs
later only if a hot path demands it. Per-op shelling fits the transient-
executor model fine.
- **netd** carries `CAP_NET_ADMIN` (transient, per-op via systemd-run),
NOT machined. machined signs tokens; netd enforces. The agent process
stays unprivileged throughout. See Privilege + trust model above.
- Cross-build on archdev: the Rust binaries cross-compile as today; BPF
programs are ELF-for-BPF (arch-independent), host-built, not cross-
compiled to aarch64.
## Decisions (resolved 2026-07-17)
1. **Daemon shape: transient executor, NOT standing.** netd runs per-
operation (one signed token, one mutation, exit), modeled on
`redflag-helper`. BPF programs persist as kernel state pinned to bpffs,
decoupled from the privileged process. Reactivity via an unprivileged
map-writer. RedFlag's "no standing elevated daemon" doctrine holds.
Detail in the Privilege + trust model section above.
2. **Loader: aya.** Pure Rust, no C libbpf in the privileged TCB (RedFlag
"small auditable privileged surface"), cross-friendly for archdev.
Requires kernel BTF — the rebuild adds it (see NET_KERNEL_GAPS.md).
3. **Signing authority: machined signs netd tokens** under domain
`netd-policy`. Single root of trust, already provisioned. Agent step-up-
requests; machined signs; netd verifies.
4. **Agent may change policy, via step-up.** Tied to the ambient/personal/
stepUp capability tiers — a policy mutation is a step-up request (gated,
acked, verified), same machinery as the session capability work.
5. **Metered budget: soft warn at N + hard deny at M** (Android
NetworkPolicyManagerService model). Honest about enforcement precision.
## Still open (genuinely, not deferred for safety)
- **Per-Android-app granularity inside Waydroid.** Whole-container works
now; per-app needs Waydroid's per-app cgroups surfaced to the host. Defer
until Waydroid is installed. The cgroup-subtree layout is designed so a
per-app subtree carves later without rework. (Separately interesting as
its own thread: running Android apps semi-natively — noted, not this
build.)
## What's deliberately NOT deferred to a "v2"
- The Consumer/LinkPolicy/Enforcement model is built whole, not phased in.
- Kernel classification (BPF + ip rule + nft) is part of this build, not a
follow-on — it's the entire reason the product is worth making.
- Per-consumer metered budget is part of this build.
- The capability-token trust model (machined-signed, RedFlag-inherited) is
part of this build, not bolted on later.
- The health surface shows consumer verdicts, not just link state.
What IS staged by necessity (not by safety): Waydroid per-app granularity
(needs Waydroid installed); the policy-editing UI (needs the engine live);
the kernel rebuild (needs the flash). These are gated on external facts,
not on us playing it safe.