Watch
1
0
Fork
You've already forked SouveraineOS
0
SouveraineOS/saf/device/modem.md
Fimeg bde961c6f2 saf: one spine — device, state, and work under the index
PAF becomes saf/device (history kept), STATE.md dissolves
into saf/state.md with the dated era archived, the substrate
SAF moves up from souveraine, and every agreement points at
saf/INDEX.md and nowhere else. one map, nothing to remember
2026-08-18 09:47:30 -04:00

837 lines
64 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# PAF · Modem bring-up
## ARCH STATUS (2026-07-10): WORKING — verified live on the Arch install.
`mmcli`: state connected, LTE, operator Fido, packet service attached,
signal 73%. The whole solved stack made the jump to Arch: `modemsmem`
is in our kernel tree (0c2eb6bf3 "soc: qcom: add modemsmem, wire up on
blueline"), rmtfs/tqftpserv/pd-mapper/ModemManager run as services, and
`blueline-clat.service` handles 464XLAT. Everything below is the research
record of HOW the 52 wall fell (pmOS era, 2026-06/07) — the mechanism is
identical on Arch; read it before touching anything modem-adjacent.
Deep evidence docs referenced below live in `evidence/`.
## RESOLVED 2026-07-03 — modem ONLINE, registered, passing LTE traffic. Root cause of the 52 wall: missing modemsmem driver (SMEM item 134).
Google Pixel modem firmware reads SoC info from SMEM_ID_VENDOR0 (item 134) during boot;
downstream writes it via `drivers/soc/qcom/modemsmem/`. Mainline never had the driver, so
the MPSS RF task found nothing there and parked op-mode OFFLINE → every RF transition
illegal → QMI 52. Fix: Richard Acayan's sdm670 mainline port of the driver (commit
d9afe44b) + `google,modemsmem`/`qcom,smem = <134>` DT node — canonical tree `4ef9b6943`,
pmaports `9d52b47813` (CONFIG_QCOM_MODEMSMEM=y). Verified live: op-mode online, Fido LTE
registration home, IPv6 bearer passing packets. Full details in HANDOFF.md (2026-07-03).
Everything below this line is the (closed) investigation record; the leads/ledger remain
correct as history. Remaining open threads are identity (SFS/test-IMEI) and polish
(IPv4 bearer, calls/SMS), not RF.
## TOP LEAD 2026-07-02 #15 — DIAG_DATA glink channel: AP CAN request it, modem acks but never completes its own open. Real firmware-behavior finding, not a kernel gap.
**Correction to this doc's own earlier framing tonight**: an initial pass concluded mainline's glink
core lacked any AP-initiated ("locally initiated") channel-open primitive at all, and that reaching
DIAG_DATA would need core `qcom_glink` infrastructure work. **That was wrong** — re-read
`drivers/rpmsg/qcom_glink_native.c` properly: `qcom_glink_create_local()` (a complete, working
implementation of exactly this, sends `GLINK_CMD_OPEN`, waits for both acks) already exists and is
already wired into the standard `.create_ept` op every transport uses. The real, and much smaller,
entry-point gap was just: no *driver* was calling `rpmsg_create_ept()` against a fresh channel name
from an already-bound `rpdev`. Fixed with **zero new kernel code** — userspace's existing
`RPMSG_CREATE_EPT_IOCTL` on `/dev/rpmsg_ctrl2` (the modem-edge control device; `rpmsg_ctrl0`/`1` are
CDSP/ADSP respectively, confirmed via `/sys/devices/.../4080000.remoteproc/.../rpmsg_ctrl2`) reaches
this exact path already. Built a 60-line test binary (`diag_open_test.c`, cross-compiled via
pmbootstrap chroot) to drive it.
**Live result, confirmed clean and repeatable across multiple fresh single-attempt modem boots**:
- `GLINK_CMD_OPEN` sent for `DIAG_DATA` → modem.
- Modem sends `GLINK_CMD_OPEN_ACK` back, consistently in <10ms when it responds at all.
- **The modem never sends its own reciprocal `GLINK_CMD_OPEN` back** — the other half of glink's
bidirectional handshake `qcom_glink_create_local()` requires (`wait_for_completion_timeout(&channel
->open_req, ...)`) before a channel counts as usable.
- Tested whether this was just slow: bumped the wait from 5s→60s (kernel commit `7441ca3dc`), rebuilt,
reflashed, ran ONE clean attempt on a fresh modem boot, waited the full 63.84s. **Still nothing.**
Falsified the "just needs more time" theory outright. Reverted the bump (`b4185e71b`) — no reason
to ship a 60s stall for a case that structurally doesn't complete.
- Cross-checked against downstream (`references/los-kernel-blueline-4.9/drivers/soc/qcom/glink.c`):
its `GLINK_CONNECTED` notification requires the exact same bidirectional
`ch_is_fully_opened(ctx)` condition (checked from both the incoming-open and the open-ack RX
handlers) — downstream isn't looser about what counts as "open," it's just event-driven with no
hard timeout instead of blocking-with-timeout. So this isn't mainline enforcing an extra
requirement Android doesn't have; it's the same requirement, and the modem doesn't satisfy it
either way under mainline's bring-up.
**What this means**: the modem's DIAG service is reachable and responsive at the glink-negotiation
level (it correctly acks a request for a channel it recognizes by name) but will not commit to
opening it — a "deliberate partial refusal" shape that rhymes with the *entire* rest of this frontier
(SIM/QMI all healthy, DMS answers reads, refuses RF-arming writes; modem parks `offline` not `LPM`).
Whatever state gate blocks RF-arm plausibly also gates DIAG. **F3 access via this channel is not
reachable by any change on the AP side alone** — this is now a real, evidenced, narrow finding, not
an assumption. Retest if/when the LPM-vs-offline or any other frontier item moves; a modem that
reaches a healthier init state might complete this handshake as a side effect.
Test binary + method preserved: `diag_open_test.c` (cross-built via `pmbootstrap chroot -b aarch64`),
`RPMSG_CREATE_EPT_IOCTL` against `/dev/rpmsg_ctrl2`, then `open()` the resulting `/dev/rpmsgN` — same
approach works for probing ANY named glink channel on ANY edge (swap `/dev/rpmsg_ctrl0/1/2` for
cdsp/adsp/modem) if a similar question comes up for a different service.
---
## TOP LEAD 2026-07-02 #14 — frontier #8's 3-precondition test finally run: provisioning-churn fixed & shipped, LPM-vs-offline confirmed unreachable from AP side
**Fixed and shipped precondition (b) from frontier #8 below**: `blueline-uim-select.sh` was
unconditionally deactivating+reactivating the primary-gw provisioning session even when it already
matched the slot/AID it was about to select — churn on an already-correct session, and exactly the
class of write the raw-QMI RESULTS.md proved times out during the modem's not-ready window. Patched
to skip the deactivate/reactivate when the existing binding already matches
(`device-google-blueline-modem-tools` 7-r18→7-r21, committed `1ed2df1788`). **Verified live**: the
log now reads "primary-gw session already bound to slot 1 app 1 ... skipping stock
deactivate/reactivate" — the session is genuinely left alone across a modem restart for the first
time this whole investigation.
**Result: does NOT flip DMS online.** Immediately after the clean-session restart,
`--dms-set-operating-mode=online` still returns `DeviceNotReady (52)`. Polled operating-mode every
3s for 18s post-restart: **steady `offline` the entire window, never once observed `LPM`.** So
precondition (c) — "modem in LPM (not offline) for the online flip" — is now confirmed, not just
suspected, to never occur spontaneously on pmOS: the firmware decides `offline` vs `LPM` before any
AP-side QMI write is even possible, so no host script/wrapper/choreography (this fix included) can
create the window frontier #8 needed.
**Where this leaves it:** frontier #8's 3-precondition experiment (configs active + clean
provisioning + LPM-not-offline) is no longer "untested" — it's tested, with (a) and (b) now cheaply
satisfiable and (b) fixed for real, and (c) confirmed firmware-internal and unreachable from any
AP-side lever. This converges with TOP LEAD #13's conclusion (SLPI/ADSP/fastrpc/choreography all
closed, F3 narration is the remaining instrument) via a different path — both land on "the
offline-vs-LPM decision happens inside the firmware before Linux gets a vote." **Next: same F3
instrument #13 named, now aimed at a sharper question** — not generic RF-init narration, but
specifically why the firmware selects `offline` over `LPM` at its own boot, since that's now the
single remaining precondition standing between here and an online modem.
Also confirmed same session: mainline blueline has **no DIAG transport at all** (`/dev/diag` absent,
no `diagchar`-equivalent config, no `DIAG` glink channel among the live rpmsg devices) — so
`diag_capture.md`'s Side B is a real, confirmed dead end on the AP side as of kernel
`7.1.1_p20260702080259-r2`; F3 access (if pursued) needs either a from-scratch DIAG-over-glink
kernel path or stays Android-side-only (Side A) for comparison purposes.
---
## TOP LEAD 2026-07-01 #13 — fastrpc `pageslen` kernel bug FIXED & shipped; 52 wall now proven independent of adsprpcd/ADSP health too
**Real, named, diff-attached kernel defect — the first of the whole investigation — found, fixed, shipped, verified.** `fastrpc_init_create_static_process()` in `drivers/misc/fastrpc.c` advertised `inbuf.pageslen = 0` while actually passing one remote_heap page in `args[2]`; the sibling dynamic `fastrpc_init_create_process()` correctly sets `pageslen = 1` for the identical one-page layout. The DSP trusts the metadata field, believes zero pages were passed, and aborts `INIT_CREATE_STATIC` — surfacing in userspace as `Broken pipe` and the adsprpcd coredump cascade (`dsp information is incorrect err: -1` for ADSP+CDSP). One-line fix (`0``1`); matches `FisiFla/nabu-cachyos` `0005-fastrpc-fix-init-create-static-pageslen.patch` (SM8250, independently shipped for the identical signature, commit claims it also killed an SLPI crash loop there). Committed `e3c47d44e` in `references/linux-7.1.1-sdm845-CANONICAL`. **Also committed the previously-floating-uncommitted RTC read-only fix** (`68ffb014b` — the load-bearing half of the 0-SPMI-writes result; it had never been committed, only sitting in the worktree). Both shipped in kernel `7.1.1_p20260701093924-r1`, flashed to `boot_a`.
**Process correction (saved to memory `pmos-src-skips-apkbuild-patches`):** `pmbootstrap --src` builds the **working tree as-is**, not committed-only. A first build silently included the two then-uncommitted RTC/DTS files; caught by inspecting `chroot_native/mnt/pmbootstrap/source-override/`. "Bake as commits" is necessary but not sufficient — the worktree itself must be in the intended state at build time.
**Verified live post-flash:**
- **adsprpcd: no longer a fatal crash-loop.** Previously coredumped every boot with `dsp information is incorrect` / `Broken pipe`. Now: coredumps ONCE on first probe, systemd restarts it (NRestarts=1), then runs stably — **successfully opens `adsp_default_listener` handles on ADSP (domain 0)**. Old fatal signatures absent from the journal. Residual non-fatal noise remains: `remote_handle64_invoke ... method 2 ... Operation not permitted` + `open_mod_table_handle_invoke ... Error 0xe` on domain 0 — a *separate* ADSP-invoke-permissions issue, not the pageslen bug; next fastrpc thread if pursued.
- **SPMI write flood: still 0** — RTC read-only fix preserved across the rebuild.
- **`/dev/fastrpc-sdsp` still ABSENT** → `hexagonrpcd-sdsp` still "Could not attach to FastRPC node: Operation canceled". The pageslen fix targets ADSP/CDSP static-process creation; the SDSP node is a different path (depends on SLPI remoteproc fastrpc edges), so the SLPI/sensor-island story is **not** changed by this fix.
- **dms-ui-ready (r17 C binary): ran clean**, same `DMS 0x0064 RESULT: FAILURE (result=1 error=58)` as before (unchanged, expected — orthogonal to fastrpc).
**Modem 52 wall: STILL THERE — now proven independent of adsprpcd/ADSP health too.** First `mmcli --enable` of the entire session with a non-crashing adsprpcd (all chain units active: modem-qrtr-ready, qcril-autoselect, qcril-prov, dms-ui-ready, ModemManager): identical `QMI protocol error (52): 'DeviceNotReady'` on `Set Operating Mode=online`, op-mode stays `offline`, state `disabled`, lock `sim-pin2`, SIM on slot 2 (correct). Combined with TOP LEAD #12's SLPI proof, **the modem's RF-ready gate is now robust to the entire host-side DSP coordination layer** (SLPI state × 3 boots, adsprpcd/ADSP health × 1 boot). The ADSP-PD/servreg candidate is substantially weakened. The wall is the modem firmware's own internal decision; nothing host-side we can move has touched it.
**Where that leaves the candidate space:** of `los-mss-translated.c`'s four deltas — (1) EFS NV item, (2) bootloader/ABL verified-boot state, (3) subsystem coordination [SLPI/PD], (4) host-QMI-replay timing — #3 is now closed and #4 is substantially exhausted (DMS 0x0064 sends+responds but error 58; NAS mode-pref fails DeviceUnsupported pre-RF). #1/#2 were researched to ground with no static delta. The one instrument that could show the modem's *own* RF-init reasoning — F3 narration — remains uncaptured (QCSuper's blanket arm skips the NAS/RRC SSID ranges; per-range SET_RT_MASK arms cleanly but yields no traffic, and we lack a Qualcomm SSID name table to confirm those ranges *are* NAS/RRC). That F3 capture + SSID identification is the most likely remaining source of new signal.
---
## TOP LEAD 2026-07-01 #12 — SLPI independence now PROVEN (not just argued); DMS 0x0064 "UI ready" built, untested
**SLPI-state dependency is now definitively closed, with hard repeated evidence** (2026-06-30
night → 2026-07-01, three separate boots, live hardware). Fixed a real regression along the way:
`blueline-slpi-ready.sh` disabled SLPI crash-recovery for a "single-shot" start; when the SAR/CHRE
firmware fault crashed it, the disabled-recovery path left the glink-edge channel permanently torn
`intent request timed out` storming every ~10s indefinitely, which took the modem's own QMI/QRTR
transport down with it (`endpoint hangup`, `Connection reset by peer`). Fix: recovery must stay
**enabled** (matches Android's `ssr_setup`/`ssr_diag`, which never disables it) so the kernel's own
crash handler tears down/rebuilds glink cleanly; the script then does ONE clean attempt and parks
SLPI (`echo stop`) once caught in a healthy `running` window, instead of crash-looping forever.
Verified over three clean boots: SLPI crashed (0 times / 50 times / 0 times depending on timing),
**every single time `DeviceNotReady` on RF-enable was byte-for-byte identical** regardless of SLPI's
state (crashed continuously, running, offline, or cleanly parked). Also wired `modem-qrtr-ready.service`
`After=`/`Wants=blueline-slpi-ready.service` to match Android's real architecture — `vendor.qcom.devup`
(what `modem-qrtr-ready` mirrors) is **literally set by** `init.qcom.devstart.sh` right after it kicks
SLPI (`echo 1 > /sys/kernel/boot_slpi/boot; setprop vendor.qcom.devup 1`), confirmed by reading the
real Android script off the mounted LOS vendor image. Architecturally correct, safe, zero cost — but
does not move the 52 wall. **Candidate #4 ("Subsystem coordination — SLPI PD state?") in
`los-mss-translated.c`'s delta list is now closed.**
**Checked and closed a real diagnostic gap**: `qmicli --dms-get-operating-mode`'s output includes an
optional "Offline reason:" line (libqmi `qmi_message_dms_get_operating_mode_output_get_offline_reason`,
bitmask of `HOST_IMAGE_MISCONFIGURATION`/`PRI_IMAGE_MISCONFIGURATION`/`PRI_VERSION_INCOMPATIBLE`/
`DEVICE_MEMORY_FULL`) that no prior capture had actually looked for. Checked the existing captured
output (`tools/modem-online-provisioner/modem-online-provision-20260622-063154.log`): the field never
appears — the modem doesn't populate any of the four documented reasons. Dead end, but a clean one.
**Built AND tested: `dms-ui-ready.py`/`.service` + a raw-C fallback. Sends clean, modem responds,
52 unchanged.** `PAF/qcril_online_sequence.md` (prior session, full disassembly of
`libril-qc-hal-qmi.so`) identified DMS message **0x0064** ("telephony UI ready") as the one
load-bearing pre-online step QCRIL performs that pmOS had never replayed. The Python/GI
implementation (`Qmi.Message` + raw TLV writers) **failed at runtime**: this libqmi build does
not expose a generic raw-message type via GObject-Introspection at all — only per-message typed
classes exist in the typelib, confirmed by enumerating `dir(Qmi)` live on the phone. Built a
small C program instead, cross-compiled against the same libqmi (openimsd fork, 1.39.0) via
`pmbootstrap chroot -b aarch64`, linked directly against the `.so` already on the phone. Ran it
live: the modem answered 0x0064 with a real, understood protocol response (not "unsupported
message"), `RESULT: FAILURE (result=1 error=58)`. Error 58's standard libqmi meaning
(`WMS_ENCODING`, an SMS-service code) can't be its true meaning for an OEM DMS message — almost
certainly a vendor-private error space reusing that numeric slot. Immediately retried
`mmcli --enable`: **`DeviceNotReady` (52), byte-identical to every attempt all night.** Full
detail + the C source path in `qcril_online_sequence.md`. Kept `dms-ui-ready.service` wired in
the boot chain (best-effort, harmless, now proven to do something real) — the Python version in
the package needs porting to the same raw-C approach before it'll actually work at boot; that's
the concrete next step, not another blind guess.
---
## TOP LEAD 2026-06-26 #11 — THE 52 IS A MODEM XPU (SECURE-MEMORY) CRASH-LOOP
> **🔴 SUPERSEDED 2026-06-29.** Captured on the **pre-DELTA-1 r2 LLVM kernel**, where the modem
> genuinely crash-looped on an XPU violation. **DELTA-1 (`c2923cb49` MPSS assign-ordering) fixed the
> load** — on the current kernel the modem MSS runs *steady* (MBA boots, MPSS loads, `remoteproc
> running`, **no XPU crash-loop, zero 22 on `4080000`**; the only crash-lines are the SLPI's
> `SNS_REG_INIT`). Yet it **still parks OFFLINE / model 0 / 52**. So **52 is NOT an XPU crash-loop**
> — it's a firmware-internal RF-init decision, kernel-independent (6.16.7 and 7.1-rc1 both give 52,
> see #10). The 22 referenced below is the SLPI's fastrpc heap. Kept as the XPU/load investigation
> record.
**Captured live on the rebuilt r2 LLVM kernel (slot A, full system, ssh).** The
`DeviceNotReady (52)` wall is now explained at the mechanism level: the modem
**boots then crash-loops on a hardware secure-memory protection violation**, so it
can never complete RF-init and parks `offline`. dmesg (`qcom-q6v5-mss
4080000.remoteproc`, remoteproc3):
```
Booting fw image qcom/sdm845/Google/blueline/mba.mbn ... MBA booted ... loading mpss
remote processor 4080000.remoteproc is now up
crash #1: fatal error received: SFR Init: wdog or kernel error suspected.
crash #2: fatal error received: modem_xpu_int.c:456:[modem_sec_xpu_I]XPU violation
... port failed halt ... modem offline ... starting ... up ... crash ... (loop)
```
**XPU = Qualcomm eXtreme Protection Unit** — hardware memory guard programmed by
TrustZone. `modem_sec_xpu XPU violation` = the modem firmware accessed a memory
region the XPU does not permit it → the modem's **secure memory regions are not
assigned/programmed correctly** for mainline. The modem comes up, touches a
protected region, XPU faults it, q6v5 recovers it, repeat. The "running" state in
a one-shot `remoteproc/state` read is just a momentary up-window between crashes —
do not trust a single-sample `running`; watch `journalctl -b -p err` over ~30s.
**This is the same TZ layer as the SLPI/fastrpc failure.** Every modem AND slpi
restart also logs `qcom_scm firmware:scm: Assign memory protection call failed -22`
(hyp-assign / `qcom_scm_assign_mem` → -EINVAL). The fastrpc patch *tolerates* it;
the modem cannot — it needs its secure regions actually assigned. **Shared root:
this blueline's TZ rejects mainline's secure-memory-assign, → modem XPU violation
→ 52.** See `secure_load_tz_frontier.md`.
**Ruled OUT as the cause of 52 (this session, hard evidence):**
- *Provisioning / MCFG* — software MCFG `Rogers_Canada` (Casey's real carrier) is
`Active` (PDC `--pdc-list-configs=software`). Selecting the carrier config — the
thing qcril/Set-Feature does on Android — is already done, and 52 reproduces
anyway. Platform configs = 0 (blueline EFS gap) and PDC Set Feature 0x2E times
out, but neither is the 52: the modem crashes below the QMI layer.
- *Kernel config regression* — 52 reproduces on the corrected 1135-module r2 LLVM
kernel with SLPI/SSC in. Not the GCC-stripped-config disaster.
- *libqmi* — confirmed the openimsd fork is installed (apk desc: "openimsd fork,
adds PDC Set Feature msg 0x2E"; ships `Qmi-1.0.typelib`). Not the blocker.
**[RETIRED 2026-07-01 — see HANDOFF CORRECTIONS]** ~~The one-per-boot online attempt was
spent this boot ... Modem now poisoned until power-cycle.~~ Online attempts are retry-safe
within a boot; QMI `Set Operating Mode = RESET (4)` cleanly reinitializes, no reboot needed
(`tools/modem-online-provisioner/RESULTS.md`). Kept for history: that boot's attempt gave 52,
mode stayed `offline`, serving-system not-registered/detached, signal InformationUnavailable.
**Next frontier:** make TZ accept the modem's secure-memory assignment — i.e. the
XPU/hyp-assign region setup. Compare LOS pil-tz secure-load (mba region
`qcom,mba-mem`, mpss region assign, `qcom_scm_assign_mem` VMID lists) vs mainline
`q6v5_mpss` region handling on this device. The `assign -22` is the thread to pull.
(qmicli also **segfaults** on `--pdc-list-configs` — separate libqmi/fork bug to log.)
## TOP LEAD 2026-06-21 #10 — KERNEL-REGRESSION HYPOTHESIS: DISPROVEN. Kernel layer ruled out airtight.
**Result (clean-boot test, 2026-06-21 eve):** built blueline on `sdm845-6.16.7-r0` (last pre-regression
stable, 2025-09-19), flashed boot_b, booted clean. The modem boots **fully** on 6.16.7 — firmware
loads (`Booting fw image …/mba.mbn``remote processor is now up`), ALL QMI services register
(qrtr-lookup shows DMS/NAS/UIM/Voice/WDS/PDC…), UIM is healthy (card present, USIM `ready`, Primary-GW
slot1/app1), `remoteproc2` = `running`. Yet `--dms-set-operating-mode=online` → **DeviceNotReady (QMI
52)**, mode stuck `offline`, Networks `''`, Max TX/RX 0 — **byte-for-byte identical to 7.1rc1.** A
second run on a *clean* boot (modem auto-started at boot-time, UIM freshly provisioned) reproduced the
same 52, removing the manual-start confound. Per this lead's own triage rule (still-52 on the
pre-regression kernel) = **the kernel is ruled out airtight.** The online wall is NOT a kernel
regression; it is device-internal (Route B) — the modem's own firmware RF-init never self-completes, so
it parks `offline` and refuses `online` (Android never has to ask; see frontier #8 REFRAME).
6.16.7 is now a known-good working kernel for blueline (modem boots clean); either 6.16.7 or 7.1rc1 is
fine going forward (both give 52). **The live frontier is Route B — see `frontier #8 REFRAME` below.**
### Firmware-path gotcha found + fixed (RECORD — bites any upstream-tree blueline DTB)
The 6.16.7 *upstream* `sdm845-google-blueline.dts` requests firmware at
`qcom/sdm845/pixel3/{mba,modem,adsp,cdsp,ipa_fws,a630_zap}.mbn`, but pmaports packages blueline
firmware at `qcom/sdm845/Google/blueline/`. So the upstream-tree DTB silently fails `request_firmware`
(ENOENT -2) → modem/ADSP/CDSP never boot (`remoteproc: Boot failed: -2`). The 7.1rc1 *pmaports* DTB
uses `Google/blueline/` (correct) — which is why 7.1rc1 loaded firmware and 6.16.7 did not until fixed.
**Fix applied on the rootfs (persists across reboots):**
`ln -snf /lib/firmware/qcom/sdm845/Google/blueline /lib/firmware/qcom/sdm845/pixel3`
Any future build using the upstream blueline DTB needs this symlink (or a DT patch correcting
firmware-name paths) or the modem will not load firmware. (Stray: the 6.16.7 DTB also references
`qcom/sdm845/oneplus6/venus.mbn` for venus — harmless, video-decoder only.)
### Kernel-swap procedure (RECORD — the modular-modem-stack gotcha)
The blueline modem stack is **fully modular** on this config: `qcom_q6v5_mss`, `qrtr`/`qrtr-smd`,
`qcom_glink_smem`, `qcom_sysmon`, `mhi*` all `=m`. So swapping kernels needs the matching modules on the
rootfs or the modem never loads (a false 52). UFS + USB-gadget-NCM are built-in (`=y`), so a
module-mismatched boot still reaches rootfs + USB. Working procedure (validated this session):
build kernel .apk → `tar xzf` it → extract `lib/modules/<kver>/` → scp to phone → `sudo cp -a` into
`/lib/modules/` + `depmod <kver>` → reboot. Boot.img = `cat vmlinuz <blueline>.dtb` + the standing
initramfs, `mkbootimg` (header v0, base 0x0, kernel_off 0x8000, ramdisk_off 0x1000000, tags 0x100, the
real cmdline `quiet splash … pmos_boot_uuid=2CE8-836F pmos_root_uuid=26733d78…`). USB iface
`enp0s20f0u6` (cdc_ncm) needs manual bring-up each boot: `nmcli dev set … managed no; ip addr add
172.16.42.2/24 dev …; ip link set … up`.
### Original #10 thesis + bisect-window analysis (HISTORICAL — hypothesis disproven above)
Thesis (Casey): the 7.1rc1 bump (pmaports MR !8446, merged 2026-04-28, "basic smoke testing… Pixel 3")
never validated blueline's modem; the modem may have regressed there. Corroboration: beryllium wiki
flags 7.1.0-rc1-sdm845 for modem breakage on the same SoC; Joel Selvaraj had SDM845 modem working on
mainline back at 5.11, so the capability predates the regression. **Running test:** rebuild blueline
kernel against `sdm845-6.16.7-r0` (last pre-regression stable, 2025-09-19), flash boot_b, retry
`--dms-set-operating-mode=online`. Outcomes: online → regression confirmed, bisect the bump; flaky
like beryllium → kernel layer validated, inherit beryllium recipe; still DeviceNotReady(52) → kernel
ruled out airtight, it's device-internal firmware/EFS state (Route B).
### Kernel-tree pull (2026-06-21) — what the AP-driver diff says about the bisect window
Pulled `references/sdm845-linux-20260610-ONHOLD-DAVID-TREE` (codeberg `sdm845-next-20260610`, base **7.1.0**) and the
`sdm845-next-20260306` tarball (base **7.0.0**) and diffed the modem-relevant subsystems:
- **`drivers/remoteproc/qcom_q6v5_mss.c`: functionally identical** 7.0→7.1 (brace-whitespace only).
- **`drivers/soc/qcom/pdr_interface.c` + `pdr_internal.h`: rename-only** — `SERVREG_LOCATOR_SERVICE`
`QMI_SERVICE_ID_SERVREG_LOC` (both 0x40), `SERVREG_NOTIFIER_SERVICE``QMI_SERVICE_ID_SERVREG_NOTIF`
(both 0x42), `reason[257]``reason[SERVREG_PFR_LENGTH+1]`(=257). No behaviour change.
- **`qcom_pd_mapper.c` present in BOTH trees** → the in-kernel pd-mapper is NOT new to 7.1 (already
in-kernel by 20260306 / 7.0.0; it landed ~mainline 6.13, so 6.16.7 has it too). The "in-kernel
pd-mapper is the regression" sub-hypothesis is weak; the **sdm845-mainline rebase** is the live one.
- `qmi_interface.c`, `qmi_encdec.c`, `rmtfs_mem.c`: identical 7.0→7.1.
**Implication for the bisect:** in the codeberg 7.0→7.1 window there is NO functional change in
mss/pdr/qmi/rmtfs — so if the regression is real and in that window, it is NOT in those drivers
(look at clk/regulator/interconnect/scm/smem/glink/IPA instead). **Caveat:** pmaports tracks
`gitlab.com/sdm845-mainline/linux`, NOT codeberg `sdm845/linux` (the two repos the Matrix thread
argues over), and my window (7.0→7.1) does not cover the 6.16.7→7.0 jump. So this scopes, it does
not pin — Casey's 6.16.7 build is still the right swing.
### Upstream blueline DTS confirms our DT is already correct (kills #9 a second way)
The 20260610 tree ships an **in-tree `sdm845-google-blueline.dts` + `sdm845-google-common.dtsi`**
(model "Google Pixel 3"). Its modem nodes: `&mss_pil` = `blueline/mba.mbn`+`blueline/modem.mbn`,
`&ipa` = `blueline/ipa_fws.mbn`, `mpss memory@8e000000` size `0x9800000`, `rmtfs-region@f2700000`
vmid MSS_MSA. **`dt-work/blueline-pmos.dts` already matches all of these byte-for-byte.** Zero
WTR/QFE/RFFE/FEM/GRFC regulator nodes in the upstream blueline DTS — same as fajita, same as pmOS.
So the pmOS modem DT is upstream-aligned, not a fajita guess. RF power is modem-firmware/RFFE driven
(see `dt-work/rf-rail-diff-FINAL.txt`). **Lead #9 (add 11 pm8998 LDOs) is refuted** by (a) the
upstream blueline DTS not declaring them either, and (b) all 11 candidate LDOs showing use=0 from the
AP on live working Android. Do NOT flash `dt-work/blueline-live.dtb` for this reason.
### qrild finding folded in — there is no hidden online handshake
`references/qrild` (abandoned AOSPM RIL): its modem bring-up (`qrild_android_modem.cc` ctor) is
DMS Set Operating Mode RESET → powerup → NAS register-indications, with `setRadioPower` a no-op stub.
Same levers ModemManager already pulls — **no secret QMI sequence** that arms RF. This corroborates
the #8 REFRAME (`qcril_bringup_capture.md`): on Android the modem firmware SELF-onlines at its own
RF-init; replaying a "QCRIL/RIL online handshake" on pmOS is a dead end because there isn't one.
---
## (REFUTED 2026-06-21) #9 — pmOS DT under-defines the pm8998 regulator set
Refuted by #10 above (upstream blueline DTS does not declare these LDOs either; all 11 are use=0 on
live Android) and by `dt-work/rf-rail-diff-FINAL.txt` (RF power is firmware/RFFE-driven, not AP-DT).
Kept for the record; do NOT flash the LDO DTB. Original reasoning below.
Direction (Casey + GLM): the fix is in OUR pmOS implementation/DT, make it mimic Android. Concrete,
evidenced bug found this session:
**Our mainline blueline DT instantiates only ~14 of the 24 pm8998 LDOs. Android defines all of them.**
Missing from `dt-work/blueline-pmos.dts` rpmh-regulators (present in downstream Android DTB):
`ldo4, ldo6, ldo8, ldo9, ldo10, ldo11, ldo15, ldo16, ldo18, ldo23, ldo27`.
Voltages (from `dt-work/android-dtb/android-base.dts`, full diff in `dt-work/regulator-diff-android-vs-pmos.txt`):
ldo6=1.856V ldo8=1.2V ldo9=1.704V ldo10=1.704V ldo11=1.0V ldo15=1.8V ldo16=2.704V ldo18=2.704V ldo23=3.0V.
Why this matters even though all 11 are `deviceless`/use=0 in BOTH regulator_summaries: **the modem
votes its RF-frontend rails (WTR5975/PA/tuners) directly via RPMh, not the AP regulator framework.**
If the AP DT doesn't define a rail in the rpmh-regulators node, it isn't in the Linux RPMh regulator
set — so when the modem's RF-init requests it, the vote can't be satisfied/arbitrated. Modem RF-init
fails → modem parks in `offline` and refuses LPM/online (DeviceNotReady / silent no-op). This fits
every symptom: modem boots (DSP up) but RF never inits; `offline` won't flip; configs don't help.
Supporting evidence this session:
- Modem FATAL crash in pmOS dmesg: `qcom-q6v5-mss: fatal error received: ...Asserted in
wlan_vdev.c:_wlan_vdev_create` + `port failed halt` on stop (subsystem not halting cleanly →
modem can't flush NV → MCFG selection lost across restart → 0 configs every boot).
- Modem comes up `offline` (not LPM like Android); `set low_power`/`online` return rc=0 but mode
never changes; PDC autoselect (Set Feature) does NOT persist across modem restart (RIL-layer, not
firmware NV).
### Executable plan (start here next session — do NOT re-derive):
1. **Diff our blueline DTS rpmh-regulators block vs UPSTREAM linux-next `sdm845-google-blueline.dts`**
(Heidelberg's). Our pmaports kernel build may predate the full regulator set. If upstream already
defines all LDOs, bump/patch the kernel DTS; if not, add them.
2. **Add the 11 missing LDOs** to the `pm8998-rpmh-regulators` node in the blueline DTS at the Android
voltages above; mark the RF-candidate analog rails `regulator-always-on` so they're up at RF-init.
3. Build via pmbootstrap (`linux-postmarketos-qcom-sdm845`), flash `boot_b` (see platform.md), boot.
4. `dmesg | grep -iE 'regulator|vreg|q6v5|fatal'` — confirm no rail/halt errors; then
`--dms-set-operating-mode=online` → `--dms-get-capabilities` Networks non-empty = RF armed.
5. **Leverage other devs (Casey): do NOT silo.** Cross-check the SDM845 mainlining wiki, sargo/bonito
(Google sister devices) DTS, and the `#sdm845:postmarketos.org` channel for how RF rails are
declared. Heidelberg's blueline series on lore.kernel.org is the authoritative reference.
Artifacts: `dt-work/regulator-diff-android-vs-pmos.txt`, `dt-work/android-dtb/android-base.dts`
(decompiled downstream DTB), `dt-work/blueline-pmos.dts`.
---
The modem is a self-contained DSP (MPSS) running Qualcomm firmware. Linux only loads it, powers
it, and talks QMI to it. RF is the modem's own internal state machine — Linux cannot drive RF
directly. Read this stack bottom-up; the blocker is at L6.
## The bring-up stack
| L | Layer | What it is | Status (2026-06-19) |
|---|---|---|---|
| L0 | Silicon / RF | SDM845 SoC; **SDR845 + WTR5975** transceivers; pm8005/pm8998 PMICs | fixed HW |
| L1 | Firmware load | `mba.mbn` (238 KB) → `modem.mbn` (67 MB MPSS), via remoteproc `q6v5-mss` | healthy — "remote processor is now up" ~12.8 s |
| L2 | Power | rpmhpd domains `cx`/`mx`/`mss`; `mss`→`mss.lvl` ARC→`pm8005_s2_level` | healthy — `q6v5` votes `mss` INT_MAX; verified vs Android |
| L3 | Transport | QMI over QRTR; `glink-edge` to the modem PID | healthy — all QMI services register |
| L4 | Host services | `rmtfs` (EFS), `pd-mapper`, `tqftpserv`, `qrtr-ns` | **CRITICAL 06-19**: rmtfs binary was replaced (69KB→141KB) and config changed from `-r -P -s` to `-P -s`, causing EFS corruption + modem crash + system freeze at "[stack] dump". Original binary at `/usr/bin/rmtfs.orig-20260618` must be restored. **CORRECTION (06-19 #2): `-r` = read-only is WRONG for arming the modem — Android serves EFS read-write and the modem writes modemst2 on boot. Use `-P -s` (read-write). The 02:38 corruption was the wrong binary, not the rw flag.** |
| L5 | SIM / UIM | card detect, app select, primary-GW provisioning | healthy — `usim (2) ready`, slot1/app1 |
| **L6** | **RF arm** | `DMS Set Operating Mode → online` arms the RF stack | **BLOCKED — DeviceNotReady (QMI 52).** The 06-19 #2 "rmtfs read-only" root-cause is **DISPROVEN (2026-07-02)**: rmtfs runs `-P -s` read-write live (drop-in enforced, all EFS partitions present) and 52 persists. rmtfs ruled out. Real cause is the modem's own boot-time RF-init decision (Route B), not an AP EFS-write path. |
| L7 | Data / voice | IPA/rmnet, ModemManager, calls/SMS/data | gated behind L6 |
**The blocker, precisely:** `DMS Set Operating Mode=online` → **DeviceNotReady (QMI 52)**;
`DMS Get Capabilities` → zero channel rates / empty networks = the RF stack never initialized
inside the DSP. Everything upstream (L1L5) is proven healthy. ModemManager reaches `enabling`,
hits this, falls back to `disabled`, cycles.
## Eliminated leads — DO NOT RE-CHASE
Each killed with proof. The "proof" column is why it's dead, not a hypothesis.
| Lead | Verdict | Proof |
|---|---|---|
| rmtfs NV-files / `msm-cros-efs-loader` (serve EFS from file copies) | DEAD ( UPDATED 06-19 ) | **REAL CAUSE FOUND**: rmtfs binary was replaced (69KB→141KB) on 06-19 02:38 and service config changed from `-r -P -s` to `-P -s`. This broke EFS partition access, causing modem crash → "[stack] dump" → system freeze. Fixed by restoring `/usr/bin/rmtfs.orig-20260618` (69192 bytes) and resetting `/etc/systemd/system/rmtfs.service.d/writable.conf` to `-r -P -s`. The NV-files approach is dead; the original partition-mode rmtfs MUST be used. |
| `removed-region@88f00000` carveout (was "lead #1") | DEAD | Android iomem: `88f00000-8aafffff : System RAM`. The device whose modem **works** leaves it allocatable. OnePlus-specific quirk. (06-19) |
| cmd-db maps `mss.lvl` to wrong PMIC (was "lead #3") | DEAD | Android regulator tree: `4080000.qcom,mss → pm8005_s2_level` (correct rail). cmd-db is the same firmware blob on both OSes. pm8005 absent from pmOS `regulator_summary` is **expected** — mainline drives it via rpmhpd/genpd, not the regulator framework. (06-19) |
| `vdd_mss-supply` / SPMI `pm8005@4` in DT | DEAD | Mainline `q6v5-mss` is power-domains-only for SDM845 (no `devm_regulator_get`). `qcom,pm8005` has **no SPMI driver** in mainline → adding `pmic@4/5` panics the kernel. Only `smps1/2` under the rpmh-regulators-2 block is safe. (06-18/19) |
| "blueline lacks `soc-qcom-sdm845-modem` dep" | DEAD (stale doc) | Current pmaports pkg is `soc-qcom-modem`; blueline depends on it just like enchilada. Device-pkg deps + firmware manifests are equivalent. (06-19) |
| fcc-unlock missing | DEAD | `fcc-unlock` exists nowhere in pmaports; no SDM845 device uses one. Modem is not FCC-locked-waiting. (06-19) |
| Modem DT node differs from working device | DEAD | blueline modem node is byte-identical to fajita (same power-domains/clocks/resets); only firmware path + phandles differ. (06-18) |
| Storage / EFS calibration is missing | DEAD | rmtfs writable + Android-primed EFS (modem online in Android first) → still DeviceNotReady in pmOS. EFS content is not what's missing. (06-18) |
**Net:** every kernel / DT / power / packaging layer is proven equal to blueline-**Android**
(which runs the modem online, Fido LTE) on the **same silicon**. The blocker is not there.
## The frontier (REVISED 2026-06-20 #8 — QCRIL does more at bring-up; mainline stack under-initializes the modem)
Casey's standing thesis (correct, and the through-line of this whole effort): **QCRIL powers and
initializes more of the modem than ModemManager does** — and mainline pmOS, as a minimal port,
under-powers/under-initializes the board. We are hunting the specific delta. Evidence gathered:
1. **QCRIL sends a `'UI ready'` / power-on-optimization handshake to the modem that MM never sends.**
Android cold-boot capture (`android-live-20260619/deep/radio-boot.txt`): `qcril_qmi_nas_telephony_ready:
poweron_opt: send 'UI ready' to modem` (qcril_qmi_nas.cpp:17527), plus an active radio-power
state machine (`qcril_qmi_nas_update_radio_power_request_info` → `BOOTUP_RADIO_POWER_OFF`,
`qcril_qmi_nas_radio_power_transaction_handler`, state line `telephony LPM | ril NONE | card UP |
modem ONLINE`). ModemManager has no equivalent of this poweron_opt / UI-ready signalling.
2. **No DT RF node on either OS** (see retracted #6) — so the delta is QMI-level (modem-side
init/power), not a Linux regulator we can add. The mainline DT brings up cx/mx/mss + the basic
stack; it does NOT replicate QCRIL's QMI bring-up choreography.
3. **Provisioning race (real bug, partially validated).** `msm-modem-uim-selection` fails every boot:
it fires during the early QRTR window ("node with id 0 not found"), then its primary-GW provisioning
ACTIVATE times out → service exits FAILURE → the live `primary-gw-provisioning` session loops
`end-with-failure`/`app-reset`. Run manually AFTER boot it SUCCEEDS (`Successfully changed
provisioning session`). KEY: with provisioning freshly+cleanly active, `--dms-set-operating-mode=
online` **no longer returns DeviceNotReady(52)** (returns rc=0) — though the modem still stays
`offline` because PDC configs were 0 that boot. We have never had configs-active AND
clean-provisioning AND the QCRIL init handshake all at once.
4. **QCRIL online path captured (airplane off→on, `android-capture-20260620/qcril-online-capture.txt`):**
`qcril_qmi_nas_dms_handle_modem_operating_mode` (qcril_qmi_nas.cpp:18289) → "modem's operating mode
needs to be changed" → ril SETTING_OP_MODE → **modem LPM → ONLINE** in ~30ms (a plain DMS
set-operating-mode=online). post-transition: `set_bootup_power_optimization_state=2`, ATEL UI-READY,
enable networking/sys-info indications. **No hidden RF-power QMI command** — the modem is already
powered (LPM) and the flip arms RF. **KEY STATE DIFFERENCE:** Android modem starts in **LPM**
(low_power) and does LPM→online; the pmOS modem sits in **`offline`** (deeper state) and
offline→online is what fails. So the gate is the combination: (a) MCFG configs loaded, (b)
primary-GW provisioning active, (c) the modem in LPM (not offline) for the online flip. We have
NEVER had all three at once on pmOS — that is the decisive untested experiment.
**The decisive instrument (next):** a QMI-message-level capture of QCRIL's bring-up on the rooted
Android (enable QMI verbose radio logging, cold boot or airplane-toggle), then diff the QMI request
sequence against ModemManager's (we already have MM's via `--debug`). The delta = exactly what to
replay on pmOS (UI-ready/poweron_opt, any NAS/DMS/RF config QCRIL issues). Then replay on pmOS with
configs active + provisioning clean. Rooted Android (slot B) is the instrument; keep this doc current
as the capture lands.
---
## The frontier (REVISED 2026-06-20 #7 — live RF-init trace: blocker reframed to primary-GW provisioning, NOT RF cal)
Ran the decisive pmOS-side RF-init trace (slot B, kernel 7.1.0-rc1, modem over qrtr://0). Two
big results — one closes a thread, one opens a sharper one.
**The modem is not faulting and nothing it requests is missing.** dmesg: clean boot, mba→mpss→
running→handover, NO crash/coredump/fatal in the RF-arm window; an `online` attempt produces ZERO
dmesg delta. tqftpserv serves only the readonly MBN tree (read OK) + an empty `mcfg.tmp` scratch
file the modem RRQ/WRQ-loops on — and `mcfg.tmp` being 0 bytes is NORMAL (tqftpserv is bind-mounted
to the **persistent** persist partition `/dev/sdf2[/rfs]` — NOT volatile tmpfs as first feared —
and persist is SHARED with Android, so if the working modem needed content there it would persist;
it's empty on both). So serving is correct and complete. **The modem boots healthy and simply
REFUSES `online` at the QMI layer (DeviceNotReady = clean refusal, not a load failure).**
**NEW lead — primary-GW SIM provisioning session fails.** ModemManager's log during the refusal:
`refresh indication received: session type 'primary-gw-provisioning', stage 'end-with-failure',
mode 'app-reset'` — looping every MM retry (~72s), alongside a NAS system-info indication that
returns `FAILURE: InformationUnavailable`. Operating mode `offline`, **HW restricted `no`** (not
FCC/HW-locked), serving system fully detached, radio interface `none`. AND `msm-modem-uim-selection`
**failed this boot**: "Application was already selected" → `operation failed: Transaction timed out`
→ service exit FAILURE. Yet the static card read is healthy: Primary GW bound (slot1/app1), USIM
app `ready`, personalization `ready`, PIN1 `disabled`. So the binding exists but the live
provisioning SESSION keeps failing + app-resetting. This is distinct from (and more proximate than)
the RF-cal lead — the modem may refuse `online` because primary-GW subscription provisioning never
completes, not because RF is uncalibrated. (mBands=[] is non-diagnostic per #6; generic IMEI works
on Android per Casey — so cal/QCN lead is weak.)
**OPEN — causality not yet proven.** Why does the provisioning session end-with-failure while the
GW binding shows bound+ready? Candidates: (a) `msm-modem-uim-selection` timing out leaves the
subscription half-selected; (b) the modem triggers an app-reset refresh (post-MCFG) that nothing
on the AP side completes the way Android's RIL does; (c) downstream symptom of offline state.
**Decisive next test (fresh boot, one clean shot):** on a fresh slot-B boot, before MM enables,
(1) confirm whether the primary-gw-provisioning refresh fails on its own; (2) re-run UIM selection
cleanly (the service timed out) and retry `online`; (3) capture the UIM/provisioning QMI sequence.
Compare to what Android's RIL does on the same refresh. Known: a remoteproc reset (online is retry-safe, see CORRECTIONS)
both clear UIM provisioning, so this needs a clean boot.
---
## The frontier (REVISED 2026-06-20 #5 — MCFG SOLVED; online blocker isolated to device-specific RF bring-up)
Live pmOS run this session SOLVED the MCFG provisioning blocker AND, by doing so, proved MCFG was
never the RF-arm blocker. The online wall (`DeviceNotReady` QMI 52) survives every config/stack/SIM
fix. Full elimination, all on hardware (slot B, kernel 7.1.0-rc1-sdm845):
**SOLVED — persistent MCFG self-provisioning.** `qcril-autoselect.py` (PDC Set Feature 0x2E,
Selection Mode=3) made the modem self-load 2 platform (SS-LA-SDM845) + 25 software (Rogers_Canada)
configs from the tqftpserv tree; survives reboot (self-loads at RF-init). Both ACTIVE; software
version `0x801481A` == Android's exact value. The "0/0 every cold boot" blocker (frontier #3/#4)
is dead. Tooling: `tools/qcril-autoselect.py` + `tools/build-openimsd-libqmi.sh` (openimsd libqmi
fork, has Set Feature). Believed a Pixel 3 mainline first.
**RULED OUT as the online (RF-arm) blocker — every one tested on hardware:**
- MCFG/PDC — both configs ACTIVE, matching Android. online still 52.
- UIM/SIM — card present, USIM app `ready`, Primary GW slot1 app1 provisioned.
- timing/lock — MM stopped, 10 patient online retries over 30s, all 52.
- ModemManager — its full 9-step enable (power-state → op-mode indication register → Set Operating
Mode 0x2E=online) returns the same 52. Not a manual-qmicli artifact.
- missing stack — pd-mapper, hexagonrpcd, q6voiced, rmtfs (`-P -s` rw), tqftpserv ALL running;
`soc-qcom-modem` installed; PD JSONs (modemr.jsn/modemuw.jsn) + firmware present; no errors.
Matches working SDM845 devices (enchilada/fajita) component-for-component.
- EFS not served — modemst1/2, fsg, fsc all populated (~260K nonzero) + served. (modemcal sdd4 is
all-zero; persist sdf2 has data.)
- firmware mismatch — modem sw version `g845-00194-210812-B-7635520` IDENTICAL to Android;
revision `MPSS.AT.4.0.c2.17-00194-SDM845_GEN_PACK`.
**THE REMAINING BLOCKER = device-specific RF bring-up.** Same firmware + same shared EFS + complete
stack + DT byte-identical to fajita, yet: NAS `get-system-info` empty, `nas-get-rf-band-info` →
no bands, mode stuck offline, `DeviceNotReady` on online. The modem boots but its RF never inits.
Two leads:
1. **RF cal / QCN** — generic test IMEI `990012001581294`, model `0`, no bands. Classic
no-calibration signature. UNCONFIRMED: does the WORKING Android show the REAL IMEI (→ pmOS isn't
loading cal/QCN) or the same generic one (→ red herring)? Decisive next test: reboot slot A,
read IMEI + DMS caps + bands; compare. (Not done autonomously — avoids leaving the phone
off-pmOS while unattended.)
2. **Board RF-frontend** — fajita (OnePlus 6T) works on mainline; blueline (Pixel 3) shares the
q6v5-mss DT node but has a different PMIC/RF frontend (WTR/QFE/antenna tuners). If blueline's RF
frontend needs power/GPIO/cal mainline doesn't provide for THIS board, the modem can't arm RF
even with identical firmware. (Distinct from the DEAD pm8005/mss-rail lead — this is the RF
frontend, not the modem-subsystem rail.)
### 2026-06-20 #6 — RETRACTED: do NOT treat the board-power lead as closed
I previously wrote this section as "lead #2 CLOSED with positive proof." That was an overclaim and
Casey was right to reject it. What the regulator/DT evidence below actually shows is narrower: there
is no *labeled RF-frontend regulator node* in the device tree — confirmed on BOTH mainline AND the
**downstream Android DTB** (`dt-work/android-dtb/android-base.dts`, decompiled from slot-A boot_a;
its `qcom,mss@4080000` declares only vdd_cx/vdd_mx/vdd_mss, no RF/WTR/antenna node anywhere). That
means we can't fix this by copying an "RF rail" node — NOT that board power is irrelevant. Mainline
blueline is a **minimal fresh port**: it instantiates only rails with explicit consumers and leaves
every other board rail off, whereas the downstream stack + QCRIL bring up far more (subsystems,
power-on handshakes, NV/RF init via QMI). The live lead remains: **what does the modem need powered/
initialized — via Linux rails the mainline port omits, or via QMI commands QCRIL sends — that pmOS
does not provide.** See the #8 frontier (top of file) for the QCRIL-does-more evidence.
(original text retained below for the evidence it contains)
Brought the "power Android provides" thread full circle on live hardware (Android slot A, rooted):
- **No Linux-side RF rail exists on EITHER OS.** In both the Android (`rf-compare/regulators-android.txt`)
and pmOS (`regulators-pmos.txt`) regulator dumps, the modem `4080000.qcom,mss` consumes ONLY the
Q6 DSP level rails (s6_level=CX, s9_level=MX, s2_level=MSS). There is no WTR/PA/tuner regulator
in Linux on Android either — RF-frontend power is driven by the modem firmware over RFFE + GRFC,
not the Linux regulator framework. So a "missing regulator in the pmOS DT" was never possible.
- **DT modem node is correct + complete** (`dt-work/blueline-pmos.dts`): power-domains cx/mx/mss,
xo/bus/mem clocks, glink "modem" transport, `gpio-reserved-ranges = <0 4, 81 4>` (Linux leaves
the modem-owned GPIOs alone). No regulator-supply phandles — correct for mainline q6v5-mss.
- **Pinctrl: no AP/modem GPIO conflict.** Android TLMM (`3400000.pinctrl`) owns only camera-CCI,
gpio-keys, and USB-extcon pins; the modem's GRFC/RF-control pins are MSS-internal and not
AP-routed. The AP cannot be stepping on an RF pin.
- **"No bands" is NOT a distinguishing signal.** The fresh working-Android radio capture
(`android-capture-20260620/android-radio-log.txt`) shows `mBands=[]` in CellIdentityLte while
fully attached to Fido LTE (REG_HOME, EARFCN 2000 = B4). So the empty mBands we saw on pmOS does
not by itself mean RF is dead — weakens (does not kill) lead #1's reasoning.
**Verdict:** every AP-side power/GPIO/pinctrl/DT input the modem receives is provably equal to
blueline-Android (which arms RF on the same silicon). Lead #2 is closed the same way IMEI closed
lead #1. The blocker lives strictly inside the modem firmware's RF-init, which the AP only feeds
via: firmware blobs, EFS (rmtfs), persist/rfs (tqftpserv), power-domains, clocks, QMI. All proven
present/equal — EXCEPT we have never traced what the modem *requests* from rmtfs/tqftpserv during
its RF-init, nor captured the modem's own failure reason.
**Decisive next experiment (pmOS-side, not yet run):** boot slot A, attempt `online`, and capture
(a) `tqftpserv`/`rmtfs` request trace during RF-init — which file/NV path the modem asks for and
whether it's served (frontier #4 showed Android rmt_storage *reads* on boot; we never logged the
pmOS read side), and (b) the modem remoteproc coredump / SSR reason on the failed online. This is
the one byte-for-byte trace we still lack; it points at the specific RF-init resource that differs.
Route B (autoselect EFS NV item) is DEPRIORITIZED: it would only persist the config state we
already achieved, and config-active does not fix online.
---
## The frontier (REVISED 2026-06-20 #4 — Android boot is VERIFY-ONLY; the modem persists MCFG itself)
Source: analysis of the captured Android radio logs in `android-live-20260619/deep/`
(`radio-boot.txt`, `radio-full.txt`, `ssr-capture.txt`), 2026-06-20. This **overturns** both
frontier #3's "just make the platform config active" and the `qcril_full_provisioning.md §D.4`
"load+activate every boot" model. Read this before either.
### What the captures actually show
Across all three logs (~12K PDC log lines, including a full modem SSR) QCRIL made **only query
calls** — never load/select/activate, never `pdc_enable_auto_selection`:
- `ssr-capture.txt:24293` — the decision point:
`qcril_qmi_pdc_is_modem_mbn_updated: prev_ver_info: MCFG-g845-00194-220517-B-8604334,
cur_ver_info: MCFG-g845-00194-220517-B-8604334` → **prev == cur → reload SKIPPED.**
- `ssr-capture.txt:23949` — `get_active_config_info_ind: config_version = 0x801481a`,
`Update CNV from 0x48 to 0x48` (0x48 = **72** = `cnv=[72]` = `persist.vendor.radio.mbn0=72`).
- Only PDC functions logged anywhere: `get_active_config_info`, `get_selected_mbn_config`,
`get_selection_carrier`, `is_modem_mbn_updated`. **Zero** `*_load_config`, `*_select_config`,
`*_activate_config`, `enable_auto_selection`, `enable_modem_update`.
### What that means
**Android's steady-state boot does almost nothing to PDC.** The modem holds its MCFG persisted in
its own EFS (`modemst1/2` — the 2 MB write to `modem_fs2` the `rmt_storage` capture caught),
**self-applies it at firmware RF-init every boot**, and QCRIL just confirms the active version
matches `persist.vendor.radio.cnv.ver_info` and stops. The full load→select→activate path runs
**only on first provision or an MCFG version change** (the `prev_ver != cur_ver` branch) — which
we have **never captured**, because every capture was an already-provisioned device. That is
exactly why `pdc_enable_auto_selection` never appears: it lives on the unobserved path.
### Consequences for the two prior theories
- **Frontier #3 ("just need an active platform config")** — necessary but NOT sufficient. PROGRESS
L165172 had BOTH configs active (SS-LA-SDM845 + Rogers) and DMS Networks stayed empty. Active
≠ armed. Don't chase "make it active" as the finish line.
- **`§D.4` ("load+activate every boot, accept two restarts")** — that is a *reconstruction of the
unobserved first-provision path*, not what Android does in steady state. It remains a valid
pmOS *workaround* (see `tools/qcril-prov.py`), but it is unproven on hardware and is not the
mechanism that actually arms RF on Android.
### The pmOS problem, restated from the evidence
The "0/0 configs every cold boot" volatility (PROGRESS L140144) is **not** something QCRIL fixes
per-boot. It's that **the modem cannot persist/restore its PDC selection in EFS the way Android
does.** The earlier "rmtfs read-write didn't fix it" result (ledger L38) is a chicken-and-egg:
writable EFS is useless if the modem never had a valid MCFG to commit, and a one-shot userspace
activate is useless if it cannot persist to survive the next firmware RF-init. **Both are
required: a valid config loaded AND modem-side persistence that survives reboot.**
### Two candidate paths forward
- **(A) Per-boot provisioner workaround** — `tools/qcril-prov.py` loads HW+SW, selects, activates
each boot. Mirrors the reconstructed first-provision path. Unproven; "active ≠ armed" is the
open risk.
- **(B) Provision-once + persistence (the Android mechanism)** — get a valid MCFG into the modem
once, then ensure `modemst` persistence so the modem self-restores + self-applies at RF-init,
and userspace reverts to verify-only. This is what the captures show actually works on hardware.
### The capture we still need (next experiment)
A normal-boot capture only re-shows the skip path. To record the **real first-provision load
sequence** (and settle whether `pdc_enable_auto_selection` / `enable_modem_update` is a required
precondition), **force a first-provision on Android**: clear `/data/vendor/modem_config/` and
`persist.vendor.radio.cnv.ver_info`, reboot, then capture. Runbook:
`PAF/first_provision_capture.md`.
---
## The frontier (SUPERSEDED 2026-06-19 #3 — see #4 above — PDC platform config is the blocker)
After a full day of Android-vs-pmOS forensics, the blocker is identified:
### What's actually wrong
The modem's **PDC (Policy Decision Control)** stores MBN configs in modem NV:
- **25 software (carrier) configs** loaded — **Rogers_Canada is ACTIVE** (133KB, v0x801481A)
- **2 platform (hardware) configs** loaded — **SS-LA-SDM845 is INACTIVE** (56KB)
Without an **active platform config**, the modem doesn't know its RF hardware capabilities.
DMS reports `Networks: ''`, NAS home network returns `NotProvisioned (QMI 16)`, and
`Set Operating Mode=online` returns `DeviceNotReady (QMI 52)`.
On Android, QCRIL (`libril-qc-hal-qmi.so`) loads MBN files from `/vendor/rfs/msm/mpss/readonly/vendor/mbn/`
via QMI PDC, activates both platform and software configs, and the modem restarts with full RF.
ModemManager does none of this — it assumes the modem is pre-provisioned.
### Attempted: direct PDC activation
`qmicli --pdc-activate-config=platform,6FE571A9...` returned **"Successfully requested config
activation"** but triggered a modem restart that **cleared the platform config** (0 configs after
restart). The activation request succeeded at the QMI level but the modem's internal handling
removed the config during restart. Needs Android QCRIL to re-load it.
`qmicli --pdc-load-config=...` **segfaults** on MBN binary files (qmicli 1.39.0 / libqmi 1.38.0).
Cannot load MBN files from pmOS side — libqmi bug or MBN format incompatibility.
### Key facts confirmed (06-19)
- **IMEI: identical** on Android and pmOS (`990012001581294`). NOT the blocker. Android works on
Fido LTE with this same test IMEI.
- **Modem firmware: identical** SHA256 (`97c1d19c...`) between slots A and B.
- **Modem crashes** (SFR Init / XPU violation): **intermittent** — some boots 0 crashes, some 2.
Even clean boots get DeviceNotReady. Not the blocker.
- **rmtfs DOES serve correct partitions**: partlabels are GPT-stable; UFS LUN renumbering changes
`/dev/sdX` names between boots but `by-partlabel` symlinks always resolve correctly.
- **DMS Get Capabilities** consistently shows `Networks: ''` and `Max TX/RX: 0` on pmOS.
Android shows populated networks (Fido LTE, Rogers 302720, full band support).
- **NAS Get System Selection Preference** shows FULL band tables on both OSes (43 LTE bands +
extended). NAS has the data; DMS doesn't. The bridge is the platform MBN activation.
- **Android rmt_storage**: reads EFS on boot (no writes), writes 2MB to modem_fs2 AFTER going
online. Normal operation — not initial provisioning.
- **pmOS rmtfs journal**: empty (rmtfs doesn't log successful reads/writes). No errors either.
### Fix path
1. Boot Android to let QCRIL re-load the platform MBN into PDC (it was there before activation)
2. Boot pmOS — do NOT touch PDC, do NOT run `qmicli` activation
3. Verify: `qmicli --pdc-list-configs=platform` shows SS-LA-SDM845 as Active
4. Verify: `qmicli --dms-get-capabilities` shows `Networks: 'gsm, wcdma, lte'`
5. Let ModemManager drive enable→online (NEVER manual qmicli online)
6. If platform config is still Inactive after Android: need to fix `qmicli --pdc-load-config`
segfault, or write a minimal MBN loader using libqmi directly
### What we have on device
- rmtfs: read-write (`-P -s`), serving correct EFS partitions via partlabels
- tqftpserv: bind-mounted to `/mnt/vendor/persist/rfs` (persist partition)
- pd-mapper: installed and running from boot
- MBN files: 179 files (11MB) at `/mnt/vendor/persist/rfs/readonly/vendor/mbn/`
(extracted from Android `/vendor/rfs/msm/mpss/readonly/vendor/mbn/`)
- MCFG version: `g845-00194-220517-B-8604334`
- Rogers Canada MBN: `mcfg_sw/generic/NA/Rogers/Commercial/CA/mcfg_sw.mbn`
- SDM845 LA/SS platform MBN: `mcfg_hw/generic/common/SDM845/LA/SS/mcfg_hw.mbn` (56KB)
- Kernel: building from `sdm845/7.1-dev` HEAD (`85f1df2a`), config fixed to `sdm845`
### Android vs pmOS service comparison
| Component | Android | pmOS | Notes |
|-----------|---------|------|-------|
| EFS storage | rmt_storage r/w | rmtfs `-P -s` r/w | Parity |
| TFTP | tftp_server → vendor/rfs | tqftpserv → persist bind | Parity |
| pd-mapper | running (maps PDs) | running | Parity |
| RIL | QCRIL (`libril-qc-hal-qmi.so`) | ModemManager | **Missing PDC provisioning** |
| ipacm | running (`/dev/ipa` FDs) | not installed | Data path only |
| Keymaster | decrypts persist blobs | none | IMEI encrypted at rest |
| QMI services | comparable (50+) | comparable (50+) | Parity |
### (superseded) earlier frontiers
**Read-only rmtfs theory (06-19 #2):** rmtfs read-only prevented modem from persisting EFS
writes. Fixing to read-write didn't resolve DeviceNotReady. EFS r/w is necessary but not
sufficient — the modem needs valid MCFG in EFS first, which comes from PDC provisioning.
**QMI-handshake theory (original):** "MM doesn't send QCRIL's init sequence." Dead —
same MM binary works on OnePlus 6. The missing sequence is PDC MBN loading, not a
handshake gap. Other SDM845 devices likely have platform config pre-loaded from factory.
**fsg_oem theory:** Modem requesting `modem_fsg_oem_1/2`. Patched rmtfs to serve them.
Modem never actually requests these when EFS is properly provisioned. Symptom, not cause.
**XPU violation theory:** Modem crashes twice on some boots. Even on clean (0-crash) boots,
DeviceNotReady persists. Not the blocker.
## Useful commands (pmOS)
```sh
qrtr-lookup | awk '$1==11' # UIM service → note Port
qmicli -p -d qrtr://0/<port> --uim-get-card-status
qmicli -d qrtr://0 --dms-get-operating-mode
qmicli -d qrtr://0 --dms-set-operating-mode=online
qmicli -d qrtr://0 --dms-get-capabilities # zero rates = RF uninit
# clean modem re-init (rmtfs runs -P -s read-write; see 2026-07-02 correction):
echo stop | sudo tee /sys/class/remoteproc/remoteproc2/state ; sleep 3
echo start | sudo tee /sys/class/remoteproc/remoteproc2/state
# preferred: QMI Set Operating Mode = RESET (4) reinitializes without touching UIM/provisioning.
```
Forcing a remoteproc reset clears UIM provisioning; `msm-modem-uim-selection` only re-runs at
boot. (The old note here claimed `-r -P -s` read-only was required for re-provisioning — stale
and wrong; rmtfs runs `-P -s` read-write live and the modem re-provisions fine across a
stop→start. For a clean reinit that does NOT drop provisioning, use QMI RESET, not remoteproc.)
## Cross-references worth mining
- **Working SDM845 (same SoC):** OnePlus 6/6T (enchilada/fajita), Poco F1 (beryllium),
SHIFT6mq (axolotl) — data + SMS work via the same mainline kernel + pmaports packages.
Their *AP-side* config is already in `~/pmos/pmaports`; nothing to pull. Their working
firmware/EFS is device-specific and not transplantable.
- **Pixel 3a (sargo) / 3a XL (bonito):** the closest *Google* analogs with modem support in
pmaports (SDM670). Best target if a Google-specific RIL quirk is suspected.
- **Android on this exact device** is the cleanest reference: same silicon, modem works.
Diff blueline-Android vs blueline-pmOS, not against OnePlus.
## TOP LEAD 2026-07-02 #16 — locally-initiated DIAG_CTRL/DIAG_DATA open CRASHES the modem (glink_channel_migration assertion), repeatably
**Real, causal, previously-unknown finding — not a theory.** Following up #15's ack-but-no-reciprocal-
open result: tried the same locally-initiated open against `DIAG_CTRL` (not just `DIAG_DATA`).
Confirmed via `dmesg -T` (not my own client's stderr, which kept getting lost across SSH session
churn — the KERNEL'S OWN log line is the ground truth here): `rpmsg rpmsg0: failed to open DIAG_CTRL`
fires cleanly at the expected 5s mainline timeout, exactly as designed. But roughly 18-40 seconds
LATER, every single time this was attempted (three separate occasions, two of them minutes apart on
different modem boots), the modem hard-crashes:
```
qcom-q6v5-mss 4080000.remoteproc: fatal error received: [disable_ssr]glink_channel_migration.c:602:
[IST336]Assertion status == GLINK_STATUS_SUCCESS failed
remoteproc remoteproc3: crash detected in 4080000.remoteproc: type fatal error
remoteproc remoteproc3: handling crash #N in 4080000.remoteproc
```
This is a **repeatable causal chain**, confirmed across crashes #5 and #6 in the same dmesg buffer,
both immediately following a `failed to open DIAG_CTRL` line, both with the identical
`glink_channel_migration.c:602` assertion. `glink_channel_migration.c` is the FIRMWARE's own glink
channel-transport-migration logic (this is a modem-firmware-side log line, reported via SFR/fatal
error, not an AP-side kernel message) — meaning the incomplete/abandoned local-open handshake left
channel/migration bookkeeping in a state the modem's own firmware asserts on and can't tolerate.
**Practical, urgent implication**: DO NOT locally-initiate an open for `DIAG_DATA` or `DIAG_CTRL`
casually or repeatedly during further investigation — it reliably crashes the whole modem MSS a
short time later, disrupting whatever else was being tested (this crashed WLAN's firmware too, via
the shared XPU/hyp-assign path, matching TOP LEAD #11's earlier finding about shared secure-memory
fragility). If this path is pursued further, it needs to be done RARELY, with full crash-recovery
tolerance, ideally on a boot with nothing else running that a modem SSR would disrupt.
**What this actually tells us, functionally**: the modem's glink/migration layer is NOT robust to an
AP-initiated open attempt that times out — downstream never hits this because (a) it doesn't time out
in 5s (event-driven, no hard deadline, per #15's downstream comparison) and/or (b) Android's DIAG
daemon presumably completes successfully before any abandonment would occur. This is now a second,
independent data point (alongside offline-vs-LPM, #14) that mainline's bring-up leaves this modem's
firmware in a MORE FRAGILE state than Android's does — not just "missing a feature" but "structurally
less tolerant of imperfect AP-side timing/sequencing." Worth weighing against #14's LPM finding: both
point at the same underlying idea -- whatever Android's fuller bring-up sequence does before touching
these channels, mainline's minimal sequence skips, and the modem firmware was never hardened against
being touched without it.
**Honest gap, not yet resolved**: unclear if the crash is specific to DIAG_CTRL/DATA's particular
migration-eligible status, or would happen for ANY locally-initiated open of ANY channel abandoned via
timeout. Worth testing later (carefully, expecting a crash) against a channel we know is NOT normally
migration-sensitive, to isolate whether this is DIAG-specific or a general "don't abandon a local
open" hazard on this modem.