PAF becomes saf/device (history kept), STATE.md dissolves into saf/state.md with the dated era archived, the substrate SAF moves up from souveraine, and every agreement points at saf/INDEX.md and nowhere else. one map, nothing to remember
61 KiB
PAF · SLPI bring-up (sensor island) — SOLVED 2026-07-10
Status update — 2026-07-25: SLPI recovers, hexagonrpcd-sdsp does not
Sequencing regression, live for hours before it was noticed. The CHRE fatal documented below recurred on an ordinary running system (not a cold boot) at 06:04:33:
kernel: qcom_q6v5_pas 5c00000.remoteproc: fatal error received:
err_qdi.c:456:EF:sensor_process:1:fatal_error.cc:22:CHRE
fatal@platform_sensor_manager.cc:158
remoteproc remoteproc2: crash detected in slpi / handling crash #1
remoteproc remoteproc2: remote processor slpi is now up
remoteproc recovery worked — remoteproc2: slpi = running afterwards. But
blueline-hexagonrpcd-sdsp.service took
Could not fetch next FastRPC message: Broken pipe and "Deactivated
successfully": systemd read a clean exit and did not restart it. SLPI came
back; the HexagonFS daemon that serves it did not. The unit sat inactive (dead) while blueline-hexagonrpcd-adsp (audio PD) stayed up, which is why
audio was unaffected and the loss was invisible.
Symptom, ~100s apart, restart counter climbing to 5:
iio-sensor-proxy: 'registry' sensor unavailable, is hexagonrpcd running?
net.hadess.SensorProxy ->
HasAccelerometer: false HasAmbientLight: false HasProximity: false
Every sensor was down — no autorotate, and blueline-proximity-lock kept
running against a proxy with nothing behind it. A standalone consumer of a
dead sensor is indistinguishable from one whose sensor is merely quiet, so
nothing reported the outage.
Fix belongs in rootfs-overlay/, not on the device. The recovery order
is the bug: SLPI restarts itself, so its filesystem daemon must follow it
(Restart= plus a binding/ordering relationship that survives remoteproc
crash-recovery, not just boot). A fix applied live is erased by the next
reprovision — see SouveraineOS/docs/tasks/25-one-repo-all-packages.md.
Consumed-by note: this was the concrete argument for evidence source health
being device state, and it is now built — SouveraineOS/docs/DEVICE-STATE-MACHINE.md
§10 (2026-07-26). A source that reported and then went silent for 90 s is
recorded down, not quiet, and every forensic snapshot taken during the outage
is stamped sensors_degraded. The half that made it possible is here:
blueline-proximity-lock re-sends its last reading every 30 s, because
monitor-sensor only emits on change and a phone on a table is otherwise
indistinguishable from a CHRE fatal. (TASK-26 itself was never written; the
requirement outlived the placeholder.)
Status update — 2026-07-20: iio-sensor-proxy is always-on now
The wake-lockup caveat below drove blueline-autorotate to start/stop the
libssc client on demand — that hazard is confirmed gone. iio-sensor-proxy
now runs continuously from boot via blueline-sensors-enable.service
(After=blueline-hexagonrpcd-sdsp.service), and AutoRotateToggle in the
shell only starts/stops the rotation daemon, not the sensor proxy. Proximity
(TMD2725) is live and consumed by blueline-screen-toggle (DT2W gate) and
the new blueline-proximity-lock user service (blanks an already-on locked
screen on proximity-near). See touch.md and README.md's Tap-to-wake row.
The historical on-demand/no-autostart record below is superseded by this but
left as-is per PAF doctrine — see docs/tasks/paf-reconciliation.md.
Reliability caveat — 2026-07-11: libssc wake lockup, root cause unproven
New, separate cold-boot incident — 2026-07-13: delayed CHRE sensor-manager fatal
On Casey's confirmed unplugged cold boot, SLPI came up at kernel monotonic
17.692s and then, with iio-sensor-proxy.service inactive, crashed at
331.667s: err_qdi.c:456:EF:sensor_process:1:fatal_error.cc:22:CHRE fatal@platform_sensor_manager.cc:158. Remoteproc recovery immediately booted
SLPI again. This is not the historical immediate sar.cc:27 signature and
must be tracked independently; its five-minute delay makes a runtime
CHRE/SEE-sensor-manager timeout or deferred initialization path plausible.
Active firmware was /lib/firmware/qcom/sdm845/Google/blueline/slpi.mbn, SHA-256
83e16ce8bffd8e4b963280dd30fa782becd005b40a3fd4dd6199093f98b82826.
The raw journal sequence is retained by journalctl -k -b; do not clear the
kernel ring before copying a newly observed fatal into this record.
Source correlation (2026-07-13): AOSP CHRE's SLPI
platform_sensor_manager.cc explicitly schedules crashAfterSensorNotFound
for 5 * 60 seconds when required SEE sensors are missing at initialization.
The observed delay is 313.975 seconds from SLPI-up to fatal. This pins the
failure to the required-SEE-sensor discovery set, rather than a random runtime
crash or the audio stack. Identify the absent SUID(s) from the live SSC/SEE
registry before touching firmware or remoteproc state.
Required-SUID sweep (2026-07-13): The bounded read-only
tools/slpi-suid-probe found SUIDs for
sensor_temperature (LIS2MDL), accel/gyro (BMI160), mag (LIS2MDL),
pressure (BMP380), and ambient_light/proximity (TMD2725). amd alone
returned no matching SUID. The preserved Pixel Android sns_amd group is
byte-for-byte reproduced by the OEM/soc-321 source wrapper. The reverse-file
trace then proved the failure mode: after startup the DSP reads the added
runtime config/sns_amd.json, opens the valid static group with write/truncate,
and emits no payload for that group, leaving it zero bytes. This is not a host
permission issue (the daemon successfully writes many other groups). Pixel's
captured Android vendor config has no sns_amd.json, so the staged repair keeps
the static group but omits that dynamic config input. Cold-boot result: the
served group remained byte-identical to Android (239 bytes) and the DSP read
it without a write-open, yet SLPI still fataled at 322.907s. Therefore the
truncation path was a genuine regression but is not sufficient to restore the
AMD SUID or satisfy CHRE. The probe was made null-safe after libssc returned an
unavailable data type without a GError; no SLPI crash was caused by the probe.
AMD provisioning correction (2026-07-13): Qualcomm's public split runtime
pair (sns_amd.json plus sns_amd_sw_enabled.json) was also tested. The
correct vendor-input location is config/registry/sns_amd.json, never the
served flat registry and never a physical registry/registry tree. Its trace
is decisive: SEE opens that input, opens the valid static sns_amd group in
write mode, allocates write-FD 256, and closes it without an fwrite; it
then writes sns_reg_config normally. O_TRUNC in hexagonrpcd therefore
turned the valid 239-byte Android-identical group into zero bytes without any
DSP replacement payload. A clean 0014-lazy-truncate-reverse-writes.patch is
installed for the next boot: w opens are truncated only at their first actual
fwrite, while append mode is unchanged. This preserves no-write static
groups, but it is a prerequisite—not yet evidence that AMD will publish.
Validation: the subsequent boot kept sns_amd at the exact 239-byte
Android SHA-256 (4756a6…2959d2) through the complete no-write open/close
sequence, yet CHRE still fataled at 325.595s in
platform_sensor_manager.cc:158. Registry placement and host-side truncation
are therefore fixed prerequisites, not the remaining AMD-publication cause.
Required-SUID timeout cleared (2026-07-13): the public SDM845 AMD split
was retried only after the no-write truncation repair: dynamic
config/registry/sns_amd.json supplies parameters while
sns_amd_sw_enabled.json supplies enablement; the preserved 239-byte Android
group remains the static registry seed. SEE opened both dynamic files,
opened sns_amd in write mode without sending a payload, and the lazy-truncate
daemon retained the exact reference hash. At 375 seconds uptime—well past the
former 325-second platform_sensor_manager.cc:158 deadline—SLPI remained
running, hexagonrpcd remained active, and no fatal was logged. This is the
first boot to clear CHRE's required-sensor watchdog. Do not immediately start
iio-sensor-proxy/ssccli: their libssc client connection is the separately
documented wake-lock hazard and needs a bounded follow-up test.
Bounded client-path validation (2026-07-13): On that stable boot,
iio-sensor-proxy was started explicitly (the package's boot-time udev
autostart remains masked). It discovered the SSC TMD2725 proximity and ambient
light sensors plus BMI160 acceleration. The desktop D-Bus API successfully
claimed and delivered acceleration (face-up) and ambient-light samples
(~14 lux); proximity also reports normally. The device rule now explicitly
selects ssc-accel ssc-proximity ssc-light on demand, while preserving the
no-autostart policy. Direct ssccli reads also returned live TMD2725 light,
BMI160 acceleration, and LIS2MDL magnetometer values. ssccli's compass
client alone still returns UNKNOWN. SLPI and hexagonrpcd remained running,
with no CHRE fatal, after these clients. This validates the practical sensor
path without reclassifying the historical boot-time libssc incident as solved.
Service-loop regression fixed (2026-07-13): After the one legitimate
five-minute fatal at 326.652s, hexagonrpcd correctly exited on its broken
FastRPC pipe with status 0. Its unit unfortunately used Restart=always, so
systemd reattached during remoteproc recovery six times in seven seconds. Each
reattach immediately crashed SLPI in SNS_REG_INIT:39, then the service hit
systemd's start limit. This was a userspace retry-policy bug, not six
independent firmware failures. The deployed and overlay unit now uses
Restart=on-failure with a five-second delay; no service/remoteproc restart
was performed while installing it. A trace-enabled hexagonrpcd was also
deployed for the next cold boot: HEXAGONRPC_TRACE_FILES=1 correlates virtual
open paths with the previously opaque Could not read file: Is a directory
FD. The binary and service changes take effect on the next genuine cold boot.
DMA-heap permission regression fixed (2026-07-13): The documentation said
the FastRPC allocation heap was owned by root:fastrpc 0660, but the actual
cold-boot device node was root:root 0600: the overlay had a rule for
/dev/fastrpc-* only and never contained the claimed dma_heap rule. Added
11-fastrpc-dma-heap.rules, narrowly assigning only
/dev/dma_heap/system to fastrpc:0660; firmware-reserved heaps remain
root-only. It was installed live with udevadm trigger and verified as the
fastrpc user before any new SLPI action. This removes a genuine local
permissions blocker for non-root FastRPC clients, but does not by itself
explain an EPERM returned from an already-open DSP RPC session.
CORRECTION — cold-boot validation failed (2026-07-13): The following
paragraph describes a staged-path hypothesis, not a completed repair. On
the next cold boot the registry was readable but SLPI again fataled at 323.7s.
The corrected root-owned listener was then installed and enabled; on its first
validation boot it attached before SLPI's delayed window but the fatal still
occurred at 326.97s. The listener deployment is necessary infrastructure, not
a sensor fix. Its remaining sns_reg_config virtual-path error points to
hexagonrpcd's mapping, and CHRE required-sensor discovery remains open.
Concrete boot-path repair staged and installed (2026-07-13): before the
fatal, hexagonrpcd logged that
/mnt/vendor/persist/sensors/registry/sns_reg_config did not exist, even
though the staged registry under /usr/share/qcom/sensors was complete. All
required types subsequently resolve through SSC (BMI160 accel/gyro and
motion/stationary; LIS2MDL magnetometer; BMP380 pressure; TMD2725 ambient
light/proximity), so this is an initialization-path race rather than a
permanently missing physical sensor. The enabled
blueline-sns-registry-stage.service now creates the guarded persistent link
/mnt/vendor/persist/sensors → /usr/share/qcom/sensors before
hexagonrpcd-sdsp.service. The next unplugged shutdown/power-on is the
validation: retain the kernel journal and confirm that the delayed fatal does
not occur after five minutes.
Correction — the tracked VFS fix was not in the installed binary
(2026-07-13 evening): The fresh cold boot again logged
Could not open /mnt/vendor/persist/sensors/registry/sns_reg_config, despite
the flat staged host file existing. Inspection of /usr/local/bin/hexagonrpcd
and ~/hexagonrpc/hexagonrpcd/rpcd_builder.c proved that the live
hexagonrpc-blueline 0.4.0-2 binary had only the old sns_reg_version map;
the recipe tracked 0008 but its installed artifact did not contain it. Applied
0008 to the live source, rebuilt, and installed a backed-up replacement. The
new binary embeds maps for sns_reg_config, sns_reg_version,
sns_reg_ctrl, file1, and file2; /usr/share/qcom/sensors/registry/
remains physically flat (no registry/registry). No DSP lifecycle action was
taken. The next unplugged shutdown/power-on must confirm the initial ENOENT is
gone and then observe CHRE past its five-minute required-sensor timeout. The
package database still says 0.4.0-2; rebuild/install the tracked pkgrel=3
package before treating this live binary repair as image-persistent.
Cold-boot validation — VFS control maps fixed; CHRE still missing a required
SEE sensor (2026-07-13 evening): On the next unplugged shutdown/power-on,
SLPI reached running at 13.014s. hexagonrpcd successfully opened virtual
registry/sns_reg_config, sns_reg_version, and sns_reg_ctrl; the former
sns_reg_config ENOENT was absent. At 327.479s (314.465s after SLPI-up), CHRE
again fataled at platform_sensor_manager.cc:158, then remoteproc recovered
SLPI and the daemon exited cleanly. This conclusively fixes the VFS-control
prerequisite without fixing required-SUID discovery. Remaining trace noise is
benign: testsig*.so is absent test firmware, and the reported EISDIR is
CHRE's expected read-then-opendir probe of virtual registry/registry.
SLPI itself remains up; this is not a return of the sar.cc crash. Toggling
sensors off/on reproducibly identifies the libssc/iio-sensor-proxy SSC QMI
connection as the trigger for a wake lockup: kernel and SSH survive, display
commits stop completing without DSI/DPU errors, FTS becomes calibration-blind,
and GENI I2C GPI DMA times out concurrently. This narrows the mechanism to an
AP-side shared-bus/power interaction on SLPI wake. An rpmh/interconnect vote
race is the leading hypothesis, not a proven root cause.
HACK shipped: SLPI stays booted, but the auto-rotate toggle starts/stops its iio-sensor-proxy/libssc client instead of leaving that client enabled from boot. It reduces exposure only; it is not a fix. Netconsole and pstore_blk are armed for an instrumented reproduction before any permanent change is attempted.
Boot-safety correction, 2026-07-11: the initial HACK was ineffective at
boot: /usr/lib/udev/rules.d/80-iio-sensor-proxy.rules starts the proxy on
every fastrpc-sdsp add event via SYSTEMD_WANTS. The overlay now supplies a
comment-only /etc/udev/rules.d/80-iio-sensor-proxy.rules, which shadows that
vendor rule while leaving the explicit polkit/toggle start path intact. Live
udevadm test confirmed /etc rule precedence and the running boot-started
proxy was stopped. The next reboot must still verify that it remains off.
Instrumented reproduction, 2026-07-11 10:45 EDT: netconsole over the USB
gadget was configured dynamically (usb0 172.16.42.1:6665 → laptop
172.16.42.2:6666) and packet-capture verified before the proxy was started.
With the proxy active, physical sleep/wake reproduced the stale-frame lockup:
the glass remained at 08:02 while SSH, low system load, Hyprland, and
Quickshell remained alive. No panic or pstore record occurred. Starting 28
seconds after the proxy-start marker, the FTS driver emitted repeated invalid
Mutual/Self Strength calibration events; no DSI/DPU error, GENI timeout, or
rpmh/ICC diagnostic was logged in this run. Netconsole captured only the
explicit console-level start marker, so it proves the transport but not the
missing low-priority driver messages. The proxy was stopped afterwards. This
is further trigger evidence, not proof of the rpmh/interconnect hypothesis.
Next capture: raise the temporary console loglevel and add narrowly scoped
rpmh/ICC tracing before another deliberate reproduction.
ARCH STATUS (2026-07-10 afternoon): THE sar.cc WALL IS DOWN. SLPI runs.
SLPI boots, stays up (0 fatals, was deterministic-fatal at ~10s), SEE fully
initializes (node 9 publishes svc 400 + services that never appeared in the
crash era), the accelerometer streams real data to userspace (ssccli), and
iio-sensor-proxy reports HasAccelerometer=true over D-Bus. Boot-started by
blueline-slpi.service (rootfs-overlay).
Root cause of sar.cc:27 (found by source archaeology, kernel 25d56ec84):
the SSC firmware's remote_proc_state sensor does
smp2p_register(SMEM_APPS, "sleepstate") in its init thread and NEVER
publishes if that fails (Qualcomm source: Rivko/android-firmware-qti-sdm670,
adsp_proc/ssc/sensors/remote_proc_state/). Mainline's sdm845 smp2p-slpi node
has no "sleepstate" entry, so registration failed, remote_proc_state never
published, and CHRE's SeeHelper::init exhausted its ~10s findSuidSync retry
budget on it (AOSP chre platform/slpi/see/see_helper.cc — initCalSensors
passes, initRemoteProcSensor fails) → "Failed to initialize ASH SEE helper"
→ chre_utils FATAL at sar.cc:27 → whole sensor_process aborts. The SAR
module was only ever the messenger. FIX: sdm845.dtsi sleepstate smp2p entry
- new drivers/soc/qcom/smp2p-sleepstate.c asserting bit 12 (= SMP2P_AWAKE 0x1000 in SSC source; = downstream PROC_AWAKE_ID msm-4.9 smp2p_sleepstate.c), dropped across suspend. CONFIG_QCOM_SMP2P_SLEEPSTATE=y.
Dead ends, so nobody repeats them: (1) qrtr ns announce-replay patch (3fe96cf2d, REVERTED in 94ac4a5e1) — replaying all nodes' services on HELLO wedged node 9's own service registration (flow control on the ns worker); sar.cc never needed modem services. (2) The whole 2026-07-08 elimination list stands (registry, socinfo, chrecd, proxy-power, sessions — all real fixes, none the wall). (3) Registry-diff vs LOS persist: the 42 missing entries are wrong-device pollution + factory cal, not the cause.
Five-fix SNS_REG_INIT chain (hw_platform OEM socinfo, module tree, chrecd,
sscregistrygen registry, socinfo sysfs) remains prerequisite and correct.
Sensor stack persists across re-provision since cd82f29. AP-side gesture
dt2w works independently (touch.md); the SLPI gesture-nanoapp path is now
UNBLOCKED for future work (CHRE runs). Auto-rotate WORKS on glass
(2026-07-10 ~13:00: blueline-autorotate + ACCEL_MOUNT_MATRIX udev fix,
see rootfs-overlay). Remaining: sx932x "sar" SUID still absent
(grip sensor hardware probe — separate, non-fatal now).
Everything below is the pmOS-era record — the archaeology that fed the
chain above; capture docs it cites live in evidence/.
2026-06-30 — chronic RTC-write flood found + fixed; wrong listener used in a failed test; dma_heap permissions fixed
Context: a fresh SLPI one-shot attempt (manual sdsprpcd sensorspd + recovery=disabled
echo start, not viablueline-slpi-ready.service) wedged the phone. Investigating the wedge surfaced three unrelated, real bugs — none of them SLPI firmware behavior — that were corrupting every capture attempt and blocking the correct listener from working at all.
1. Chronic disallowed SPMI write to sid=0, addr=0x6046 flood (~40/sec, continuous,
present even with SLPI offline) — unrelated to SLPI, was drowning every capture.
Traced to drivers/spmi/spmi-pmic-arb.c:1364 (pmic_arb_offset_v5): a hardware EE-ownership
check the AP fails every time. sid=0/0x60xx is the pm8998 RTC peripheral
(pm8998.dtsi:104, rtc@6000). Root cause: drivers/rtc/rtc-pm8xxx.c:622-631 force-overrides
allow_set_time = true regardless of the DT allow-set-time property, on the premise
(comment at :624-628) that "the modem reads real wall-clock time during RF-init" via this
write. That premise is wrong — cross-referenced against LineageOS on the same silicon:
los-blueline.dts:2371 sets qcom,qpnp-rtc-write = <0x00> (write disabled),
qpnp-rtc.c:474,587-588 defaults to read-only ops and only switches to read-write if that
property is true, and init.sdm845.rc:233-237 never touches rtc at all — only
start time_daemon. Android does not write this register either; it can't, for the same
hardware-ownership reason. The actual (real, working, already-verified — see file.txt
2026-06-30, "FIRST ACCEPT... modem User time is now real wall-clock") mechanism for feeding
the modem wall-clock time is tools/blueline-timed.service, a QMI TIME-service-22 push
straight to the modem (modem-time-set.py), completely unrelated to the PMIC RTC. The
rtc-pm8xxx.c patch never achieved its stated goal either way: __pm8xxx_rtc_set_time()
(:294-343) fails at its very first register access (alarm_ctrl, same address) every call,
before ever reaching the actual time-write registers.
Fix: reverted the force-true override (falls through to the driver's existing
pm8xxx_rtc_update_offset() software-offset path, which is what LOS's read-only mode is the
equivalent of); dropped allow-set-time; from sdm845-google-blueline.dts. Kernel rebuilt
2026-06-30 ~21:56.
2. The SLPI wedge test used the wrong listener with no ordering — not a firmware crash we
observed correctly. device-google-blueline ships blueline-slpi-ready.service, which
Requires=+After= blueline-sns-registry-stage.service and hexagonrpcd-sdsp.service
— exactly the registry-staged/listener-ready ordering the S6 session below already proved is
required. The failed manual test instead started the proprietary sdsprpcd (from the
fastrpc apk, installed this session) directly, with zero registry staging and zero
wait-for-listener. Next SLPI attempt: systemctl start blueline-slpi-ready.service,
not a hand-rolled sequence. All three services were disabled (not masked) on this
rootfs — never actually run, not deliberately blocked.
3. /dev/dma_heap/* was root:root 0600, silently blocking hexagonrpcd-sdsp.service
(User=fastrpc) from allocating FastRPC memory at all. 60-fastrpc.rules's fix for this
(RUN+="/bin/setfacl -m g:fastrpc:rw $env{DEVNAME}", and only for the system heap, not
adsp_region/default_cma_region/reserved) silently no-ops on every boot because
setfacl/the acl package was never installed. Confirmed live: fastrpc user got
Permission denied opening /dev/dma_heap/system before the fix, clean open() after.
Fixed both live (chgrp fastrpc; chmod 0660 on all four) and durably in the rules file
(replaced the setfacl RUN directive with GROUP="fastrpc", MODE="0660", matching the
fastrpc-* device rules already in the same file) — not yet verified across a real cold
boot.
4. Separately, /vendor/dsp/{adsp,cdsp,sdsp} didn't exist on the rootfs at all (needed
by adsprpcd/cdsprpcd/sdsprpcd/hexagonrpcd alike). The live vendor_b partition
(/dev/sda10) mounts but reads as directory-corrupted via both the kernel VFS and
debugfs -R "ls /" — likely a raw-mount-bypasses-dm-verity artifact, not real corruption of
the working LOS install (left untouched, not fsck'd). Used the already-extracted, genuinely
readable copy from an earlier session, pixel3-work/los-vendor/vendor.img (531MB ext2,
volume name "vendor"), and copied just dsp/{adsp,cdsp,sdsp} (17MB) to /vendor/dsp/ on
the phone. This did not fix adsprpcd_audiopd (its real failure is Qualcomm QDI status
0x80000414, not literally ENOENT despite the log text — needs an audio-PD servreg
registration the vendor tree alone doesn't provide; out of scope, re-masked), but the tree
is now in place for whatever hexagonrpcd/SDSP actually needs from it.
**Net effect: the next SLPI attempt is the first one with quiet logs, the correct listener
- ordering, correct dma_heap permissions, and the vendor DSP tree all present simultaneously.** No prior attempt (including tonight's wedge) had all four.
2026-06-30 (cont.) — first clean-ish trace: two racing FastRPC-sdsp consumers, one legit, one not
After the RTC fix (kernel flashed, rebooted, SPMI flood confirmed 0 hits — first quiet
boot ever), ran a real test: staged the registry (blueline-sns-registry-stage.service,
179 files), then a poller script raced hexagonrpcd -f /dev/fastrpc-sdsp -d sdsp -s onto
the device the instant it appeared (/dev/fastrpc-sdsp doesn't exist pre-boot; neither
hexagonrpcd nor any udev rule retries for it — the poller was necessary, see script in
tools/ history / this session). echo start on remoteproc2. Captured via
journalctl -f to a persistent file (works — unlike tail -f /dev/kmsg, which silently
dies on the ring-buffer overflow the RTC bug used to cause; this is now the standard
capture method, not kmsg tail).
What actually happened, in order (journal-confirmed, slpi-test2/journal-live.txt):
remoteproc remoteproc2: powering up slpi→ boots clean, heap already preassigned (no-22, confirms the earlier adsprpc memory conversion still holds)./dev/fastrpc-sdspappears. Two things react simultaneously, not one:systemd[1]: Starting sDSP RPC daemon...→sdsprpcd.service, auto-started by60-fastrpc.rules'sKERNEL=="fastrpc-sdsp" ... SYSTEMD_WANTS+="sdsprpcd.service". This is the Qualcomm proprietary daemon from thefastrpcapk installed earlier this same session — an unintended side effect of installing that package to getsdsprpcdfor testing. It attached first (attachguestoson domain 2 succeeded, openedadsp_default_listenerhandles), then hitError 0x4e ... errno Operation not permittedon an RPC invoke.hexagonrpcd(the poller-launched, intended listener) tried immediately after and gotCould not attach to FastRPC node: Operation not permitted— plausibly becausesdsprpcdhad already claimed the attach slot, not necessarily a firmware-side rejection ofhexagonrpcdspecifically.- A third thing also started at the same instant:
systemd[1]: Starting IIO Sensor Proxy service...(iio-sensor-proxy.service, pre-existing base-image package, installed 2026-05-08, unrelated to tonight — see below).
- ~1s later:
qcom_q6v5_pas ...: watchdog received: SFR Init: wdog or kernel error suspected.→remoteproc2crashes.
sdsprpcd.service is a real bug (masked) — but not a new one tonight. Checked the
history properly rather than assume: the 2026-06-29 session also installed the fastrpc
apk and ran sdsprpcd sensorspd manually as its one-shot test — and that is the exact
command HANDOFF's own notes say caused that session's wedge too ("the same one-shot
sdsprpcd sensorspd + SLPI start caused SSH/network/UI to wedge"). file.txt's later
session was also checking for sdsprpcd specifically. So sdsprpcd-as-manual-test is a
repeated pattern across at least three sessions, not something introduced tonight — this
session continued that same habit (reinstalling the fastrpc apk to get sdsprpcd back
after the rootfs rewrite lost it) before realizing hexagonrpcd was the actually-intended
listener. hexagonrpcd, and the blueline-slpi-ready.service/hexagonrpcd-sdsp.service
ordering, come from device-google-blueline (a separate, independently-developed track
within this same project) and were never actually exercised in isolation until tonight's
sdsprpcd-masked re-test. systemctl mask sdsprpcd.service done this session — first
time this collision has been diagnosed rather than just repeatedly hit.
iio-sensor-proxy.service is NOT a bug and should NOT be masked long-term (briefly
masked mid-session, then unmasked once this was understood — leave it unmasked).
It's built against libssc.so.2 (qmi_client_ssc_control, "SSC QMI client") — a genuine
QMI Sensor SubSystem Client, i.e. exactly the downstream consumer this whole SLPI
effort is trying to feed ("no SNS QMI service exposed" has been the standing symptom
description). Its reacting to SLPI coming up is systemd/udev doing its job correctly, not
a coincidence. The real fix here is ordering, not suppression: it needs to start
after SLPI/hexagonrpcd are confirmed stable, not the instant the remoteproc powers up —
right now nothing enforces that, so its QMI connection attempt lands during the same
fragile registry-init window the S6 finding already identified as timing-critical, on top
of the sdsprpcd collision. Untested whether iio-sensor-proxy alone (with sdsprpcd
masked) still perturbs the window enough to matter — that's the next isolated test.
Next step: re-run the identical poller+hexagonrpcd test with sdsprpcd.service now
masked. If iio-sensor-proxy still auto-fires and coincides with a crash, that tells us
the ordering fix is needed for real; if hexagonrpcd attaches cleanly this time, the
sdsprpcd collision was the whole story and today's Operation not permitted /
SFR Init crash was never a hexagonrpcd/firmware problem at all — just two proprietary
and community daemons fighting over one singleton resource.
2026-06-30 (cont. 2) — isolated re-test: sdsprpcd wasn't the story either. The SAR/CHRE fault is real and reproduces clean.
Re-ran the identical poller+hexagonrpcd test with sdsprpcd.service masked (previous
section). iio-sensor-proxy.service still auto-fired on SLPI coming up (as expected, it's
legitimate) but did not coincide with any crash this time — no evidence it perturbs
anything on its own.
Clean timeline (journalctl, unfiltered, this run):
22:25:26 remoteproc remoteproc2: powering up slpi
22:25:27 remoteproc remoteproc2: remote processor slpi is now up
22:25:27 [hexagonrpcd] Starting hexagonrpcd (INIT_ATTACH_SNS) on /dev/fastrpc-sdsp
22:25:27 [hexagonrpcd] Could not attach to FastRPC node: Operation not permitted
22:25:36 qcom_q6v5_pas ...: fatal error received: err_qdi.c:456:EF:sensor_process:1:sar.cc:27:chre_utils fatal
22:25:36 remoteproc remoteproc2: crash detected in slpi: type fatal error
No AP-level instability this time — remoteproc2 settled into crashed (recovery still
disabled, no relaunch spam) and the phone stayed fully reachable throughout. The hard
wedge in the first test tonight was very likely the sdsprpcd/hexagonrpcd collision
itself (two clients contending for one singleton attach), not simply "SLPI crashes." A
clean, isolated SLPI crash does not wedge the AP.
Conclusion: this is the same sar.cc:27 SAR/CHRE fatal documented in the "2026-06-29
live result" section above, now reproduced with everything else this session found and
fixed already in place — quiet logs (RTC fix), correct listener alone (no sdsprpcd
collision), correct dma_heap permissions, /vendor/dsp/ present, registry staged
(179→182 files). None of tonight's plumbing fixes were the actual blocker. They were
real bugs and worth fixing (nothing else could be tested cleanly until they were), but the
live frontier is exactly where it was before tonight: INIT_ATTACH_SNS returns EPERM,
and ~9s later the firmware faults inside CHRE's SAR path. Next real work is still the
already-identified target: convert the remaining Lineage sensor/SAR FastRPC path
(sscrpcd sensorspd equivalent, sensors PDR/service-location behavior, whatever makes
INIT_ATTACH_SNS attach cleanly) — see "Status / next" above. That work has not started;
tonight only cleared the ground for it to be tested honestly.
2026-06-30 (cont. 3) — the "proper" pmOS chain has a real, previously-undiagnosed timing bug
Per Casey's explicit direction: stop working around the shipped mechanism, make the
shipped mechanism itself work. The gap: hexagonrpcd-sdsp.service has
ConditionPathExists=/dev/fastrpc-sdsp, which can never be true at the moment systemd
evaluates it (the device doesn't exist until SLPI starts) — device-google-blueline never
shipped a udev rule to re-trigger it once the device appears. Added one, as a local admin
override (not yet upstreamed into the package):
# /etc/udev/rules.d/61-hexagonrpcd-sdsp.rules
ACTION=="add", SUBSYSTEM=="misc", KERNEL=="fastrpc-sdsp", TAG+="systemd", ENV{SYSTEMD_WANTS}+="hexagonrpcd-sdsp.service"
With this in place, systemctl start blueline-slpi-ready.service — the actual, sole,
intended entry point — runs end-to-end with no manual scripts at all: registry stages,
hexagonrpcd-sdsp.service auto-triggers via udev the instant the device appears (via a
blueline-hexagonrpcd-sdsp-wait.sh wrapper this project hadn't inspected before), SLPI
starts, hexagonrpcd attaches. This is real progress — the shipped mechanism now actually
runs as designed, on the current rootfs, for the first time.
But it lost the registry-init race — a real bug, not the earlier sdsprpcd collision.
Clean journal, unfiltered:
22:48:43 blueline-slpi-ready.sh: disabling recovery, starting SLPI once
22:48:43 remoteproc remoteproc2: powering up slpi
22:48:44 remoteproc remoteproc2: remote processor slpi is now up
22:48:44 blueline-hexagonrpcd-sdsp-wait.sh: starting hexagonrpcd on /dev/fastrpc-sdsp
22:48:44 hexagonrpcd: Starting hexagonrpcd (INIT_ATTACH_SNS) on /dev/fastrpc-sdsp
22:48:44 qcom_q6v5_pas: fatal error received: err_qdi.c:456:EX:sensor_process:1:SNS_REG_INIT:39:PC=b20226c0
22:48:44 remoteproc remoteproc2: crash detected in slpi: type fatal error
This is SNS_REG_INIT failing — the exact same registry-init ordering race the S6
session (2026-06-27) diagnosed and believed auto_boot=false had solved — reproducing
in the same second SLPI came up, before hexagonrpcd's attach could complete.
Compare to the earlier isolated test tonight (this doc, "isolated re-test" section):
manually racing hexagonrpcd onto the device with a raw 20ms-poll bash loop (no systemd
involved) won that same race reliably — SLPI ran a full ~9s before failing later, in
sar.cc:27:chre_utils (deeper into CHRE/SAR init, past registry-init). The manual
poller was faster and more reliable than the real systemd service-start chain.
hexagonrpcd-sdsp.service's startup path — udev event → systemd service activation →
blueline-hexagonrpcd-sdsp-wait.sh → exec hexagonrpcd → device open → INIT_ATTACH_SNS
— has enough cumulative latency to lose a race a tight bash loop wins.
FIXED, same session, same night. Checked ~/pmos/pmaports/device/testing/device-google-blueline/
(the actual local source for this package) and found the design was already correct — the
hexagonrpcd-sdsp.service drop-in (hexagonrpcd-sdsp-root.conf) already clears
ConditionPathExists, already redirects to a wrapper (blueline-hexagonrpcd-sdsp-wait.sh)
that runs a persistent while :; do wait for device; exec hexagonrpcd; wait for teardown; done loop — i.e. it was already starting before SLPI, via the normal
Requires=/After= chain in blueline-slpi-ready.service, exactly as the S6 finding
prescribed. (My first instinct — that a missing udev rule was the gap — was wrong; added
/etc/udev/rules.d/61-hexagonrpcd-sdsp.rules then confirmed via systemctl show hexagonrpcd-sdsp.service -p ConditionResult → yes that it was never needed, and removed
it again.) The actual bug was one number: the wrapper's poll interval,
blueline-hexagonrpcd-sdsp-wait.sh 0.2 (200ms), was too coarse to reliably detect
/dev/fastrpc-sdsp and get hexagonrpcd attached inside SLPI's registry-init window.
Changed to 0.02 (20ms, matching the interval that won the race in the manual poller
test above). Retested through the pure, unmodified entry point —
systemctl start blueline-slpi-ready.service, zero manual scripts — and it now reliably
reproduces the same deep sar.cc:27:chre_utils fatal crash tonight's best manual test
got, not the shallow SNS_REG_INIT race-loss. Fix committed to
hexagonrpcd-sdsp-root.conf in the local pmaports source (not yet built into a package/
flashed — currently live-patched onto the running rootfs only).
Net: the shipped mechanism is now correct and reliable. No more manual scripts, no
udev workaround, no competing daemons — systemctl start blueline-slpi-ready.service
alone gets to the real wall every time. The remaining problem is exactly what it's been
all along: INIT_ATTACH_SNS → EPERM, firmware faults in CHRE's SAR path ~10s later.
That's real kernel/source work (the Lineage sensor/SAR FastRPC conversion), not another
timing or plumbing bug.
(Also observed, unrelated background noise in this capture: qcril-autoselect PDC
Set Feature requests timing out repeatedly (Transaction timed out (2)) — pre-existing,
not investigated tonight, flag for later.)
Note: qrtr-lookup is genuinely absent from this rootfs (find / -xdev -iname 'qrtr-lookup*' finds nothing at all — not a $PATH issue). It's used routinely across
modem.md, tools/slpi-capture.sh, and tools/phone-slpi-sdsprpcd-collect.sh, so it was
clearly present and working in earlier sessions, but there is no apk package or
reference source for it anywhere in this project (qrtr-libs only ships the shared
library, not the CLI tool) — it must have been a manually-built binary on the pre-wedge
rootfs with no packaging record, same category of loss as the SLPI service masks. Needs
building from the upstream qrtr tools source (not currently cloned into references/)
to restore. Did not block tonight's test since it wasn't needed, but real and unsolved.
🟠 REFRAME 2026-06-29 (Casey) — the SLPI is a CANDIDATE ROOT-CAUSE for the modem 52, NOT a side-quest. The prior "SLPI proven independent of the 52" conclusion (
modem_consumer_chain_20260626.mdSESSION-3) is invalid. It was tested only with an SLPI that was either crash-looping OR merely remoteproc-not-crashing — never a functioning sensor stack (that doc itself notes: "qrtr-lookup empty — the SLPI never exposes an SNS QMI service before it dies"; "CHRE nanoapp crash INSIDE the SLPI firmware at init"). "Stop SLPI → still 52" only proves stopping doesn't worsen it: if the modem's RF-init (~12 s window) queries the SLPI for sensor/NV/registry data and gets silence, it parks OFFLINE identically whether the SLPI is crashing or stopped. The SLPI has never been healthy in any modem test. Path to actually test it: get the SLPI genuinely healthy —auto_boot=false+ start after the correct SDSP listener is ready + Lineage-compatible FastRPC/adsprpc behavior so the sensor stack initializes and SNS QMI exposes — then test the modem. That is the first real modem test with a functioning SLPI.
2026-06-29 late session — current SLPI frontier and phone state
Phone is stable only with SLPI held offline. Current handoff state:
- Kernel:
7.1.1-sdm845, packagelinux-postmarketos-qcom-sdm845-7.1.0_p20260629134644-r1.apk. - WiFi regression fixed live:
CONFIG_ATH10K_SNOC=m,CONFIG_QCOM_PD_MAPPER=m,wlan0present. - MSS/modem remoteproc running; modem still DMS op-mode
offline, HW restrictedno. - SLPI is offline, recovery disabled when touched.
blueline-slpi-ready.serviceandhexagonrpcd-sdsp.serviceare masked via/etc/systemd/system/*.service -> /dev/null. Do not unmask/start SLPI on boot until the next capture path is planned.
Conversions landed:
qcom_q6v5_pas:sdm845_slpi_resource_init.auto_boot = false.- Blueline Lineage/downstream adsprpc memory conversion:
fixed no-map
pil_adsp_mem@8c500000foradsp_pas, reusableadsp_regionfor adsprpc, SDSP FastRPC VMIDsHLOS + MSS_MSA + SSC_Q6 + ADSP_Q6, and a newqcom,msm-adsprpc-mem-regionhelper indrivers/misc/fastrpc.cthat preassigns the heap. - Qualcomm FastRPC userspace installed:
sdsprpcd,adsprpcd,libsdsp_default_listener.so. blueline-slpi-ready.shnow disables remoteproc crash recovery before its one-shot SLPI start.
What the tests proved:
- The old TZ/heap failure is fixed. With the Lineage adsprpc memory conversion, the boot log shows
the heap preassigned and SDSP FastRPC skipping duplicate hyp-assign. The previous
qcom_scm Assign memory protection call failed -22/SDSP fastrpc hyp-assign rejected by TZwall is no longer the current failure. - Crash recovery had been hiding first-failure evidence.
auto_boot=falseonly prevents PAS probe boot; remoteproc still relaunches after a crash unlessrecoveryis set todisabled. - Clean one-shot test before the SDSP capability-query patch:
sdsprpcd sensorspdwas started, SLPI recovery disabled, then SLPI started once. Result:remoteproc2becamecrashed. First useful kernel failure:qcom,fastrpc ... Error: dsp information is incorrect err: -1, followed bywatchdog received: SFR Init: wdog or kernel error suspected. - Kernel compatibility patch then added: for SDSP only, if the DSP utilities capability query is
rejected with
-EPERM, cache zero capabilities and continue. Rationale: Qualcomm userspace treats missing capabilities as "not supported"; Lineage downstream adsprpc does not make this optional query a hard sensors-PD gate. - One-shot test after that patch caused SSH/network/UI to wedge during the SLPI start window
before post-start logs could be collected. Host later saw the phone in fastboot. So the patch
likely moved execution past the old
dsp informationgate, but we do not have a clean next failure line yet.
Next real work: do not keep hand-starting SLPI over SSH and hoping logs flush. Add a capture path
that survives the wedge (pstore/ramoops, serial/console, persistent journal sync, or earlier in-kernel
tracepoints), then repeat the one-shot sdsprpcd sensorspd test. In parallel, source-diff mainline
FastRPC SDSP/sensors-PD against Lineage drivers/char/adsprpc.c, especially INIT_ATTACH_SNS,
sensors static PD, listener1/geteventfd, service-location/PDR, and the domain/session handling around
sdsprpcd sensorspd.
2026-06-29 (cont.) — FastRPC sensors-PD/servreg gap isolated (move #2 design)
Source-diff of mainline drivers/misc/fastrpc.c vs Lineage drivers/char/adsprpc.c around the
sensors PD (INIT_ATTACH_SNS). This is the concrete next-conversion target (move #2); it is gated
behind making the SLPI-start wedge survivably observable (move #1).
Mainline sensors attach is a bare RPC — references/linux-7.1.1-sdm845-CANONICAL/drivers/misc/fastrpc.c:1735:
fastrpc_init_attach(fl, SENSORS_PD) sends FASTRPC_RMID_INIT_ATTACH (method 0, tgid arg), sets
fl->pd = 2, returns. grep confirms mainline fastrpc.c has NO pdr/servreg/service_location/
spdname wiring — only the generic INIT_CREATE_STATIC (method 8) ioctl (:1340/:2193).
Lineage wraps the sensors PD in a servreg/PDR scaffold mainline lacks:
INIT_ATTACH_SENSORStags the client (adsprpc.c:2250):fl->spdname = SENSORS_PDR_SERVICE_LOCATION_CLIENT_NAME;(="sensors_pdr_adsprpc"),fl->pd = 2.- At probe, gated on DT prop
qcom,fastrpc-adsp-sensors-pdr(adsprpc.c:4371), it registers a service-location notifier:get_service_location("sensors_pdr_adsprpc", "tms/servreg", &get_service_nb);
So the production SLPI firmware expects the AP to track the sensors PD via tms/servreg
(PDR/service-location) around INIT_ATTACH_SNS. Mainline supplies servreg via pd-mapper
(userspace DB) + the glink reverse channel — so the open question is whether the sensors-PD servreg
domain is present in pd-mapper's DB AND whether the reverse listener is up when INIT_ATTACH_SNS
fires (ties to the S6 auto-boot ordering race).
Conversion target (move #2): register a PDR service-location lookup for the SDSP channel in
fastrpc_rpmsg_probe (fastrpc.c:2432) when domain_id == SDSP_DOMAIN_ID, using mainline's existing
pdr_handle_alloc() (drivers/soc/qcom/pdr_interface.c, already used by apr.c/pmic_glink.c) —
mirroring Lineage's get_service_location. Service path TBD from the SLPI firmware expectation /
pd-mapper config (likely an "msm/slpi/sensors"-style servreg domain).
Capture (move #1) — state:
- Running kernel already has
CONFIG_PSTORE_RAM=y/PSTORE_CONSOLE/PSTORE_PMSG/PSTORE_COMPRESS(verified on phone via/proc/config.gz). pstore backend is currently(null)because there is noramoopsreserved-memory node in our DTS (only apq8016/msm8992 boards carry one). - So ramoops capture needs only a reserved-memory node — NO config change. BUT the physical address
must be bootloader-preserved across the watchdog→reset (the wedge lands the phone in fastboot).
Do NOT guess an address (brick risk, no EDL net). Candidate preserved region to verify against the
live map in
/proc/device-tree/reserved-memory/: Lineageremoved_region@85fc0000, or a top-of-RAM carve. Verify ABL preserves it across reset before flashing. - Safe, no-flash capture helper staged now:
tools/slpi-capture.sh— snapshots dmesg to a timestamped file on the UFS rootfs and tails/dev/kmsgcontinuously through the SLPI start, so a partial capture survives the wedge on reboot even before ramoops is wired. Inert until invoked; does NOT start SLPI.
Coordination note (2026-06-29): another model is baking the RTC time-fix (pm8998 rtc@6000
label + blueline &rtc { allow-set-time; }) — uncommitted in the canonical worktree. Do not edit
pm8998.dtsi/sdm845-google-blueline.dts until that lands; the fastrpc conversion (a different
file) and the ramoops node (sdm845.dtsi shared reserved-memory) can proceed collision-free.
CORRECTION 2026-06-29 (post-Opus, supersedes the "pd-mapper missing" premise above):
pd-mapper is IN-KERNEL here, not userspace — CONFIG_QCOM_PD_MAPPER=m (module
qcom_pd_mapper, aux devices qcom_common.pd-mapper.{0,1,3}), and it is serving
(wlan0 up proves the wlan PD is advertised). The earlier bug was that the in-tree
arch/arm64/configs/sdm845.config carried upstream CONFIG_QCOM_PD_MAPPER=n and,
being merged last by merge_config.sh -m, overrode pmaports' =m on every build
→ wifi/sound/SLPI-sensor all lost their PD broker. Opus flipped that line to =m,
rebuilt, reflashed → wifi + sound came back. See build.md gotcha #0.
Crucially, the in-kernel sdm845_domains[] table includes slpi_root_pd AND
slpi_sensor_pd — so the sensors PD is being advertised now. Therefore:
- The "servreg server missing → install userspace pd-mapper" premise in this section
was WRONG (I installed then reverted the userspace
pd-mapperapk; userspace pd-mapper is explicitly NOT used on blueline). Do not install it. - The in-kernel fastrpc
pdr_add_lookupconversion (move #2 above) is NOT the gap — the sensors-PD servreg server is present in-kernel. Drop that conversion. - Every prior SLPI conclusion (auto-boot ordering, fastrpc-attach EPERM, sar.cc:27)
was drawn on a substrate with no sensors PD advertised. Re-test SLPI on the
post-pd-mapper (
=m) kernel before any more fastrpc/SLPI source work.
2026-06-29 live result — Lineage adsprpc memory conversion moved the wall
This was real progress, not log suppression. We ported the blueline-specific Lineage/downstream adsprpc memory behavior into the 7.1.1 mainline tree:
adsp_mem/fastrpc_memare split like Lineage: fixed no-mappil_adsp_mem@8c500000foradsp_pas, dynamic reusableadsp_regionfor adsprpc.- The SDSP FastRPC node uses the Lineage VMID set:
HLOS + MSS_MSA + SSC_Q6 + ADSP_Q6. - A new mainline
qcom,msm-adsprpc-mem-regionhelper preassignsadsp_regionbefore SDSP FastRPC probes, matching downstreamqcom,msm-adsprpc-meminstead of letting mainline assign the heap late.
Verified on hardware after deploy to slot A:
qcom_scm Assign memory protection call failed -22count: 0.- Boot log says the adsprpc heap was preassigned and SDSP FastRPC skipped duplicate hyp-assign.
- With corrected manual ordering (listener watcher running before SLPI start),
SNS_REG_INITcount after start: 0. - The failure moved later to:
EF:sensor_process:1:sar.cc:27:chre_utils fatal.
So the old wall is fixed: the TZ/heap -22 and SNS_REG_INIT are no longer the
current explanation. The current SLPI wall is downstream of attach/init:
hexagonrpcd -f /dev/fastrpc-sdsp -d sdsp -s still reports
Could not attach to FastRPC node: Operation not permitted, and the firmware later
dies in SAR/CHRE. The next Lineage conversion target is therefore the sensor/SAR
FastRPC attach path, not another attempt to mask the crash-loop.
Symptom: SLPI remoteproc loads and reaches running, then the firmware
faults in CHRE / SAR sensor init (sar.cc:27, chre sar_utils) and
crash-loops every ~10s. hexagonrpcd ends up STOPPED. SLPI runs fine on
the LineageOS/Android side of the same phone, so the gap is on ours.
Reference sources (corrected 2026-06-22 — was wrong before)
Earlier "Android reference" diffs were against references/msm-google-crosshatch
= Google AOSP kernel/msm 4.9.270 (android.googlesource.com/kernel/msm).
That is STOCK Google, not what the phone runs. The phone runs LineageOS 22.2
(slot B). The authoritative references are now on disk and provenance-verified:
- LOS kernel SOURCE:
references/los-kernel-blueline-4.9/=LineageOS/android_kernel_google_msm-4.9@lineage-22.2, 4.9.337, commit b929452c5. The blueline device tree (LineageOS/android_device_google_crosshatch,lineage_blueline.mk, PRODUCT_MODEL "Pixel 3") pulls this kernel vialineage.dependencies→kernel/google/msm-4.9. Provenance match: boot.img string is4.9.337-gb929452c5c49(built 0615) — same sublevel, same commit hash. This is the literal source of the kernel running on slot B. - LOS DTB:
/home/casey/Downloads/los-blueline.dts(decompiled). - LOS vendor partition:
pixel3-work/vndlos-mnt(confirmed LOS viavendor.lineage.health-service,hal_lineage_livedisplay_qti— not stock).
Note: LineageOS does NOT ship its own qcom remoteproc/PIL/fastrpc drivers —
it inherits google/msm-4.9. So crosshatch 4.9 is the correct driver-source
base, just an older point release; the LOS-specific deltas are DTB + vendor HALs.
Mainline vs LOS SLPI lifecycle — what is and isn't a real gap
LOS uses the downstream qcom,pil-tz-generic driver
(subsys-pil-tz.c + subsystem_restart.c + peripheral-loader.c).
Mainline uses qcom,sdm845-slpi-pas (qcom_q6v5_pas.c + qcom_q6v5.c).
The DT property lists differ, but most differences are spelling, not behavior:
NOT gaps (verified):
- Identity — PAS id 12, crash-reason SMEM 424, ssctl 0x16, firmware "slpi"
all match mainline match-data (
qcom_q6v5_pas.c:1300-1312). - Lifecycle handshake — LOS
qcom,gpio-err-fatal/err-ready/stop-ack/force-stopresolve toqcom,smp2pgpio-ssr-smp2p-3-in, entry "slave-kernel", remote-pid 3 (los-blueline.dts:10453) — i.e. SMP2P, not raw GPIO. Mainline wires the same SMP2P asslpi_smp2p_infatal/ready/handover/stop-ack +slpi_smp2p_outstop (blueline-pmos-slpi.dts:3460-3469). Same hardware. Behavioral proof: the subsystem reachesrunning(thereadySMP2P fired) — a handshake gap could not produce a post-boot crash. - signal-aop / AOP load_state — LOS
pil_notify_aop()sends QMP{class: image, res: load_state, name: slpi, val: on}(peripheral-loader.c:949-955,980). Mainlineq6v5_load_state_toggle()sends the BYTE-IDENTICAL string viaqmp_send(qcom_q6v5.c:31-32), gated onqmp_getreturning non-NULL (:28,335). Our DTS wires it: SLPI node hasqcom,qmp = <0x2f>(blueline-pmos-slpi.dts:3464) →qcom,sdm845-aoss-qmp(:4912), andCONFIG_QCOM_AOSS_QMP=y. So this signal genuinely fires. - The LOS GPIO/proxy/mbox/pil-tz DT props are bindings of
subsys-pil-tz.c/subsystem_restart.c, which mainline does not have. Pasting them onto our PAS node is inert — the PAS driver never reads them. (This was the earlier "reconcile the DT" dead end.)
THE REAL GAP (runtime power, SLPI-specific):
qcom,keep-proxy-regs-on. Mainlineqcom_pas_handover()drops ALL proxy power the instant the remote signals handover — px/cx regulators, xo/aggre2 clocks, and the lcx/lmx power domains (qcom_q6v5_pas.c:386-397, unconditional). LOS deliberately keeps SLPI's vdd_cx/vdd_mx rails enabled for the subsystem's whole life:disable_regulators(..., reg_no_disable=true)skipsregulator_disable()whenkeep_proxy_regs_onis set (subsys-pil-tz.c:501-502, viapil_remove_proxy_vote:583). Strongest signal it's load-bearing:qcom,keep-proxy-regs-onis set on the SLPI node and NOT on adsp (los-blueline.dtsssc node ~3108 vs adsp ~3053). Hypothesis: SLPI's CHRE/SAR runtime needs the AP to hold CX/MX after boot; mainline yanking them post-handover collapses the rails →sar.cc:27fault. Matches the boot-then-crash signature exactly, and matches every prior win on this project (Android holds a resource mainline drops; port it).
(Secondary, same direction: qcom,proxy-timeout-ms = 10000 — LOS holds proxy
power 10s post-boot then unvotes; mainline releases at handover with no grace.)
FastRPC: LOS qcom,msm-fastrpc-compute (adsprpc.c, 7 CBs + per-CB IOMMU) vs
mainline qcom,fastrpc (vmids + shared-dma-pool) — different drivers, but our
mainline node already attaches (/dev/fastrpc-sdsp exists). Not the blocker.
The patch (2026-06-22)
pmaports .../linux-postmarketos-qcom-sdm845/slpi-keep-proxy-power-on.patch
(wired into APKBUILD source=, pkgrel 3→4, checksums updated).
Adds a per-SoC opt-in keep_proxy_pds_on flag to qcom_pas_data/qcom_pas,
set true only for sdm845_slpi_resource_init, and makes qcom_pas_handover()
return early (skipping the proxy teardown) when set. adsp/cdsp/modem unchanged.
This mirrors downstream keep-proxy-regs-on: SLPI's cx/mx (lcx/lmx) stay
powered past handover. Patch verified to apply against the build tree
(tag 85f1df2a, identical to references/linux-7.1.1-sdm845-CANONICAL); dry-run rc=0.
Known trade-off (acceptable on a throwaway test bench): power is held across stop/SSR-restart too, so refcounts are not balanced on a clean shutdown — same "never turns the rails off" behavior downstream has. One online/boot attempt per cycle anyway.
Status / next
- [done] References corrected + provenance-verified.
- [done] Behavioral diff (LOS source vs mainline) — gap isolated to proxy power.
- [done] Patch written and proven to apply.
- [done] Folded into the unified 7.1.1 tree as a commit (no longer a
source=.patch):references/linux-7.1.1-sdm845-CANONICAL, alongside the SLPI-enable DTS commit and the fastrpc SDSP tolerance commit. Built via pmbootstrap (see build.md). - [done] Deploy unified 7.1.1 kernel to slot A, reboot (2026-06-26, kernel r2 LLVM).
- [DISPROVEN 2026-06-26] LIVE TEST: keep-proxy-power-on does NOT stop the
crash-loop. On the booted r2 kernel SLPI still faults
err_qdi.c:456:EF:sensor_process:1:sar.cc:27:chre_utils fataland crash-loops every ~10s (observed crash #66→#72 in one capture). So the proxy-rails-held hypothesis is wrong — holding cx/mx past handover did not save it. Theqcom_ssc_sensorsmodule (CONFIG_QCOM_SSC_SENSORS=m) is present and loaded this build, and SLPI still crashes — so the SSC sensor transport being absent was not the (whole) cause either. The fault is inside the SLPI firmware's SAR (CHRE sar.cc) path itself: mainline isn't giving the firmware the SAR-sensor hardware/config it expects, and neither rails nor the SSC glink transport fix it. - [RESOLVED 2026-06-29] The
qcom_scm ... Assign memory protection call failed -22/SDSP fastrpc hyp-assign rejected by TZ (-22)wall is fixed by the Lineage adsprpc memory conversion above. Do not chase this as the current failure. - [current] Next: convert the remaining Lineage sensor/SAR FastRPC path:
sscrpcd sensorspd,libsdsprpc.so, sensors PDR/service-location behavior, and whatever makesINIT_ATTACH_SNSattach cleanly instead of returningEPERM.
2026-06-27 (S6, Opus) — VERBOSE TRACE settles the registry theory: it's an auto-boot ORDERING race, not file serving
Built a verbose hexagonrpcd (the -Dhexagonrpcd_verbose=true meson option only
exists in the newer source pmos/_src/hexagonrpc @ dd9ac70, NOT in the v0.4.0 tarball;
our two patches — registry-nodes + method-29 fstat — are already applied in that working
tree). Snapshotted it to temp/hexagonrpcd/hexagonrpc-0.4.0.tar.gz and built with verbose
on. Deployed, captured two clean traces of what the SLPI actually requests.
THE FINDING (instrument-proven, not inferred):
- Across both captures the SLPI made zero registry file requests (
openat/opendirof real paths = 0) before crashing atSNS_REG_INIT:36. hexagonrpcd's reverse listener was attached and blocked in the listen loop, and received nothing. - The ~981 log lines were all hexagonrpcd's OWN failures to hold the device node:
972
No such file or directory(/dev/fastrpc-sdspvanishes while SLPI is down mid crash-cycle), 139Broken pipe, 94Operation not permitted, 9Resource busy. - So
SNS_REG_INIT:36fails BEFORE the SLPI requests any file. Serving the registry (staging path, dsp/sdsp libs, the gate, RestartSec=0) was answering a request that, at this crash point, never happens. The registry-serving theory is DEAD for SNS_REG_INIT.
Root cause (circular, self-reinforcing): the SLPI auto-boots at qcom_q6v5_pas module
probe (~3s) → its registry-init needs a reverse-RPC to the AP → no listener registered yet
(hexagonrpcd is userspace, ~33s) → fails instantly at SNS_REG_INIT → crashes → the crash
tears down /dev/fastrpc-sdsp → reappears ~0.3s later → vanishes again. hexagonrpcd can
never hold a stable attach long enough to register its listener; the node thrashes faster
than it can attach + enter the listen loop. The rare earlier glimpses of sar.cc:27 were a
hammering hexagonrpcd momentarily aligning with one restart's request window — not a stable
state. (Note: crash line drifts :36/:39 boot-to-boot; same failure, pre-file-request.)
Real fix direction (Casey's architectural read, now proven): the SLPI must NOT auto-boot
at module probe. Hold it stopped until hexagonrpcd is attached AND in its listen loop, then
start it once so its first registry request hits a ready listener. That's a kernel/remoteproc
ordering change (auto_boot=false for the SLPI PAS, or defer module bind), NOT more userspace
plumbing. Mirrors Android: SLPI firmware boots but sensor_process is triggered by the AP
(chrecd chre_slpi_start_thread) AFTER the listener/daemon is up — confirmed by subagent that
INIT_ATTACH_SNS does NOT itself start sensor traffic; chrecd does.
Still-valid hygiene from S5/S6 (real, but does NOT fix SNS_REG_INIT): the registry staging
path bug (cp -a .../registry/. was nesting registry/registry/; fixed to flatten) and the
blueline-slpi-ready.service gate pattern. libchre_slpi_skel.so IS needed for chrecd and was
found at pixel3-work/vndlos-mnt/dsp/sdsp/ (whole dsp/sdsp/ tree: skel + nanoapps +
fastrpc_shell_2); it's static vendor firmware, belongs in a firmware package, not a one-time copy.
2026-07-01 — crash-loop-vs-modem interaction found + fixed; SLPI ruled OUT of the modem 52 wall
Wired blueline-slpi-ready.service into the real boot chain (modem-qrtr-ready.service
After=/Wants= it, mirroring Android's init.qcom.devstart.sh which sets
vendor.qcom.devup — the property modem-qrtr-ready was modeled after — immediately after
kicking SLPI). This surfaced a real regression: blueline-slpi-ready.sh disabled SLPI crash
recovery for a "single-shot" start. When the SAR/CHRE fault crashed it with recovery disabled,
the driver had nowhere to go — glink-edge was left half torn down, storming
intent request timed out every ~10s indefinitely and taking the modem's own QMI/QRTR
transport down with it (endpoint hangup). Android never disables recovery for this
(ssr_setup/ssr_diag in init.sdm845.rc exist precisely to handle ongoing SSR cleanly).
Fix: recovery stays enabled through the boot attempt (so any crash tears down/rebuilds
glink cleanly via the kernel's own SSR path — verified: crash→recovering→running cycling with
zero storm across 50 crashes in one test), then the script catches SLPI in its brief running
window and issues a clean echo stop to park it — one attempt, not an infinite crash-loop,
since the modem does not benefit from SLPI staying up (see below).
Decisive result, three separate live boots: with SLPI crash-looping 50 times in ~7 minutes,
with SLPI crash-free and running continuously, and with SLPI cleanly parked/offline —
mmcli -m 0 --enable returned byte-for-byte identical QMI protocol error (52): 'DeviceNotReady'
every time. SLPI state has zero measurable effect on the modem RF-arm wall. This closes
"Subsystem coordination (SLPI PD state?)" from los-mss-translated.c's delta list as a modem
blocker. SLPI's own SAR/CHRE crash (this file, above) remains real and unfixed, but is now
fully decoupled from modem bring-up — see modem.md TOP LEAD #12 for where the modem hunt
went next (dms-ui-ready.py, DMS 0x0064).