research log

PSI signal validation and rate-guided zram experiments

Published 2026-07-14 · Updated 2026-07-21

The long-term question behind this work is practical: can memory-management policy keep a constrained machine responsive? It is tempting to begin by implementing a controller and comparing performance. That would put the policy ahead of the evidence needed to justify it.

Origin question

Can Linux Pressure Stall Information identify early, repeatable pressure windows before reclaim and swap activity dominate?

This began as a narrower question than whether a particular policy improves performance. The first task was to establish whether the proposed input to such a policy was timely and stable enough to be useful at all.

Experimental context

The harness runs fixed workloads under controlled memory limits and records rolling PSI windows alongside memory availability, reclaim counters, page faults, swap, zswap, zram, cgroup observations, and phase-aligned responsiveness timing. Runs use stable raw JSONL output, and failed runs are retained rather than silently removed.

Workload compressibility is an explicit control. A policy can appear effective because the workload compresses well, not because its pressure response is generally useful. Disk-backed swap and zswap must also be treated as different experimental conditions.

Evidence to date

Controlled V1 and V2 reruns established that the measured PSI counter becomes observable before material swap-slot growth under the tested workloads. The available lead time changes with the pressure ramp and memory configuration; it is not treated as a universal constant.

The original aggregate latency probe was not a valid responsiveness outcome. A replacement canary separates wake delay, service time, and end-to-end response on a monotonic schedule. Under pressure, it records worsening response and service-time tails rather than the misleading improvement reported by the original aggregate probe.

Compression-sensitivity work also confirmed that page content matters: structured and sparse pages occupy zswap very differently from random pages. Those results characterize the compression path; they do not themselves show a policy advantage.

Decision

The first milestone was signal characterization. The evidence needed three properties:

An isolated early PSI reading is not a policy result. The validation work established enough observable signal to justify a bounded policy screen, not a claim that PSI alone prescribes an optimal action.

Guided-zram policy screens

The policy study used separate foreground and background cgroups. A 384 MiB memory boundary made background pressure attributable to the cgroup, while the foreground canary remained outside that cap. The controller observed sustained background-cgroup PSI-some over one-second windows and activated zram once when a threshold-and-hold predicate qualified. PSI-full remained diagnostic telemetry.

The first three-host gradual-onset screen covered disk swap, static zram, and early/late sustained-PSI-guided zram across structured and random workloads at moderate and gradual pressure rates. It retained phase-aligned response and service-time tails, deadline misses, swap growth, backend occupancy, faults, reclaim, and valid controller logs.

That screen motivated a narrower question: could a signal available before sustained PSI distinguish the two tested pressure regimes and choose a better activation time? PSI onset slope could not do that under the current cgroup boundary, but the one-second slope of background/memory.current could. The V5 controller classified the tested moderate ramp at about 0.57 seconds, compared with about 5.75 seconds for PSI-only qualification, and produced no false fast classification under the gradual ramp. Gradual cases fell back to PSI at essentially the same time as the PSI-only controller.

This is a mechanism result, not yet a policy win. It establishes that cgroup-memory growth rate can recognize the tested moderate regime roughly five seconds before PSI-only qualification, while preserving the gradual fallback path.

Focused telemetry confirmation

The first rate-gated comparison had incomplete zram observability, so a focused three-host confirmation repaired the agent telemetry before interpreting outcomes. Every policy sampled zram occupancy, compression, backing-device activity, and writeback throughout the run.

The storage path differed sharply by workload. Structured pages compressed at roughly 8.4× and occupied about 32–38 MiB of zram without backing writeback. Random pages compressed at roughly 1.8–2.1×, occupied about 156–167 MiB, and produced about 40.5k–56.2k backing-write pages. Compressibility is therefore not a cosmetic workload attribute: it changes the path that an activation policy is controlling.

Earlier rate-gated activation did not consistently outperform static zram. In moderate structured pressure, static zram had the best foreground response p99 on all three hosts despite similar zram occupancy to the rate-gated policy. In gradual random pressure, PSI-only was slightly best on average, but the host-level result was mixed; rate-gated activation lost to both alternatives on all three hosts. The focused data does not support a general claim that activating zram earlier improves responsiveness.

What worked and what did not

Several parts of the research programme are now established under the tested workloads:

What did not work is a simple timing claim: earlier one-shot activation is not reliably better than static zram. The binary controller question, “activate now or later,” does not explain the outcome by itself.

Ongoing research

The broader research question remains how memory-management policy can preserve responsiveness under constrained RAM. The next experiment is deliberately not fixed. Capacity management, writeback behavior, and workload characteristics are plausible future directions, but the current evidence does not yet select one controller design.

Limitations

Memory PSI reports aggregate time tasks spend stalled because of memory pressure; it does not identify the cause of every latency event or prescribe an action. The VPS canary's end-to-end tail is also dominated by wake delay, so observed policy differences cannot be attributed solely to compression or page-fault cost. The experiments therefore treat PSI as one signal alongside reclaim, faults, swap activity, memory availability, backend telemetry, and workload observations.