llama.cpp Mlock Failed on Raspberry Pi: Limits and Load Modes¶
If llama.cpp reports a failed memory lock on Raspberry Pi, check both the loading option supported by your binary and the locked-memory limit of the running process. Increasing a limit does not create RAM. A model that exhausts physical memory needs a smaller model, context, or workload before memory locking is useful.
This page covers lock failures and option migration. For building llama.cpp, choosing a GGUF model, and measuring inference speed, start with the Raspberry Pi llama.cpp setup guide.
llama.cpp mlock error quick reference¶
| Observation | Check first | Next step |
|---|---|---|
--mlock is not recognised |
Binary version and --help |
Use the loading option shown by your installed version. |
| Lock warning mentions memory allocation | Process limits and available memory | Check the soft memlock limit before assuming the model is too large. |
| Process is killed during startup | Kernel and service logs | Diagnose OOM separately from a recoverable lock warning. |
| Terminal launch works, service fails | Service process limits | Inspect the running PID, not your interactive shell's ulimit. |
| Locking succeeds but speed is unchanged | Comparable benchmark runs | Keep locking only if it helps the workload you measured. |
Check the current loading options¶
From your llama.cpp checkout, inspect the actual executable:
The upstream option parser checked on 2 October 2026 exposes --load-mode, including mmap, mlock, and mmap+mlock. These select different loading paths. See the current argument implementation, model loader, and upstream loading-mode documentation.
| Loading choice | Purpose |
|---|---|
--load-mode auto |
Let the runtime choose its normal loading path. |
--load-mode mmap |
Memory-map the model without requesting locking. |
--load-mode mmap+mlock |
Memory-map and request that model memory remain resident. |
--load-mode mlock |
Use the non-mmap loading path with memory locking. |
--load-mode none |
Use the non-mmap loading path without locking. |
Older builds used separate --mlock and --no-mmap switches. For a current build, mmap+mlock expresses the mapped-and-locked intention; mlock expresses the unmapped-and-locked intention. Confirm support in local help rather than mixing flags from different revisions. Do not assume that a wrapper such as llama-cpp-python has the same command-line interface.
Establish an unlocked baseline first¶
Choose a model that fits with the OS, context/KV cache, compute buffers, and other services. Start with a short context and a single workload:
Replace the model path with your own. Stop that server before starting the comparison run with --load-mode mmap+mlock. Keep the model, context, requests, and other options identical. The server remains local to the Pi in this example; you can check readiness using curl http://127.0.0.1:8080/health from another terminal on the Pi.
Memory locking can help when pages would otherwise be evicted or swapped under pressure. It does not increase CPU compute throughput. Compare startup time, request latency, and memory use rather than assuming a higher tokens/s result.
Inspect the running process's memory limit¶
For a shell-launched process, inspect the shell's soft and hard limits before launching:
In Bash, these values are in KiB. An unprivileged shell cannot raise its soft limit above its hard limit. Changing the limit in another terminal does not change an already-running server.
For an existing process, use its real PID. For a system service named llama-server.service:
/proc/PID/limits reports the process limits in bytes for locked memory; VmLck reports locked memory in kB. A successful model load is not proof that every requested lock succeeded. Read llama.cpp's startup warnings as well as the process state. The Linux mlock manual explains locked-page behaviour, limit checks, and error conditions.
Set a bounded systemd limit when required¶
A system service gets its limits from the service manager. Your SSH shell's ulimit is not its configuration. Inspect the current unit and effective settings:
If the measured workload fits RAM and needs more locked memory, create a service override with sudo systemctl edit llama-server.service. For example:
2G is an illustrative limit, not a model-size recommendation. The environment setting is supported by the checked upstream parser; if ExecStart already sets a loading mode explicitly, update that deliberate choice rather than maintaining conflicting settings. Keep the service's unprivileged user.
The systemd execution reference documents LimitMEMLOCK and the distinction between system and user-service limits. A user service cannot exceed the limits inherited by its user manager merely by adding a larger unit value.
Apply the override during a suitable restart window, then inspect the new PID:
Setting a limit permits locking up to that amount; it does not preallocate that amount or guarantee a successful lock. Avoid granting every process an unlimited lock allowance to fix one service.
Distinguish a lock warning from an OOM kill¶
Check memory and kernel logs:
A service can hit its cgroup memory limit even when the whole Pi has free memory. Increasing LimitMEMLOCK does not increase MemoryMax or the board's physical RAM. If the model only works unlocked and starts thrashing under load, reduce the model size, context, or concurrent workload before enabling locks. See the memory-pressure and OOM guide.
Record whether the change helped¶
Compare the same requests with unlocked and locked runs. Record the llama.cpp revision, loading mode, model checksum, context, concurrency, startup time, request latency, VmLck, peak memory, and thermal/power state. Separate a cold first load from a cached repeat. Retain the setting only when it improves the deployment's measured behaviour without starving other services.