Docs menu · Reference

Removes cached model engines from the runtime to free memory. Pass a model spec to unload one entry, or omit it to unload every cached generation, embedding, and ranking model. The result reports whether any cached entry was actually removed. This affects only the in-memory runtime cache; it never deletes downloaded model files from disk.

Examples

Evict cached models from the runtime (a no-op when nothing is loaded).

CLI

$ strata inference unload

Wire

{"type":"inference_unload"}

Parameters

NameTypeRequiredDescription
modelstringnoOptional model spec.

Returns

UnloadResult.

Errors

Invocation

  • CLI: strata inference unload
  • Wire type: inference_unload

agents: this page as markdown → /docs/reference/inference/unload.md