/model with no argument opens a searchable list of everything your endpoints actually serve.
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ โ search โบ gemma free 2 of 464 โ โ โฏ openrouter/google/gemma-4-26b-a4b-it:free โ โ openrouter/google/gemma-4-31b-it:free โ โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
The catalogue runs into the hundreds and the ones you actually use are a handful. /favourite pins the current model so it rises to the top of /model; running it again unpins it. While the list is open, ctrl+f pins or unpins the highlighted model without leaving, so you can mark several in one pass.
Pinning is independent of default: a favourite is a shortcut to reach a model, not a decision to use it next time.
Off by default. Turn it on with a ladder of models, largest last. When the conversation outgrows the current model, raunen moves to the next rung rather than failing โ and nothing already found is lost in the handover.
{
"auto_switch": true,
"fallback": [
"ollama/qwen3-coder:30b",
"ollama-cloud/qwen3-coder:480b"
]
}The ladder is yours to define and can mix local and hosted models freely. Adding "free_fallback": true appends every model the providers report as free, roomiest first โ see what it costs.
A ladder is only useful if it remembers what just failed. Otherwise a rate-limited model is retried every turn, fails every turn, and each one pays the same tax. The full table of failures and responses is on the cost page; the principles are:
/status shows what is being held back, which is why a ladder can look shorter than it is:
held back 2 models groq/llama-3.3-70b ยท cooling down for 1m30s nvidia/typo ยท locked out: rejected