Wally is a command-line tool that connects your coding agent to the model of your choice. With a single command, it opens a coding agent, such as OpenCode, Claude Code, Hermes, OpenClaw, DeepSeek Harness or Prime Agent, already configured for that model, and it installs the agent first if it is missing. You can use a hosted model, such as GLM 5.3 Flash, billed against your RunAnywhere credit, or download an open model and run it on your own machine without an account. Wally can also chat with a model in the terminal and serve it through an OpenAI-compatible API. It runs on macOS, Linux and Windows.
Get started with the official documentation.
On macOS with Apple Silicon (Intel Macs are not supported) and on Linux with x86-64 or ARM64 and glibc 2.39 or newer (Ubuntu 24.04+, Debian 13+):
curl -fsSL https://raw.githubusercontent.com/RunanywhereAI/wally/main/install.sh | shOn Windows (x64 and ARM64):
irm https://raw.githubusercontent.com/RunanywhereAI/wally/main/install.ps1 | iexThe installer adds wally to your PATH for new terminals. In the terminal you
installed from, run the line it prints (. ~/.zshrc or similar), or open a new
one, before your first wally command. On Windows, open a new terminal.
Get up and running with a few commands.
Log in:
wally account loginOpen a coding agent on a hosted model:
wally opencode -m glm-5.3-flashOr open it on the default model:
wally opencodeThe default model is glm-5.3-flash. To change it, run wally models default <model>.
Get help:
wally helpLearn more about open models in our official documentation.
Wally runs the models; Eve makes the calls. eve is RunAnywhere's hosted
decision model: it scores explicit labels without generating text. Ask yes/no,
choice, or ordered score questions:
wally decisions --input "Checkout is blank after Pay" \
--ask "Is this a software bug?" \
--choice "Owner=frontend,payments,account"Use --json for the raw API response, or --request FILE for the complete
typed request shape. --image FILE (PNG, JPEG or WebP, up to 8, cloud only;
on macOS a HEIC photo is sent as a full-size JPEG, upright and without its
EXIF or GPS) asks every question about the images too; each image is billed as prompt
tokens once per question, and a model that does not take images refuses the
request before anything is billed. Images are sent at full size: the service
shrinks each to what the model sees, so a larger photo scores and costs the
same. It takes a request of up to 24 MiB with its images (1 MiB outside them),
and an image of up to 64 MP as JPEG or 24 MP as PNG or WebP; past those it
refuses with HTTP 400 and says which limit. Decision models are refused by wally run and the coding
tools, which point back here.
Eve is also runnable on this machine. Pull a local checkpoint and point
-m at it; --local (or just a local -m) scores in-process through the
SDK's decision component, --cloud forces the hosted path:
wally models pull clef-flash-9b
wally decisions --local -m clef-flash-9b \
--input "Checkout is blank after Pay" --ask "Is this a software bug?"Both transports share the request flags, the human rendering and the --json
document, so the same invocation works locally and hosted.
Wally is written in Rust and built with CMake against a prebuilt RunAnywhere C++ desktop kit (not the SDK source tree). CONTRIBUTING.md has the steps.
- Releases: every version, with checksums
- Engines and platforms: what runs where and how wally picks
- Models: the full catalog
- Editors and hosted models: how each coding agent is wired and where your session lives
- docs.runanywhere.ai
- Discord
- Hugging Face
Licensed under the MIT license. See LICENSE.
