Local Inference Server · CUDA · Free
One executable. Point it at a model file and run. No installer, no background service, no setup beyond that.
Scroll
The younger brother
Parta is a full platform: identity, audit trails, policy control, production-grade infrastructure for regulated industries. Runtime Lite is not that. We stripped all of it out and kept only the inference server.
No identity layer / No audit chain / No policy engine
Compatible by design
OpenAI-compatible and Anthropic-compatible chat completion endpoints, so anything already built against those APIs works against your own local server. Vision support is included for models that accept image inputs.
OpenAI-compatible / Anthropic-compatible / Vision included
Copy it anywhere
One statically linked executable and one configuration file. No admin rights, no installer, no background service to manage. Point the config at a GGUF model file already on your machine.
Bring your own model / No account / No tracking
Roadmap
This release targets Windows with an NVIDIA GPU. Vulkan and Linux builds are on the roadmap after CUDA, for anyone not on NVIDIA hardware.
CUDA / Vulkan (soon) / Linux (soon)
Windows · CUDA · v0.1.0
Download the zip and check its SHA256 checksum against the value below before running it.