Parta Runtime Lite

Local Inference Server · CUDA · Free

One executable. Point it at a model file and run. No installer, no background service, no setup beyond that.

Scroll

The younger brother

Parta Runtime,
without the platform.

Parta is a full platform: identity, audit trails, policy control, production-grade infrastructure for regulated industries. Runtime Lite is not that. We stripped all of it out and kept only the inference server.

No identity layer  /  No audit chain  /  No policy engine

Compatible by design

Speaks the formats
you already use.

OpenAI-compatible and Anthropic-compatible chat completion endpoints, so anything already built against those APIs works against your own local server. Vision support is included for models that accept image inputs.

OpenAI-compatible  /  Anthropic-compatible  /  Vision included

Copy it anywhere

Download, unzip,
run.

One statically linked executable and one configuration file. No admin rights, no installer, no background service to manage. Point the config at a GGUF model file already on your machine.

Bring your own model  /  No account  /  No tracking

Roadmap

CUDA is first.
Vulkan and Linux next.

This release targets Windows with an NVIDIA GPU. Vulkan and Linux builds are on the roadmap after CUDA, for anyone not on NVIDIA hardware.

CUDA  /  Vulkan (soon)  /  Linux (soon)

Download and verify.

Windows · CUDA · v0.1.0

Download the zip and check its SHA256 checksum against the value below before running it.