Open-source experimental pre-release

Run and manage AI models on your own infrastructure.

Orchard provides OpenAI-compatible APIs, manages model workers on Apple Silicon, and records how each request was handled.

OpenAI-compatible Chat Completions and Responses APIs · Apache-2.0
Request evidence live
REQ_01K7C3A9

POST /v1/responses

Completed
Policy admittedTenant grant verified · deadline captured
+ 8 ms
Placed on mac-studio-03Healthy · model loaded · 3 slots available
+ 21 ms
MLX worker streamed 186 eventsTool call normalized · one terminal outcome
+ 4.2 s
Capacity releasedAttempt complete · no retry required
+ 4.3 s
model: qwen3-coder · exact revision pinned200 completed
Open sourceApache-2.0 licensed
API supportChat Completions and Responses
Current runtimeApple Silicon and MLX
Project statusSource-only experimental pre-release
Request history

See how each request was handled.

Orchard connects access checks, Node selection, model execution, and the final result in one request record.

01 · Govern

Orchard checks access before the request runs.

Orchard authenticates the client, checks its model access, and records the request deadline before sending work to a model.

REQ_01K7C3A9 / ATTEMPT_01recorded
Tenantstudio-api
Model grantAllowed
Deadline28.0s remaining
Orchard Console

Manage models, Nodes, and access.

The Console gives operators one place to manage the service used by their team.

01

Models

Find model artifacts, manage downloads, register them in the catalog, and activate them when they are ready.

See the Model Catalog
02

Nodes

Enroll a machine, admit it to the cluster, and check its current health and model-serving state.

See the runtime path
03

Workspace Access

Grant access to specific models, manage API credentials, and give developers a scoped way to use the service.

See Workspace access
Inside the Console

Operate the service without losing the details.

Inspect model readiness, test inference, and keep Workspace access explicit. These captures come from the official Orchard project documentation.

Orchard Console Playground showing a completed local model inference and request metrics
PlaygroundRun an inference and inspect its result before integrating an application.
View full size
Orchard Console Model Catalog showing imported models and lifecycle states
Model CatalogReview imported models and their lifecycle state.
View full size
Orchard Console Workspace model access showing explicit model grants
Workspace accessKeep model grants separate from catalog and runtime state.
View full size
Orchard demo: Local AI, shared by your team
See Orchard

Local AI, shared by your team.

Choose a model, use the Playground, and follow a live Qwen3.8 request from generation through its execution details.

Product walkthrough1:11
Watch on YouTube
Request path

Follow one inference request through Orchard.

Your application keeps control of its tools and data. Orchard manages the inference request and the model worker.

Request and controlEvents and status
YOUR APPLICATION

Client or agent

Sends a request. It keeps control of tools, permissions, and side effects.

1
ORCHARD CONTROL PLANE

BEAM Controller

Checks access, records the deadline, selects a Node, and tracks the outcome.

Access and policyModel routingRequest recordRetry controls
2
PRIVATE EXECUTION

Trusted Node

The Node Agent supervises model workers and reports current capacity.

Node Agenthealth · capacity
MLX workerload · generate · stream
Pinned modelexact revision
Streamed events, status, and the final outcome return to the client.
Available nowApple Silicon macOS, MLX, OpenAI-compatible APIs, model and Node operations, tenant access, and request records.
Planned workThe worker interface is designed to support other runtimes. Linux accelerator support is still in development and is not a supported deployment today.
System layers

See where each Orchard responsibility runs.

This view separates the public API, the control plane, Node supervision, and model execution.

ENTRY POINT

API surface

ApplicationsChat and Responses clients
Developer PortalCredentials and model access
Operator ConsoleCluster administration
PORTABLE SERVICES

BEAM Controller

Access and policyAuthentication and tenant grants
SchedulingPlacement, capacity, and retry
Request recordsAttempts, timing, and outcomes
TRUSTED MACHINE

Node Agent

Worker supervisionStart, stop, health, and cleanup
CapacityAdmission, cancellation, and release
Model placementFiles, loaded state, and eligibility
MODEL EXECUTION

Worker Runtime

MLX workerCurrent Apple Silicon runtime
Pinned artifactExact model revision
Streamed eventsText, tools, usage, errors, and final state
Install from source

Bring up Orchard on an Apple Silicon Mac.

Choose a coding-agent setup or follow the manual path. Both use Orchard’s pinned toolchain and lead to the same source-development environment.

Experimental source-only publication

Orchard currently provides no official binary, supported release, SLA, or maintenance commitment. The public quickstart targets Apple Silicon macOS.

Recommended

Set up with a coding agent

Give this prompt to a coding agent on the Mac that will run Orchard.

Agent prompt
Set up https://github.com/kapitan-ai/orchard on this Mac.
Clone the repository if needed, read AGENTS.md, docs/tooling.md,
and docs/local-dev.md, and inspect the machine and any existing
Orchard installation. Install and configure prerequisites using
the pinned toolchain and repository setup, including optional MLX
dependencies. Start source development with the default single-node
runtime. Prepare a compatible model bundle, create a Workspace and
API token, grant explicit model access, and verify a real API response.
If testing the Playground, grant Playground access explicitly.
Keep credentials private, preserve existing data, and report the
Console URL plus the commands to stop and restart.
Manual setup

Use the pinned development environment

First install mise and PostgreSQL 15+, then prepare a local Orchard Model Bundle.

  1. 1
    Clone and install
    Terminal
    git clone https://github.com/kapitan-ai/orchard.git
    cd orchard
    make setup
    mise exec -- uv sync --locked \
      --directory native/orchard_worker_mlx --extra mlx
    make dev
  2. 2
    Prepare access in the running IEx session
    IEx
    OrchardCLI.main(["models", "import", "/path/to/model-bundle", "--activate"])
    OrchardCLI.main(["tenants", "create", "--slug", "dev", "--name", "Dev"])
    OrchardCLI.main(["api-keys", "create", "--tenant-id", "<tenant-id>", "--name", "dev"])
    OrchardCLI.main(["models", "access", "grant", "<model_id>@<version>", "--tenant", "dev"])
  3. 3
    Verify real inference

    Securely read the one-time token, then call /v1/chat/completions with the exact imported model ID and version.

    Terminal
    printf 'Orchard API Token: ' >&2
    read -r -s ORCHARD_API_KEY
    export ORCHARD_API_KEY
    printf '\n' >&2
    
    curl -X POST http://localhost:4000/v1/chat/completions \
      -H 'Content-Type: application/json' \
      -H "Authorization: Bearer ${ORCHARD_API_KEY}" \
      -d '{
        "model": "<model_id>@<version>",
        "messages": [{"role": "user", "content": "Hello from Orchard"}]
      }'
Consolehttp://localhost:4000Node Agent gRPC127.0.0.1:50071

Starting the service alone is not inference-ready: import and activate a model, create a Workspace and API token, then grant explicit model access.

Open-source project

Orchard is available as an experimental pre-release.

The source, specifications, design decisions, and validation results are public. We qualify support against specific models, runtime versions, hardware, and workloads.