Self-hosted inference · project-operated
An OpenAI-compatible endpoint serving Qwen3.6-35B, running on a single GPU on the project's own cloud allocation. Your key, your usage, your prompts staying on infrastructure we operate.
Available now · evaluation serviceThis is a temporary resource for exploring use cases and capabilities. It is a pilot, not a production service. There is no uptime guarantee, the model and the limits may change, and the endpoint may be taken down when the exploration phase ends. Use it to find out what a self-hosted open-weight model can do for your work — do not build anything that depends on it staying up.
One L40S GPU for an evaluation period, serving one open-weight model over a standard API. This is for evaluation and testing self-hosted models for teaching, in research code, in day-to-day work — with real users rather than benchmarks.
Treat it as a service that is here now and not promised forever. The evaluation has an end; what happens after it depends partly on what this period shows. Nothing you build against it is wasted: it speaks the same API as every commercial provider, so pointing a tool somewhere else later is a change of two settings.
| Model | qwen3.6-35b-a3b — Qwen3.6-35B-A3B, open weights (Apache-2.0), served by vLLM |
|---|---|
| Endpoint | https://tokens.velella.ca/v1 over TLS |
| API | OpenAI-compatible: /v1/chat/completions, /v1/models |
| Streaming | yes — server-sent events, as you would expect |
| Tool calling | yes — function and tool calls work, including agentic loops. Give a tool-using request at least 2,048 output tokens, or the reasoning can truncate the call itself. |
| Context window | up to 65,536 tokens on the server; 32,768 on a standard key |
| Concurrency | 2 simultaneous requests per key by default |
| Claude Code and Anthropic clients |
not yet. The server already speaks the Anthropic API internally; one routing change remains. Until then, use an OpenAI-compatible client — the list below is long enough that this is rarely a real constraint. |
Every OpenAI-compatible tool asks for these, in its own dialect.
Access is per person, by key. You generate the key yourself and never send it to anyone — only a fingerprint of it, which cannot be turned back into the key. That is what makes revocation safe and attribution honest: usage is yours, and access can be withdrawn from the server side without anybody having to trust a shared secret.
mkdir -p ~/.config/spigot && umask 177
openssl rand -hex 32 | tr -d '\n' > ~/.config/spigot/key
tr -d '\n' < ~/.config/spigot/key | sha256sum
Send the 64-character output to Falk Herwig. It is a one-way digest: it enrols you and reveals nothing.
You will hear when the key is enrolled. A standard key is capped at 32,768 tokens of context and 2 concurrent requests — comfortable for editors, agents and scripts; if your work needs more, ask, and say what for.
Set a bounded max_tokens — the one setting that
trips everyone. This model reasons before it answers, so both extremes fail:
Too small (say 512): the hidden reasoning consumes the whole budget and you get an empty reply.
Unset, or too large: many clients then ask for the
model's entire context window as output, leaving no room for your prompt — and you get a
maximum context length error before the request even runs. Open WebUI and some CLI
clients do this out of the box.
Use 2,048–8,192 for most work. For short mechanical tasks you
can also switch the reasoning off by adding
"chat_template_kwargs": {"enable_thinking": false} to the request body.
Before configuring anything, ask the server what it is serving. One line, and the reply tells you more than it looks.
KEY=$(cat ~/.config/spigot/key)
curl https://tokens.velella.ca/v1/models -H "Authorization: Bearer $KEY"
What you should get back — a single line of JSON listing exactly one model:
{"object":"list","data":[{"id":"qwen3.6-35b-a3b","object":"model",
"created":1787116060,"owned_by":"vllm","root":"/models","parent":null,
"max_model_len":65536,"permission":[{...}]}]}
That confirms four things at once: the endpoint is reachable over TLS,
your key is enrolled and accepted, the model is loaded and serving, and its id
is qwen3.6-35b-a3b — the exact string every tool below must send. The
permission block and owned_by are the serving software's own
bookkeeping; ignore them.
One number there can mislead you.
max_model_len:65536 is what the server can hold — not what your key
may ask for. A standard key is capped at 32,768 tokens of context, and a request above your cap
is refused however large that number looks.
Listing models proves your key; it does not prove the model generates. This does:
curl https://tokens.velella.ca/v1/chat/completions \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"model":"qwen3.6-35b-a3b",
"messages":[{"role":"user","content":"In one sentence: what is a tide?"}],
"max_tokens":1024}'
Expect a JSON object with the answer in choices[0].message.content and a
usage block counting tokens. Ten to fifteen seconds is normal — the
model reasons before it replies.
| What you see | What it means |
|---|---|
| 401 / unauthorized | The key is not enrolled, or what you sent is not what you enrolled. Re-run the fingerprint command and compare it against the digest you submitted. |
| 404 | Usually a base URL missing its /v1 — or an Anthropic-API path, which is not
live yet. |
| An empty answer | Almost always max_tokens set too low: the reasoning consumed the whole budget. |
| A long pause, then an answer | Working as intended — and somebody else may be using the card at the same time. |
Anything that speaks the OpenAI API works. Which one you want depends on how you like to work — in your editor, in a terminal, in a browser, or in your own code.
Every example below reads your key
from the one place the setup put it: ~/.config/spigot/key. Copy and paste
them as they are — nothing to fill in.
Two tools (Continue and Cline) want the key typed into a settings file or a settings box and cannot read a file for you. Put it on the clipboard rather than retyping it:
pbcopy < ~/.config/spigot/key # macOS
xclip -sel clip < ~/.config/spigot/key # Linux
That is safe and it is not a contradiction of step 2. Your key belongs in your own config, on your own machine — that is what it is for. The rule is only that it never travels to anyone else: what you send us is the fingerprint, never this.
Two habits worth keeping: make the file private
(chmod 600 ~/.config/spigot/key), and in shell or Docker commands read it with
$(cat …) rather than typing the key out — a literal key ends up in your shell
history, the file reference does not.
An open-source assistant that lives inside your editor: a chat panel beside your code, select-and-ask edits, and autocomplete. You stay the driver — it suggests, you apply.
Best for everyday work in a file you are already looking at: explain this function, write this test, tidy this block.
Get it: continue.dev · VS Code Marketplace
# ~/.continue/config.yaml (then reload VS Code)
models:
- name: Qwen3.6-35B (tokens)
provider: openai
model: qwen3.6-35b-a3b
apiBase: https://tokens.velella.ca/v1
apiKey: 6f3c... # the 64-character contents of ~/.config/spigot/key
roles: [chat, edit]
An agent rather than an assistant: give it a task and it plans, reads and edits several files, and runs commands — asking your approval at each step. Formerly "Claude Dev".
Best for multi-file jobs you want to supervise rather than type: a refactor, a bug hunt across modules, wiring something new in. It reads a lot, so it is the heaviest user of your key's context.
Get it: cline.bot · VS Code Marketplace
Extension settings -> API Provider: OpenAI Compatible
Base URL: https://tokens.velella.ca/v1
API Key: (paste ~/.config/spigot/key -- the key itself, not the fingerprint)
Model ID: qwen3.6-35b-a3b
An open-source terminal coding agent. Like Cline it plans and edits across files, but it runs in the shell rather than in an editor — so it works over ssh, on a cluster login node, or anywhere you would use a terminal.
Best for agentic work without an editor, and for anyone who wants the key kept in a file rather than pasted into settings — it is the only client here that reads the key from disk instead of storing a copy of it.
Get it: opencode.ai ·
provider docs —
npm i -g opencode-ai (needs Node)
# ~/.config/opencode/opencode.json
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"velella": {
"npm": "@ai-sdk/openai-compatible",
"name": "tokens.velella.ca",
"options": {
"baseURL": "https://tokens.velella.ca/v1",
"apiKey": "{file:~/.config/spigot/key}"
},
"models": {
"qwen3.6-35b-a3b": {
"name": "Qwen3.6-35B",
"reasoning": true,
"interleaved": { "field": "reasoning_content" },
"limit": { "context": 32768, "output": 8192 }
}
}
}
}
}
Run a task with opencode for the interactive
interface, or headless:
opencode run --auto -m velella/qwen3.6-35b-a3b "write primes.py, run it, show the output"
Two lines in that config are doing real work.
"reasoning" with reasoning_content shows the model's thinking before
the answer, so a long turn reads as work rather than as a hang — this model thinks for a while
before it writes, and OpenCode is the client here that shows you that.
"limit": {"output": 8192} sets the max_tokens cap from the box above
once, in config, instead of per request.
Leave "output" at 8192 or higher. Lowering it is the one change
that looks like a sensible economy and is not: the thinking phase draws on the same budget, so a
small cap is spent before the answer starts and you get an empty reply. Tested — 8192 is safe.
The model id must match what
/v1/models returns — qwen3.6-35b-a3b, exactly as above. The
interactive interface is OpenCode's solid path; the headless run command has hung
on long tasks in our testing, so check on it if you script against it.
A pair-programmer in the shell. It knows your git repo, edits files in place, and commits each change, so every step is reviewable and revertible with the tools you already use.
Best for repo-scale changes from the command line, and for working over ssh on a cluster or server where an editor is not practical.
Get it: aider.chat — pip install aider-install
export OPENAI_API_BASE=https://tokens.velella.ca/v1
export OPENAI_API_KEY=$(cat ~/.config/spigot/key)
aider --model openai/qwen3.6-35b-a3b
A ChatGPT-style web interface you run yourself. No coding involved once it is up: open a browser, type, get answers, keep conversations.
Best for people who do not want a terminal or an editor at all — and for putting one shared chat window in front of a group.
Get it: openwebui.com · source (needs Docker)
docker run -d -p 3000:8080 \
-e OPENAI_API_BASE_URL=https://tokens.velella.ca/v1 \
-e OPENAI_API_KEY=$(cat ~/.config/spigot/key) \
-v open-webui:/app/backend/data \
--name open-webui --restart always ghcr.io/open-webui/open-webui:main
Then open http://localhost:3000. The model appears in the
model list.
It will ask you to create an account with a name, an email and
a password. That is not a sign-up. Open WebUI is a multi-user application that you are
running yourself, so it has its own login; the account is written to the container's database
in your open-webui volume and goes nowhere else. The email is only an identifier —
nothing is sent to it and nothing verifies it. By their design the first account created becomes
the administrator, and any later sign-up on that instance stays pending until the administrator
approves it.
Set Max Tokens before your first prompt, or it will fail.
Open WebUI does not cap the answer by default, so it asks for the entire 65,536-token
window and leaves no room for your question — you get a maximum context length
error before the request even reaches the server. In a chat, open Controls
(the sliders, top right) → Advanced Params →
Max Tokens (num_predict) and set 4096; or set it once per
model under Admin Panel → Settings → Models. This is the client's default,
not a limit of the endpoint.
Web search is the client's job, not the model's. No language model browses on its own; searching works only when your client offers a search tool and runs it. In Open WebUI: Admin Panel → Settings → Web Search, choose a provider (DuckDuckGo needs no key), then switch Web Search on in the chat. The endpoint supports tool-calling; it does not supply the tools.
Give the password real thought, though. Whoever can
open that page can spend your key — it is stored in the container, not typed per message. If you
are the only user, bind it to your own machine only, which is one edit to the command:
-p 127.0.0.1:3000:8080. Without the 127.0.0.1, Docker publishes the
port on every interface.
The mismatched numbers are not a typo. 3000:8080
means "port 3000 on your machine maps to 8080 inside the container", and the container serves
on 8080 — so 3000:3000 gives you a browser pointed at nothing. If the page does not
load: docker ps (is it running?) and docker logs open-webui (did it
start?). One honest caveat: $(cat …) keeps the key out of your shell history, but
it is still readable in the container's environment via docker inspect — inherent
to giving a container a credential, so run this on a machine you control.
The official client libraries, pointed at this endpoint instead of OpenAI's. Everything built on them — LangChain, LlamaIndex, your own scripts — works the same way.
Best for building something: batch jobs, analysis pipelines, notebooks, a tool of your own.
Get it: pip install openai ·
openai-python ·
npm i openai
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.velella.ca/v1",
api_key=open(os.path.expanduser("~/.config/spigot/key")).read().strip())
r = client.chat.completions.create(
model="qwen3.6-35b-a3b",
messages=[{"role": "user", "content": "hello"}],
max_tokens=1024)
print(r.choices[0].message.content)
import OpenAI from "openai";
import { readFileSync } from "fs";
import { homedir } from "os";
const client = new OpenAI({
baseURL: "https://tokens.velella.ca/v1",
apiKey: readFileSync(`${homedir()}/.config/spigot/key`, "utf8").trim(),
});
const r = await client.chat.completions.create({
model: "qwen3.6-35b-a3b",
messages: [{ role: "user", content: "hello" }],
max_tokens: 1024,
});
GitHub Copilot cannot be pointed at another endpoint at all — it is tied to GitHub's own models. Use Continue or Cline in the same editor instead.
Claude Code and other Anthropic-API clients are not live yet. The server already speaks that API internally and one routing change remains; you will hear when it is on.
| Speed | Around 80–100 tokens/second when you have the card to yourself. It drops as more people work at once — that is the nature of one GPU, not a fault. |
|---|---|
| A pause before the answer | The model reasons first, so a turn can take 10–15 seconds and an agentic task making several calls takes proportionally longer. Expected. |
| How busy it gets | Comfortable at roughly 5–10 actively-working streams. Agent-style tools idle most of the time, so a good many enrolled users coexist happily. |
| Privacy | Prompts transit the server's logs (rotated on the order of two weeks) on a machine the project operates. Fine for work; do not send anything you would not put in a work email. |
| Search, code execution, file access |
Features of your client, not of the model. It can only use a tool your tool gives it and runs. If a client cannot search the web, that is a setting in the client — not a restriction here. |
| If something breaks | Tell us what you sent and what came back. An empty reply is nearly always the
max_tokens point above. |