Self-hosted inference · project-operated

A large language model you can point your tools at tokens.velella.ca

An OpenAI-compatible endpoint serving Qwen3.6-35B, running on a single GPU on the project's own cloud allocation. Your key, your usage, your prompts staying on infrastructure we operate.

Available now · evaluation service

This is a temporary resource for exploring use cases and capabilities. It is a pilot, not a production service. There is no uptime guarantee, the model and the limits may change, and the endpoint may be taken down when the exploration phase ends. Use it to find out what a self-hosted open-weight model can do for your work — do not build anything that depends on it staying up.

What this is

One L40S GPU for an evaluation period, serving one open-weight model over a standard API. This is for evaluation and testing self-hosted models for teaching, in research code, in day-to-day work — with real users rather than benchmarks.

Treat it as a service that is here now and not promised forever. The evaluation has an end; what happens after it depends partly on what this period shows. Nothing you build against it is wasted: it speaks the same API as every commercial provider, so pointing a tool somewhere else later is a change of two settings.

What it serves right now

Model qwen3.6-35b-a3b — Qwen3.6-35B-A3B, open weights (Apache-2.0), served by vLLM
Endpointhttps://tokens.velella.ca/v1 over TLS
APIOpenAI-compatible: /v1/chat/completions, /v1/models
Streamingyes — server-sent events, as you would expect
Tool callingyes — function and tool calls work, including agentic loops. Give a tool-using request at least 2,048 output tokens, or the reasoning can truncate the call itself.
Context windowup to 65,536 tokens on the server; 32,768 on a standard key
Concurrency2 simultaneous requests per key by default
Claude Code
and Anthropic clients
not yet. The server already speaks the Anthropic API internally; one routing change remains. Until then, use an OpenAI-compatible client — the list below is long enough that this is rarely a real constraint.

The three settings

Every OpenAI-compatible tool asks for these, in its own dialect.

base URL
https://tokens.velella.ca/v1
model
qwen3.6-35b-a3b
api key
the contents of your own key file

Getting access

Access is per person, by key. You generate the key yourself and never send it to anyone — only a fingerprint of it, which cannot be turned back into the key. That is what makes revocation safe and attribution honest: usage is yours, and access can be withdrawn from the server side without anybody having to trust a shared secret.

  1. Make your key — it stays on your machine

    mkdir -p ~/.config/spigot && umask 177
    openssl rand -hex 32 | tr -d '\n' > ~/.config/spigot/key
  2. Send the fingerprint, not the key

    tr -d '\n' < ~/.config/spigot/key | sha256sum

    Send the 64-character output to Falk Herwig. It is a one-way digest: it enrols you and reveals nothing.

  3. Wait for the confirmation, then point your tools at it

    You will hear when the key is enrolled. A standard key is capped at 32,768 tokens of context and 2 concurrent requests — comfortable for editors, agents and scripts; if your work needs more, ask, and say what for.

Set a bounded max_tokens — the one setting that trips everyone. This model reasons before it answers, so both extremes fail:

Too small (say 512): the hidden reasoning consumes the whole budget and you get an empty reply.

Unset, or too large: many clients then ask for the model's entire context window as output, leaving no room for your prompt — and you get a maximum context length error before the request even runs. Open WebUI and some CLI clients do this out of the box.

Use 2,048–8,192 for most work. For short mechanical tasks you can also switch the reasoning off by adding "chat_template_kwargs": {"enable_thinking": false} to the request body.

First — check your key works

Before configuring anything, ask the server what it is serving. One line, and the reply tells you more than it looks.

KEY=$(cat ~/.config/spigot/key)
curl https://tokens.velella.ca/v1/models -H "Authorization: Bearer $KEY"

What you should get back — a single line of JSON listing exactly one model:

{"object":"list","data":[{"id":"qwen3.6-35b-a3b","object":"model",
 "created":1787116060,"owned_by":"vllm","root":"/models","parent":null,
 "max_model_len":65536,"permission":[{...}]}]}

That confirms four things at once: the endpoint is reachable over TLS, your key is enrolled and accepted, the model is loaded and serving, and its id is qwen3.6-35b-a3b — the exact string every tool below must send. The permission block and owned_by are the serving software's own bookkeeping; ignore them.

One number there can mislead you. max_model_len:65536 is what the server can hold — not what your key may ask for. A standard key is capped at 32,768 tokens of context, and a request above your cap is refused however large that number looks.

Then check it can actually answer

Listing models proves your key; it does not prove the model generates. This does:

curl https://tokens.velella.ca/v1/chat/completions \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model":"qwen3.6-35b-a3b",
       "messages":[{"role":"user","content":"In one sentence: what is a tide?"}],
       "max_tokens":1024}'

Expect a JSON object with the answer in choices[0].message.content and a usage block counting tokens. Ten to fifteen seconds is normal — the model reasons before it replies.

If it does not work

What you seeWhat it means
401 / unauthorized The key is not enrolled, or what you sent is not what you enrolled. Re-run the fingerprint command and compare it against the digest you submitted.
404 Usually a base URL missing its /v1 — or an Anthropic-API path, which is not live yet.
An empty answer Almost always max_tokens set too low: the reasoning consumed the whole budget.
A long pause, then an answer Working as intended — and somebody else may be using the card at the same time.

Pointing tools at it

Anything that speaks the OpenAI API works. Which one you want depends on how you like to work — in your editor, in a terminal, in a browser, or in your own code.

Every example below reads your key from the one place the setup put it: ~/.config/spigot/key. Copy and paste them as they are — nothing to fill in.

Two tools (Continue and Cline) want the key typed into a settings file or a settings box and cannot read a file for you. Put it on the clipboard rather than retyping it:

pbcopy < ~/.config/spigot/key      # macOS
xclip -sel clip < ~/.config/spigot/key   # Linux

That is safe and it is not a contradiction of step 2. Your key belongs in your own config, on your own machine — that is what it is for. The rule is only that it never travels to anyone else: what you send us is the fingerprint, never this.

Two habits worth keeping: make the file private (chmod 600 ~/.config/spigot/key), and in shell or Docker commands read it with $(cat …) rather than typing the key out — a literal key ends up in your shell history, the file reference does not.

Continue VS Code · JetBrains

An open-source assistant that lives inside your editor: a chat panel beside your code, select-and-ask edits, and autocomplete. You stay the driver — it suggests, you apply.

Best for everyday work in a file you are already looking at: explain this function, write this test, tidy this block.

Get it: continue.dev · VS Code Marketplace

#  ~/.continue/config.yaml   (then reload VS Code)
models:
  - name: Qwen3.6-35B (tokens)
    provider: openai
    model: qwen3.6-35b-a3b
    apiBase: https://tokens.velella.ca/v1
    apiKey: 6f3c...   # the 64-character contents of ~/.config/spigot/key
    roles: [chat, edit]

Cline VS Code

An agent rather than an assistant: give it a task and it plans, reads and edits several files, and runs commands — asking your approval at each step. Formerly "Claude Dev".

Best for multi-file jobs you want to supervise rather than type: a refactor, a bug hunt across modules, wiring something new in. It reads a lot, so it is the heaviest user of your key's context.

Get it: cline.bot · VS Code Marketplace

Extension settings -> API Provider: OpenAI Compatible

Base URL:  https://tokens.velella.ca/v1
API Key:   (paste ~/.config/spigot/key -- the key itself, not the fingerprint)
Model ID:  qwen3.6-35b-a3b

OpenCode terminal · agent

An open-source terminal coding agent. Like Cline it plans and edits across files, but it runs in the shell rather than in an editor — so it works over ssh, on a cluster login node, or anywhere you would use a terminal.

Best for agentic work without an editor, and for anyone who wants the key kept in a file rather than pasted into settings — it is the only client here that reads the key from disk instead of storing a copy of it.

Get it: opencode.ai · provider docs — npm i -g opencode-ai (needs Node)

# ~/.config/opencode/opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "velella": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "tokens.velella.ca",
      "options": {
        "baseURL": "https://tokens.velella.ca/v1",
        "apiKey": "{file:~/.config/spigot/key}"
      },
      "models": {
        "qwen3.6-35b-a3b": {
          "name": "Qwen3.6-35B",
          "reasoning": true,
          "interleaved": { "field": "reasoning_content" },
          "limit": { "context": 32768, "output": 8192 }
        }
      }
    }
  }
}

Run a task with opencode for the interactive interface, or headless:

opencode run --auto -m velella/qwen3.6-35b-a3b "write primes.py, run it, show the output"

Two lines in that config are doing real work. "reasoning" with reasoning_content shows the model's thinking before the answer, so a long turn reads as work rather than as a hang — this model thinks for a while before it writes, and OpenCode is the client here that shows you that. "limit": {"output": 8192} sets the max_tokens cap from the box above once, in config, instead of per request.

Leave "output" at 8192 or higher. Lowering it is the one change that looks like a sensible economy and is not: the thinking phase draws on the same budget, so a small cap is spent before the answer starts and you get an empty reply. Tested — 8192 is safe.

The model id must match what /v1/models returns — qwen3.6-35b-a3b, exactly as above. The interactive interface is OpenCode's solid path; the headless run command has hung on long tasks in our testing, so check on it if you script against it.

Aider terminal

A pair-programmer in the shell. It knows your git repo, edits files in place, and commits each change, so every step is reviewable and revertible with the tools you already use.

Best for repo-scale changes from the command line, and for working over ssh on a cluster or server where an editor is not practical.

Get it: aider.chat — pip install aider-install

export OPENAI_API_BASE=https://tokens.velella.ca/v1
export OPENAI_API_KEY=$(cat ~/.config/spigot/key)
aider --model openai/qwen3.6-35b-a3b

Open WebUI browser · self-hosted

A ChatGPT-style web interface you run yourself. No coding involved once it is up: open a browser, type, get answers, keep conversations.

Best for people who do not want a terminal or an editor at all — and for putting one shared chat window in front of a group.

Get it: openwebui.com · source (needs Docker)

docker run -d -p 3000:8080 \
  -e OPENAI_API_BASE_URL=https://tokens.velella.ca/v1 \
  -e OPENAI_API_KEY=$(cat ~/.config/spigot/key) \
  -v open-webui:/app/backend/data \
  --name open-webui --restart always ghcr.io/open-webui/open-webui:main

Then open http://localhost:3000. The model appears in the model list.

It will ask you to create an account with a name, an email and a password. That is not a sign-up. Open WebUI is a multi-user application that you are running yourself, so it has its own login; the account is written to the container's database in your open-webui volume and goes nowhere else. The email is only an identifier — nothing is sent to it and nothing verifies it. By their design the first account created becomes the administrator, and any later sign-up on that instance stays pending until the administrator approves it.

Set Max Tokens before your first prompt, or it will fail. Open WebUI does not cap the answer by default, so it asks for the entire 65,536-token window and leaves no room for your question — you get a maximum context length error before the request even reaches the server. In a chat, open Controls (the sliders, top right) → Advanced Params → Max Tokens (num_predict) and set 4096; or set it once per model under Admin Panel → Settings → Models. This is the client's default, not a limit of the endpoint.

Web search is the client's job, not the model's. No language model browses on its own; searching works only when your client offers a search tool and runs it. In Open WebUI: Admin Panel → Settings → Web Search, choose a provider (DuckDuckGo needs no key), then switch Web Search on in the chat. The endpoint supports tool-calling; it does not supply the tools.

Give the password real thought, though. Whoever can open that page can spend your key — it is stored in the container, not typed per message. If you are the only user, bind it to your own machine only, which is one edit to the command: -p 127.0.0.1:3000:8080. Without the 127.0.0.1, Docker publishes the port on every interface.

The mismatched numbers are not a typo. 3000:8080 means "port 3000 on your machine maps to 8080 inside the container", and the container serves on 8080 — so 3000:3000 gives you a browser pointed at nothing. If the page does not load: docker ps (is it running?) and docker logs open-webui (did it start?). One honest caveat: $(cat …) keeps the key out of your shell history, but it is still readable in the container's environment via docker inspect — inherent to giving a container a credential, so run this on a machine you control.

The OpenAI SDKs python · node

The official client libraries, pointed at this endpoint instead of OpenAI's. Everything built on them — LangChain, LlamaIndex, your own scripts — works the same way.

Best for building something: batch jobs, analysis pipelines, notebooks, a tool of your own.

Get it: pip install openai · openai-python · npm i openai

import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.velella.ca/v1",
                api_key=open(os.path.expanduser("~/.config/spigot/key")).read().strip())
r = client.chat.completions.create(
        model="qwen3.6-35b-a3b",
        messages=[{"role": "user", "content": "hello"}],
        max_tokens=1024)
print(r.choices[0].message.content)
import OpenAI from "openai";
import { readFileSync } from "fs";
import { homedir } from "os";

const client = new OpenAI({
  baseURL: "https://tokens.velella.ca/v1",
  apiKey: readFileSync(`${homedir()}/.config/spigot/key`, "utf8").trim(),
});
const r = await client.chat.completions.create({
  model: "qwen3.6-35b-a3b",
  messages: [{ role: "user", content: "hello" }],
  max_tokens: 1024,
});

What will not work (yet) for now

GitHub Copilot cannot be pointed at another endpoint at all — it is tied to GitHub's own models. Use Continue or Cline in the same editor instead.

Claude Code and other Anthropic-API clients are not live yet. The server already speaks that API internally and one routing change remains; you will hear when it is on.

What to expect

Speed Around 80–100 tokens/second when you have the card to yourself. It drops as more people work at once — that is the nature of one GPU, not a fault.
A pause before the answer The model reasons first, so a turn can take 10–15 seconds and an agentic task making several calls takes proportionally longer. Expected.
How busy it gets Comfortable at roughly 5–10 actively-working streams. Agent-style tools idle most of the time, so a good many enrolled users coexist happily.
Privacy Prompts transit the server's logs (rotated on the order of two weeks) on a machine the project operates. Fine for work; do not send anything you would not put in a work email.
Search, code execution,
file access
Features of your client, not of the model. It can only use a tool your tool gives it and runs. If a client cannot search the web, that is a setting in the client — not a restriction here.
If something breaks Tell us what you sent and what came back. An empty reply is nearly always the max_tokens point above.