# HES-only customer router template Deployable source template, not a hosted tenant or an activated paid business. Your operator provides their own existing HES account/key. Purchases and API calls go only to `https://api.hesoyam.business/v1`; no supplier fallback, account creation, wallet signing, payments or paid installation calls exist. Defaults expose only `hs-mini`, the small local Qwen2.5-0.5B q4_0 integration demo. Quality, capacity and the upstream free allowance are limited; each customer's local quota shares your one upstream account allowance. No premium availability or customer income is promised. ## Initial setup Docker Engine and Docker Compose v2 are prerequisites. Unpack and enter this directory. `python3 install.py` prints a plan without writes or network. Explicit setup is one command: ```sh python3 install.py --configure --brand 'My AI router' ``` Paste your existing HES key into the hidden terminal prompt. Alternatively supply `--key-file /absolute/private/hes.key` (owned regular mode-600 file). Never put keys in argv, URLs, browser storage, screenshots, compose environment or logs. The assistant creates private `data/` mode700, `upstream.secret` and `config.json` mode600, and `.env` containing only your numeric UID/GID. No upstream request occurs. It refuses overwrites. Separately start when ready: ```sh docker compose up --build -d ``` This explicit build downloads the Python base image and pinned aiohttp dependency; it does not purchase inference. The Docker image/Compose/Caddy deployment has not been executed in the supplied offline proof. The container runs as your host UID/GID (run the installer as a non-root operator), mounts only its own `data/`, and publishes `127.0.0.1:8080`. One replica only; never share the SQLite directory between deployed routers. Caddyfile.example demonstrates an HTTPS proxy for a domain you own. HTTPS and DNS are your operator step. Do not expose Docker's port directly. `GET /health` checks only this router, not HES availability. Access logs are disabled. Back up the private data directory securely, including the SQLite ledger; restoring an older ledger can restore spent quotas. Keep clocks sane. Never archive `data/`, `.env`, generated customer keys or private backups. ## Customer keys and local administration Issue a different random key for each customer; only its SHA256 hash is stored in SQLite. Plaintext is written once to a new private file, never returned again by the API. Retrieve it privately and store it in the customer's secret store; remove the transfer file after handover. HES master key is never handed to customers. ```sh docker compose exec router python router.py issue --label client-a --models hs-mini --max-requests 5 --max-units 12000 --key-out data/client-a.key ``` The command prints an opaque local key ID and file path, not the secret. Administrative issue/revoke/scope/budget/usage commands are local CLI only; there is no HTTP admin endpoint. Use the printed ID: ```sh docker compose exec router python router.py usage --id KEY_ID docker compose exec router python router.py budget --id KEY_ID --max-requests 8 --max-units 18000 docker compose exec router python router.py scopes --id KEY_ID --models hs-mini docker compose exec router python router.py revoke --id KEY_ID ``` Revocation is rechecked atomically before admission. A request already admitted upstream cannot be cancelled/refunded by revocation. Scope/budget changes persist. Budgets cannot be set below already used amounts. Revocation preserves account usage history. CLI paths are private; do not put personal details in labels. ## API integration Authenticated routes: `GET /v1/models`, `GET /v1/usage`, `POST /v1/chat/completions`. Only string text messages with system/user/assistant roles and model/messages/max_tokens/temperature/stream are supported. Tools, images, audio, `n`, external URLs, unsupported parameters and extra message fields are rejected. The model catalogue is limited to that customer's scope and checks the upstream advertised price on each request; price drift blocks a priced route. The requested model ID is passed unchanged; a different returned model is rejected. HES IDs can be service aliases: this template does not attest underlying supplier, weights, quantization or equivalence to another provider. It never silently substitutes models or routes via prompts. Smart routing/cache and reference-price discount tiers are future features, not this release. Use any OpenAI-compatible client that supports this subset: ```python import os from openai import OpenAI client = OpenAI(base_url="https://router.example.com/v1", api_key=os.environ["CUSTOMER_ROUTER_KEY"]) # This explicitly sends ONE inference request; it consumes both local admission quota and HES allowance. result = client.chat.completions.create( model="hs-mini", messages=[{"role":"user", "content":"Explain your service briefly."}], max_tokens=64, stream=False) print(result.choices[0].message.content) ``` OpenAI SDK is an integration example, not a template dependency or a runtime proven in the offline tests. With `stream=true`, the router waits for one valid NONSTREAM upstream completion, then emits buffered SSE: content, finish/usage and `[DONE]`. It is compatible formatting, not incremental token streaming or fast first-token delivery. Invalid completion, error, timeout or disconnect never fabricates successful `[DONE]`. Error text/headers from upstream are not forwarded. Client and HES key-shaped strings are redacted from validated response content. ## Quotas, pricing and payments One request and a pessimistic **estimated-token-unit reservation** are permanently consumed at admission. Units = serialized UTF8 message bytes +256 per message +128 +requested max output. This is a local abuse budget, not an exact tokenizer, invoice, money balance or maximum upstream cost. Reported upstream token usage is recorded separately and must be finite nonnegative integer counts within bounds. It does not refund/debit local units. Admitted timeouts/failures/crashes consume the reservation. No upstream automatic inference retry. Clients must understand that their own retries create new attempts. Two in-flight requests maximum by default; input/output sizes and total timeout are bounded. Crash-interrupted admissions remain consumed and are marked interrupted on restart. Requests and SQLite admission transactions are isolated per key. Single-process lock prevents a second server on this directory. Customer money billing, customer crypto checkout, credit debits, payouts and revenue split **are not implemented**. An operator may manually allocate local request/unit budgets after agreeing a service with their customer; this template never interprets that as verified payment. HES account portal is `https://api.hesoyam.business/agents`; funding there remains subject to actual service availability. This template does not create invoices/sponsor-links or activate HES payments. Default `paid_calls_enabled=false`, `quote_enabled=false`, scope hs-mini only. For future verified paid supply, an operator must deliberately configure exact allowed model IDs, `verified_paid_prices` with input/output USD per1M, both flags, and `markup_factor` (1–10). Each catalogue check must exactly match verified upstream prices; wrong/negative/nonfinite/missing prices fail closed. Retail quote = upstream published rate ×factor. This is a price quote only; variable reasoning/cache/fees and monetary customer accounting require a separate reviewed implementation. This release is not ready for an unbounded prepaid money product. ## Promotion Use truthful public demos, opt-in referrals and public campaign links to your own product. HES offers its own `/partners` and `/agents` pages; partner attribution must use links actually issued by that service. This router does not post, DM, harvest contacts, award recruiting payouts or interpret campaign query strings as payment. Hosted setup can be quoted separately with defined deliverables and acceptance. No OpenClaw runtime or automatic profitable business has been tested/promised.