Skip to main content

LLM and Model Providers

This page covers setting up inference providers for Nastech Agent — from cloud APIs like OpenRouter and Anthropic, to self-hosted endpoints like Ollama and vLLM, to advanced routing and fallback configurations. You need at least one provider configured to use Nastech.

Inference Providers​

You need at least one way to connect to an LLM. Use nastech model to switch providers and models interactively, or configure directly:

ProviderSetup
Nastech Portalnastech model (OAuth, subscription-based)
OpenAI Codexnastech model → ChatGPT or Codex Subscription (ChatGPT OAuth, uses Codex models)
GitHub Copilotnastech model (OAuth device code flow, COPILOT_GITHUB_TOKEN, GH_TOKEN, or gh auth token)
GitHub Copilot ACPnastech model (spawns local copilot --acp --stdio)
Anthropicnastech model (Claude Max + extra usage credits via OAuth; also supports Anthropic API key or manual setup-token — see note below)
OpenRouterOPENROUTER_API_KEY in ~/.nastech/.env, or nastech auth add openrouter --type oauth (browser login via OpenRouter's PKCE flow; stores a key in the credential pool)
Ramp RouterRAMP_ROUTER_API_KEY in ~/.nastech/.env (provider: router; aliases: ramp-router, ramp, router.com; Responses-native gateway, live account-scoped catalog)
Fireworks AIFIREWORKS_API_KEY in ~/.nastech/.env (provider: fireworks; aliases: fireworks-ai, fw)
NovitaAINOVITA_API_KEY in ~/.nastech/.env (provider: novita, 200+ models, Model API, Agent Sandbox, GPU Cloud)
AI GatewayAI_GATEWAY_API_KEY in ~/.nastech/.env (provider: ai-gateway)
z.ai / GLMGLM_API_KEY in ~/.nastech/.env (provider: zai)
Kimi / MoonshotKIMI_API_KEY in ~/.nastech/.env (provider: kimi-coding)
Kimi / Moonshot (China)KIMI_CN_API_KEY in ~/.nastech/.env (provider: kimi-coding-cn; aliases: kimi-cn, moonshot-cn)
Arcee AIARCEEAI_API_KEY in ~/.nastech/.env (provider: arcee; aliases: arcee-ai, arceeai)
GMI CloudGMI_API_KEY in ~/.nastech/.env (provider: gmi; aliases: gmi-cloud, gmicloud)
Nebius Token FactoryNEBIUS_API_KEY in ~/.nastech/.env (provider: nebius-token-factory; aliases: nebius, nebius-tf, tokenfactory)
Actual ComputerACTUAL_API_KEY in ~/.nastech/.env for the hosted relay; set model.base_url in config.yaml for a local daemon (no key needed on loopback). Provider: actual; aliases: actual-computer, actualcomputer, aci.
MiniMaxMINIMAX_API_KEY in ~/.nastech/.env (provider: minimax)
MiniMax ChinaMINIMAX_CN_API_KEY in ~/.nastech/.env (provider: minimax-cn)
xAI (Grok) — Responses APIXAI_API_KEY in ~/.nastech/.env (provider: xai)
xAI Grok OAuth (SuperGrok)nastech model → "xAI Grok OAuth (SuperGrok / Premium+)" — browser login, no API key. See guide
Qwen Cloud (Alibaba DashScope)DASHSCOPE_API_KEY in ~/.nastech/.env (provider: alibaba; mainland-China endpoint: alibaba-cn)
Alibaba Cloud (Coding Plan)ALIBABA_CODING_PLAN_API_KEY (falls back to DASHSCOPE_API_KEY) (provider: alibaba-coding-plan, alias: alibaba_coding; mainland-China endpoint: alibaba-coding-plan-cn with ALIBABA_CODING_PLAN_CN_API_KEY, falling back to the shared keys) — separate billing SKU, different endpoint
Alibaba Cloud (Token Plan)ALIBABA_TOKEN_PLAN_API_KEY in ~/.nastech/.env (provider: alibaba-token-plan; mainland-China endpoint: alibaba-token-plan-cn with ALIBABA_TOKEN_PLAN_CN_API_KEY, falling back to the shared key) — Model Studio flat-token tier
Kilo CodeKILOCODE_API_KEY in ~/.nastech/.env (provider: kilocode)
Xiaomi MiMoXIAOMI_API_KEY in ~/.nastech/.env (provider: xiaomi, aliases: mimo, xiaomi-mimo)
Tencent TokenHubTOKENHUB_API_KEY in ~/.nastech/.env (provider: tencent-tokenhub, aliases: tencent, tokenhub, tencentmaas)
Tencent TokenPlanTOKENPLAN_API_KEY in ~/.nastech/.env (provider: tencent-tokenplan, aliases: tokenplan, tencent-lkeap; Anthropic Messages endpoint)
OpenCode ZenOPENCODE_ZEN_API_KEY in ~/.nastech/.env (provider: opencode-zen)
CommandCodeCOMMANDCODE_API_KEY in ~/.nastech/.env (provider: commandcode, alias: commandcode-chat; Claude models via commandcode-anthropic, alias: commandcode-claude). Works with GOAT/Pro/Max/Provider plans (not the $1 Go plan — no API access).
OpenCode GoOPENCODE_GO_API_KEY in ~/.nastech/.env (provider: opencode-go)
DeepSeekDEEPSEEK_API_KEY in ~/.nastech/.env (provider: deepseek)
Hugging FaceHF_TOKEN in ~/.nastech/.env (provider: huggingface, aliases: hf)
Google / GeminiGOOGLE_API_KEY (or GEMINI_API_KEY) in ~/.nastech/.env (provider: gemini)
Google Vertex AInastech model → "Google Vertex AI" (provider: vertex; OAuth2 via service-account JSON or ADC, GCP billing)
OpenAI API (direct)OPENAI_API_KEY in ~/.nastech/.env (provider: openai-api, optional OPENAI_BASE_URL)
Azure AI Foundrynastech model → "Azure AI Foundry" (provider: azure-foundry; uses Azure OpenAI / Foundry endpoint and key)
AWS Bedrocknastech model → "AWS Bedrock" (provider: bedrock; standard AWS credentials chain via boto3)
NVIDIA BuildNVIDIA_API_KEY in ~/.nastech/.env (provider: nvidia; NIM-hosted models on build.nvidia.com)
Ollama Cloudnastech model → "Ollama Cloud" (provider: ollama-cloud; cloud-hosted Ollama API)
Qwen OAuthnastech model → "Qwen OAuth" (provider: qwen-oauth; browser PKCE login)
MiniMax OAuthnastech model → "MiniMax (OAuth)" (provider: minimax-oauth; browser PKCE login)
StepFunSTEPFUN_API_KEY in ~/.nastech/.env (provider: stepfun)
LM Studionastech model → "LM Studio" (provider: lmstudio, optional LM_API_KEY)
Custom Endpointnastech model → choose "Custom endpoint" (saved in config.yaml)

Both built-in OpenCode providers send an opaque, per-conversation x-opencode-session header on every request (main turns on every transport plus auxiliary calls such as compression, titles, approval checks, skills-hub lookups and /btw side questions — including the ones that run in the background after the turn has ended; headless Kanban specify/decompose and dashboard estimate calls use a per-task key). OpenCode uses it to pin a conversation to one backend so its prompt cache stays warm; the value is derived from the Nastech session id (or the Kanban task id) and carries no personal data.

The two built-in OpenCode providers each pin their own relay on opencode.ai (opencode-zen → /zen/v1, opencode-go → /zen/go/v1). A model.base_url left behind by the other relay is healed to the selected provider's relay, and the model you pick (-m, /model, a fallback entry or a channel override) decides which relay is used — so switching from a Zen model to a Go-only one never sends the request to Zen. A custom provider you define under providers: whose name extends a family slug (for example opencode-go-bridge) still gets the family's per-model API-mode routing and /v1 handling, but its base_url is taken as declared: name it after the relay it actually points at.

For the official API-key path, see the dedicated Google Gemini guide.

Model key alias

In the model: config section, you can use either default: or model: as the key name for your model ID. Both model: { default: my-model } and model: { model: my-model } work identically.

Nastech Portal​

Nastech Portal is Nastech Research's unified subscription gateway and the recommended way to run Nastech Agent. One OAuth login covers 300+ frontier agentic models (Claude, GPT, Gemini, DeepSeek, Qwen, Kimi, GLM, MiniMax, Grok, ...) plus the Tool Gateway (web search, image generation, TTS, browser automation) — billed against your Nastech subscription instead of separate per-provider accounts.

nastech setup --portal # fresh install — OAuth + provider + gateway in one command
nastech model # existing install — pick "Nastech Portal" from the list
nastech portal info # inspect login + routing at any time

Don't have a subscription yet? Get one at portal.nastechresearch.github.io/manage-subscription.

For full details: see the dedicated Nastech Portal integration page (what's in the subscription, model catalog, troubleshooting) and the step-by-step Run Nastech Agent with Nastech Portal guide.

Client identification. Every Portal request from Nastech Agent carries a client=nastech-client-v<version> tag (e.g. client=nastech-client-v0.13.0) auto-aligned to your installed release. This is sent on all Portal pathways — main chat loop, auxiliary calls, compression summarizer, web extraction — and lets Portal-side telemetry distinguish Nastech traffic from other clients. No config required; the tag updates automatically when you nastech update.

JWT auth (automatic). Nastech prefers scoped inference:invoke JWTs for Portal requests with the legacy opaque session-key path as a fallback. No configuration is required — credentials are managed by the OAuth flow and rotate transparently. Revoked refresh tokens are quarantined to avoid replay loops.

Codex Note

The OpenAI Codex provider authenticates via device code (open a URL, enter a code). Nastech stores the resulting credentials in its own auth store under ~/.nastech/auth.json and can import existing Codex CLI credentials from ~/.codex/auth.json when present. No Codex CLI installation is required. Automatic adoption of the Codex CLI login (when Nastech' own refresh fails) is controlled by auth.adopt_external_logins — see Borrowed CLI logins.

If a token refresh fails with a terminal error (HTTP 4xx, invalid_grant, revoked grant, etc.), Nastech marks the refresh token as dead and stops replaying it so you don't see a flood of identical auth failures. The next request surfaces a typed re-auth message instead. Run nastech auth add openai-codex (or nastech model → ChatGPT or Codex Subscription) to start a fresh device-code login; the quarantine clears on the next successful exchange.

Device login can fail with [SSL: UNEXPECTED_EOF_WHILE_READING] or a TLS handshake timeout on Python/OpenSSL 3.5+ when a middlebox rejects post-quantum groups such as X25519MLKEM768 (curl may still work). A one-off dropped connection is not fatal: while waiting for your browser approval Nastech keeps polling through up to six consecutive transport errors (and retries the device-code request and token exchange twice) before giving up, so only a persistently broken network surfaces this error. Nastech does not change default TLS policy. Point OPENSSL_CONF at a config that restricts Groups to classic curves before running nastech model, or diagnose with TLS 1.2:

openssl_conf = openssl_init

[openssl_init]
ssl_conf = ssl_sect

[ssl_sect]
system_default = system_default_sect

[system_default_sect]
Groups = x25519:secp256r1:secp384r1:x448
warning

Even when using Nastech Portal, Codex, or a custom endpoint, some tools (vision, web summarization, MoA) use a separate "auxiliary" model. By default (auxiliary.*.provider: "auto"), Nastech routes these tasks to your main chat model — the same model you picked in nastech model. You can override each task individually to route it to a cheaper/faster model (e.g. Gemini Flash on OpenRouter) — see Auxiliary Models.

Nastech Tool Gateway

Paid Nastech Portal subscribers also get access to the Tool Gateway — web search, image generation, TTS, and browser automation routed through your subscription. No extra API keys needed. On a fresh install, nastech setup --portal logs you in, sets Nastech as your provider, and turns the gateway on in one command. Existing users can enable it from nastech model or per-tool from nastech tools. Inspect routing at any time with nastech portal info.

Two Commands for Model Management​

Nastech has two model commands that serve different purposes:

CommandWhere to runWhat it does
nastech modelYour terminal (outside any session)Full setup wizard — add providers, run OAuth, enter API keys, configure endpoints
/modelInside a Nastech chat sessionQuick switch between already-configured providers and models

If you're trying to switch to a provider you haven't set up yet (e.g. you only have OpenRouter configured and want to use Anthropic), you need nastech model, not /model. Exit your session first (Ctrl+C or /quit), run nastech model, complete the provider setup, then start a new session.

Subscription plans: what your plan pays for​

Several providers let you sign in to Nastech with a consumer subscription (Claude Max, ChatGPT, SuperGrok / X Premium+, …) instead of an API key. What that subscription actually pays for — and what it doesn't — differs per provider, and it's the single most common source of billing surprises. The table below is the short version; each provider's own section has the details.

Cells marked not currently documented mean exactly that: Nastech docs do not yet specify the behavior. Don't assume — check your provider's billing dashboard, and treat these as open questions.

Plan / pathCan Nastech use it?What gets consumedWhat does NOT get consumedCommon surprise
Anthropic — Claude Max + OAuth✅ Yes — nastech model → Anthropic OAuth. Requires Max and purchased extra usage creditsThe extra/overage credits you've added on top of the Max planThe base Max plan allowance (the usage included in Claude Code by default)All Nastech usage bills as "extra usage" even while your included Max allowance sits untouched
Anthropic — Claude Pro❌ No — Pro subscribers cannot use the OAuth pathNothing (path unavailable)Your Pro subscriptionPro looks like it should work; it doesn't. Use an ANTHROPIC_API_KEY instead (pay-per-token, independent of any Claude subscription)
OpenAI Codex — ChatGPT plan OAuth✅ Yes — nastech model → ChatGPT or Codex Subscription (ChatGPT OAuth device-code login, uses Codex models)Not currently documentedNot currently documentedDocs cover auth and token refresh only; plan-quota semantics are not yet documented
xAI — SuperGrok / X Premium+ OAuth✅ Yes — browser OAuth, no API key neededYour subscription quota (documented explicitly for X Search: OAuth is preferred over an API key and "uses your subscription quota instead of API spend"). Inference quota semantics beyond that: not currently documentedXAI_API_KEY / pay-per-token API spend, when OAuth credentials are configured and preferredHTTP 403 after a successful login — xAI has restricted OAuth API access to specific SuperGrok tiers despite an active in-app subscription
Google — Gemini consumer plan (Google AI Pro / Ultra)❌ No documented path — the gemini provider is API-key only (GOOGLE_API_KEY / GEMINI_API_KEY); Vertex AI uses GCP billingYour API key's quota (free tier or billing-enabled Google Cloud project) — consumer-plan consumption not currently documentedNot currently documentedFree-tier keys can be exhausted after a handful of agent turns, because Nastech may make several model calls per user turn

Anthropic. The OAuth path routes as Claude Code against your Anthropic account and only works on a Claude Max plan with purchased extra usage credits — the base Max allowance is never consumed by Nastech, only the extra/overage credits on top. Claude Pro subscribers cannot use this path; the supported alternative is an ANTHROPIC_API_KEY, billed pay-per-token against that key's organization at standard API pricing. See Anthropic (Native) below.

OpenAI Codex. Nastech authenticates via ChatGPT device-code OAuth, stores credentials in ~/.nastech/auth.json, and can import existing Codex CLI credentials from ~/.codex/auth.json. Which ChatGPT plan tiers are eligible, and how Nastech usage counts against your plan's Codex limits, are not currently documented — the Codex note under Nastech Portal covers authentication and token-refresh behavior only.

xAI (SuperGrok / X Premium+). Browser OAuth works with either an active SuperGrok subscription or an X Premium+ subscription on the linked X account, and the same bearer token is reused by direct-to-xAI tools (TTS, image gen, video gen, transcription, X Search). If inference returns HTTP 403 after a successful login, that's a tier/entitlement restriction on xAI's side, not a stale token — the workaround is switching to an XAI_API_KEY. See xAI (Grok) below and the xAI Grok OAuth guide.

Google Gemini. There is currently no way to sign in to Nastech with a consumer Gemini subscription — the gemini provider takes an API key, and Google Vertex AI bills to your GCP project. A billing-enabled Google Cloud project is recommended for agent use; free-tier quotas are too small for long-running agent sessions. See the Google Gemini guide.

One subscription instead of five

If you'd rather not track per-provider plan semantics at all, Nastech Portal covers 300+ models under a single subscription with one OAuth login.

Anthropic (Native)​

Use Claude models directly through the Anthropic API — no OpenRouter proxy needed. Supports three auth methods:

When no explicit environment credential is selected, Nastech-owned OAuth grants in the credential pool take precedence over a borrowed Claude Code login. The borrowed login remains the fallback when no owned OAuth grant is available — unless auth.adopt_external_logins: false is set, in which case Nastech never reads or refreshes Claude Code's credentials (see Borrowed CLI logins). Auxiliary authentication recovery refreshes the credential used by the failed request, not an unrelated ambient login; rotating a borrowed login can otherwise invalidate its owner's refresh token.

Requires Claude Max "extra usage" credits

When you authenticate via nastech model → Anthropic OAuth (or via nastech auth add anthropic --type oauth), Nastech routes as Claude Code against your Anthropic account. It only works if you're on a Claude Max plan and have purchased extra usage credits. The base Max plan allowance (the usage included in Claude Code by default) is not consumed by Nastech — only the extra/overage credits you've added on top are. Claude Pro subscribers cannot use this path.

If you don't have Max + extra credits, use an ANTHROPIC_API_KEY instead — requests are billed pay-per-token against that key's organization (standard API pricing, independent of any Claude subscription).

# With an API key (pay-per-token)
export ANTHROPIC_API_KEY=***
nastech chat --provider anthropic --model claude-sonnet-4-6

# Preferred: authenticate through `nastech model`
# Nastech will use Claude Code's credential store directly when available
nastech model

# Manual override with a setup-token (fallback / legacy)
export ANTHROPIC_TOKEN=*** # setup-token or manual OAuth token
nastech chat --provider anthropic

# Auto-detect Claude Code credentials (if you already use Claude Code)
nastech chat --provider anthropic # reads Claude Code credential files automatically

When you choose Anthropic OAuth through nastech model, Nastech prefers Claude Code's own credential store over copying the token into ~/.nastech/.env. That keeps refreshable Claude credentials refreshable.

Or set it permanently:

model:
provider: "anthropic"
default: "claude-sonnet-4-6"
Aliases

--provider claude and --provider claude-code also work as shorthand for --provider anthropic.

GitHub Copilot​

Nastech supports GitHub Copilot as a first-class provider with two modes:

copilot — Direct Copilot API (recommended). Uses your GitHub Copilot subscription to access GPT-5.x, Claude, Gemini, and other models through the Copilot API.

nastech chat --provider copilot --model gpt-5.4

Authentication options (checked in this order):

  1. COPILOT_GITHUB_TOKEN environment variable
  2. GH_TOKEN environment variable
  3. GITHUB_TOKEN environment variable
  4. gh auth token CLI fallback

If no token is found, nastech model offers an OAuth device code login — the same flow used by the Copilot CLI and opencode.

Token types

The Copilot API does not support classic Personal Access Tokens (ghp_*). Supported token types:

TypePrefixHow to get
OAuth tokengho_nastech model → GitHub Copilot → Login with GitHub
Fine-grained PATgithub_pat_GitHub Settings → Developer settings → Fine-grained tokens (needs Copilot Requests permission)
GitHub App tokenghu_Via GitHub App installation

If your gh auth token returns a ghp_* token, use nastech model to authenticate via OAuth instead.

Copilot auth behavior in Nastech

Nastech sends a supported GitHub token (gho_*, github_pat_*, or ghu_*) directly to api.githubcopilot.com and includes Copilot-specific headers (Editor-Version, Copilot-Integration-Id, Openai-Intent, x-initiator).

On HTTP 401, Nastech now performs a one-shot credential recovery before fallback:

  1. Re-resolve token via the normal priority chain (COPILOT_GITHUB_TOKEN → GH_TOKEN → GITHUB_TOKEN → gh auth token)
  2. Rebuild the shared OpenAI client with refreshed headers
  3. Retry the request once

Some older community proxies use api.github.com/copilot_internal/v2/token exchange flows. That endpoint can be unavailable for some account types (returns 404). Nastech therefore keeps direct-token auth as the primary path and relies on runtime credential refresh + retry for robustness.

API routing: GPT-5+ models (except gpt-5-mini) automatically use the Responses API. All other models (GPT-4o, Claude, Gemini, etc.) use Chat Completions. Models are auto-detected from the live Copilot catalog.

copilot-acp — Copilot ACP agent backend. Spawns the local Copilot CLI as a subprocess:

nastech chat --provider copilot-acp --model copilot-acp
# Requires the GitHub Copilot CLI in PATH and an existing `copilot login` session

Permanent config:

model:
provider: "copilot"
default: "gpt-5.4"
Environment variableDescription
COPILOT_GITHUB_TOKENGitHub token for Copilot API (first priority)
NASTECH_COPILOT_ACP_COMMANDOverride the Copilot CLI binary path (default: copilot)
NASTECH_COPILOT_ACP_ARGSOverride ACP args (default: --acp --stdio)

First-Class API-Key Providers​

These providers have built-in support with dedicated provider IDs. Set the API key and use --provider to select:

# Fireworks AI
nastech chat --provider fireworks --model accounts/fireworks/models/kimi-k2p6
# Requires: FIREWORKS_API_KEY in ~/.nastech/.env

# NovitaAI Model API
nastech chat --provider novita --model moonshotai/kimi-k2.5
# Requires: NOVITA_API_KEY in ~/.nastech/.env

# Ramp Router (model IDs come from your account's live catalog)
nastech chat --provider router --model gpt-5.4-mini
# Requires: RAMP_ROUTER_API_KEY in ~/.nastech/.env

# z.ai / ZhipuAI GLM
nastech chat --provider zai --model glm-5
# Requires: GLM_API_KEY in ~/.nastech/.env

# Kimi / Moonshot AI (international: api.moonshot.ai)
nastech chat --provider kimi-coding --model kimi-for-coding
# Requires: KIMI_API_KEY in ~/.nastech/.env

# Kimi / Moonshot AI (China: api.moonshot.cn)
nastech chat --provider kimi-coding-cn --model kimi-k2.5
# Requires: KIMI_CN_API_KEY in ~/.nastech/.env

# MiniMax (global endpoint)
nastech chat --provider minimax --model MiniMax-M2.7
# Requires: MINIMAX_API_KEY in ~/.nastech/.env

# MiniMax (China endpoint)
nastech chat --provider minimax-cn --model MiniMax-M2.7
# Requires: MINIMAX_CN_API_KEY in ~/.nastech/.env

# Qwen Cloud / DashScope (Qwen models)
nastech chat --provider alibaba --model qwen3.5-plus
# Requires: DASHSCOPE_API_KEY in ~/.nastech/.env

# Xiaomi MiMo
nastech chat --provider xiaomi --model mimo-v2-pro
# Requires: XIAOMI_API_KEY in ~/.nastech/.env

# Tencent TokenHub (Hy4 preview)
nastech chat --provider tencent-tokenhub --model hy4-preview
# Requires: TOKENHUB_API_KEY in ~/.nastech/.env

# Tencent TokenPlan (Hy4 preview via Anthropic Messages endpoint)
nastech chat --provider tencent-tokenplan --model hy4-preview
# Requires: TOKENPLAN_API_KEY in ~/.nastech/.env

# Arcee AI (Trinity models)
nastech chat --provider arcee --model trinity-large-thinking
# Requires: ARCEEAI_API_KEY in ~/.nastech/.env

# Meta Model API (Muse Spark family)
nastech chat --provider meta-ai --model muse-spark-1.2
# Requires: MODEL_API_KEY in ~/.nastech/.env

# GMI Cloud
# Use the exact model ID returned by GMI's /v1/models endpoint.
nastech chat --provider gmi --model zai-org/GLM-5.1-FP8
# Requires: GMI_API_KEY in ~/.nastech/.env

# Nebius Token Factory
nastech chat --provider nebius --model deepseek-ai/DeepSeek-V4-Pro
# Requires: NEBIUS_API_KEY in ~/.nastech/.env

Fireworks uses its native slash-form catalog IDs, such as accounts/fireworks/models/kimi-k2p6. Run nastech model, choose Fireworks AI, and select from the live catalog or enter another Fireworks model ID. The default endpoint is https://api.fireworks.ai/inference/v1; configure a different endpoint through model.base_url in config.yaml, not .env.

Or set the provider permanently in config.yaml:

model:
provider: "gmi"
default: "zai-org/GLM-5.1-FP8"

Base URLs can be overridden with NOVITA_BASE_URL, GLM_BASE_URL, KIMI_BASE_URL, MINIMAX_BASE_URL, MINIMAX_CN_BASE_URL, DASHSCOPE_BASE_URL, XIAOMI_BASE_URL, GMI_BASE_URL, META_BASE_URL, or TOKENHUB_BASE_URL environment variables.

Meta contributor tier

muse-spark-1.2-contributor and muse-spark-1.3-contributor are Meta's contributor tiers — Meta may train on your prompts and completions, so interactive model selection asks for confirmation before using either. For current pricing and rate limits, see Meta Model API pricing and rate limits. Use the standard muse-spark-1.2 / muse-spark-1.3 (no training) for confidential work.

Z.AI Endpoint Auto-Detection

When using the Z.AI / GLM provider, Nastech automatically probes multiple endpoints (global, China, coding variants) to find one that accepts your API key. You don't need to set GLM_BASE_URL manually — the working endpoint is detected and cached automatically.

xAI (Grok) — Responses API + Prompt Caching​

xAI is wired through the Responses API (codex_responses transport) for automatic reasoning support on Grok 4 models — no reasoning_effort parameter needed, the server reasons by default. Set XAI_API_KEY in ~/.nastech/.env and pick xAI in nastech model, or drop grok as a shortcut into /model grok-4-fast-reasoning.

SuperGrok and X Premium+ subscribers can sign in with browser OAuth instead of using an API key — pick xAI Grok OAuth (SuperGrok / Premium+) in nastech model, or run nastech auth add xai-oauth. The same OAuth bearer token is automatically reused by direct-to-xAI tools (TTS, image gen, video gen, transcription). See the xAI Grok OAuth guide for the full flow — and if Nastech runs on a remote host, also see OAuth over SSH / Remote Hosts for the required ssh -L tunnel.

When using xAI as a provider (any base URL containing x.ai), Nastech automatically enables prompt caching by sending the x-grok-conv-id header with every API request. This routes requests to the same server within a conversation session, allowing xAI's infrastructure to reuse cached system prompts and conversation history.

No configuration is needed — caching activates automatically when an xAI endpoint is detected and a session ID is available. This reduces latency and cost for multi-turn conversations.

xAI also ships a dedicated TTS endpoint (/v1/tts). Select xAI TTS in nastech tools → Voice & TTS, or see the Voice & TTS page for config.

Retired xAI model migration (May 15, 2026): xAI is retiring grok-4*, grok-3, grok-code-fast-1, and grok-imagine-image-pro on 2026-05-15. nastech doctor and nastech chat startup both detect any config still pointing at a retired ref and print the recommended replacement. Use nastech migrate xai for a one-shot config rewrite — dry-run by default, add --apply to write changes (a timestamped copy of the previous config lands in backups/config/ first).

nastech migrate xai # preview replacements
nastech migrate xai --apply # rewrite ~/.nastech/config.yaml in place

xAI Web Search backend. When the Web Search toolset is enabled, web.backend: xai routes search through xAI's hosted search endpoint using the same XAI_API_KEY / OAuth credentials. No additional setup required if xAI is already configured as a provider.

NovitaAI​

NovitaAI is the AI-native cloud for builders and agents. Its three product lines are Model API for 200+ models, Agent Sandbox for building and running AI agents, and GPU Cloud for scalable compute, all available from one platform.

# Use any available model
nastech chat --provider novita --model moonshotai/kimi-k2.5
# Requires: NOVITA_API_KEY in ~/.nastech/.env

# Short alias
nastech chat --provider novita-ai --model deepseek/deepseek-v3-0324

Or set it permanently in config.yaml:

model:
provider: "novita"
default: "moonshotai/kimi-k2.5"
base_url: "https://api.novita.ai/openai/v1"

Get your API key at novita.ai/settings/key-management. The base URL can be overridden with NOVITA_BASE_URL.

Ollama Cloud — Managed Ollama Models, OAuth + API Key​

Ollama Cloud hosts the same open-weight catalog as local Ollama but without the GPU requirement. Pick it in nastech model as Ollama Cloud, paste your API key from ollama.com/settings/keys, and Nastech auto-discovers the available models.

nastech model
# → pick "Ollama Cloud"
# → paste your OLLAMA_API_KEY
# → select from discovered models (gpt-oss:120b, glm-4.6:cloud, qwen3-coder:480b-cloud, etc.)

Or config.yaml directly:

model:
provider: "ollama-cloud"
default: "gpt-oss:120b"

The model catalog is fetched dynamically from ollama.com/v1/models and cached for one hour. model:tag notation (e.g. qwen3-coder:480b-cloud) is preserved through normalization — don't use dashes.

Ollama Cloud vs local Ollama

Both speak the same OpenAI-compatible API. Cloud is a first-class provider (--provider ollama-cloud, OLLAMA_API_KEY); local Ollama is reached via the Custom Endpoint flow (base URL http://localhost:11434/v1, no key). Use cloud for large models you can't run locally; use local for privacy or offline work.

DeepInfra​

DeepInfra (--provider deepinfra, DEEPINFRA_API_KEY) is discovered live from its catalog. Reasoning is controlled through DeepInfra's top-level reasoning_effort field, so agent.reasoning_effort, /reasoning <level>, --reasoning and per-model agent.reasoning_overrides work in both directions: an effort turns thinking on for models that default off (DeepSeek-V4.x), /reasoning none turns it off for models that default on (GLM-4.6, Qwen3-Thinking). Leaving reasoning unset keeps DeepInfra's per-model default; xhigh is native and ultra is sent as max.

AWS Bedrock​

Anthropic Claude, Amazon Nova, DeepSeek v3.2, Meta Llama 4, and other models via AWS Bedrock. Uses the AWS SDK (boto3) credential chain — no API key, just standard AWS auth.

# Simplest — named profile in ~/.aws/credentials
nastech chat --provider bedrock --model us.anthropic.claude-sonnet-4-6

# Or with explicit env vars
AWS_PROFILE=myprofile AWS_REGION=us-east-1 nastech chat --provider bedrock --model us.anthropic.claude-sonnet-4-6

Or permanently in config.yaml:

model:
provider: "bedrock"
default: "us.anthropic.claude-sonnet-4-6"
bedrock:
region: "us-east-1" # or set AWS_REGION
# profile: "myprofile" # or set AWS_PROFILE
# discovery: true # auto-discover region from IAM
# guardrail: # optional Bedrock Guardrails
# guardrail_identifier: "your-guardrail-id"
# guardrail_version: "DRAFT"

Authentication uses the standard boto3 chain: explicit AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY, AWS_PROFILE from ~/.aws/credentials, IAM role on EC2/ECS/Lambda, IMDS, or SSO. No env var is required if you're already authenticated with the AWS CLI.

Bedrock uses the Converse API under the hood — requests are translated to Bedrock's model-agnostic shape, so the same config works for Claude, Nova, DeepSeek, and Llama models. Set BEDROCK_BASE_URL only if you're calling a non-default regional endpoint.

See the AWS Bedrock guide for a walkthrough of IAM setup, region selection, and cross-region inference.

Google Vertex AI​

Gemini models on Google Cloud Vertex AI via Vertex's OpenAI-compatible endpoint. Authentication is OAuth2 — a short-lived access token (~1 hour) minted from a service-account JSON or Application Default Credentials (ADC). There is no static API key; Nastech mints and auto-refreshes the token for you, including re-minting on a mid-session 401.

# Service account JSON (recommended for servers / gateways)
echo "VERTEX_CREDENTIALS_PATH=/path/to/service-account.json" >> ~/.nastech/.env
# or Application Default Credentials
gcloud auth application-default login

nastech model # → "Google Vertex AI" → project → region → model

Or in config.yaml (project/region are non-secret and live here; the credential path stays in .env):

model:
provider: "vertex"
default: "google/gemini-3-flash-preview" # Vertex requires the google/ prefix
vertex:
project_id: "my-gcp-project" # blank → use the project embedded in the credentials
region: "global" # required for the Gemini 3.x previews

VERTEX_PROJECT_ID / VERTEX_REGION env vars override the config.yaml values. Nastech lazy-installs google-auth on first use; run nastech setup if the managed install needs repair. See the Google Vertex AI guide for the full walkthrough, and the Google Gemini guide for the static-API-key AI Studio path instead.

Qwen Portal (OAuth)​

Alibaba's Qwen Portal with browser-based OAuth login. Pick Qwen OAuth (Portal) in nastech model, sign in through the browser, and Nastech persists the refresh token.

nastech model
# → pick "Qwen OAuth (Portal)"
# → browser opens; sign in with your Alibaba account
# → confirm — credentials are saved to ~/.nastech/auth.json

nastech chat # uses portal.qwen.ai/v1 endpoint

Or configure config.yaml:

model:
provider: "qwen-oauth"
default: "qwen3-coder-plus"

Set NASTECH_QWEN_BASE_URL only if the portal endpoint relocates (default: https://portal.qwen.ai/v1).

Qwen OAuth vs Qwen Cloud (Alibaba DashScope)

qwen-oauth uses the consumer-facing Qwen Portal with OAuth login — ideal for individual users. The alibaba provider uses Qwen Cloud (Alibaba DashScope) with a DASHSCOPE_API_KEY — ideal for programmatic / production workloads. Both route to Qwen-family models but live at different endpoints.

Alibaba Cloud (Coding Plan)​

If you're subscribed to Alibaba's Coding Plan (a pricing SKU separate from standard DashScope API access), Nastech exposes it as its own first-class provider: alibaba-coding-plan. Endpoint: https://coding-intl.dashscope.aliyuncs.com/v1. It's OpenAI-compatible like the regular alibaba provider but with a different base URL and billing surface.

model:
provider: alibaba_coding # alias for alibaba-coding-plan
model: qwen3-coder-plus

Or from the CLI:

nastech chat --provider alibaba_coding --model qwen3-coder-plus

alibaba_coding uses the same DASHSCOPE_API_KEY your alibaba entry already uses — no separate key needed, just a different routing target. Before this provider was registered, users who set provider: alibaba_coding in config.yaml silently fell through to OpenRouter routing.

For the mainland-China endpoint (alibaba-coding-plan-cn, https://coding.dashscope.aliyuncs.com/v1) set ALIBABA_CODING_PLAN_CN_API_KEY. The CN provider still falls back to ALIBABA_CODING_PLAN_API_KEY / DASHSCOPE_API_KEY, but with only the shared key set the /model picker lists just the international row — set the CN key (or provider: alibaba-coding-plan-cn in config.yaml) to surface the CN one. The same applies to alibaba-token-plan-cn with ALIBABA_TOKEN_PLAN_CN_API_KEY.

MiniMax (OAuth)​

MiniMax-M2.7 via browser OAuth login — no API key needed. Pick MiniMax (OAuth) in nastech model, sign in through the browser, and Nastech persists the access + refresh tokens. Uses the Anthropic Messages-compatible endpoint (/anthropic) under the hood.

nastech model
# → pick "MiniMax (OAuth)"
# → browser opens; sign in with your MiniMax account (global or CN region)
# → confirm — credentials are saved to ~/.nastech/auth.json

nastech chat # uses api.minimax.io/anthropic endpoint

Or configure config.yaml:

model:
provider: "minimax-oauth"
default: "MiniMax-M2.7"

Supported models: MiniMax-M2.7 (main) and MiniMax-M2.7-highspeed (wired as the default auxiliary model). The OAuth path ignores MINIMAX_API_KEY / MINIMAX_BASE_URL.

MiniMax OAuth vs API key

minimax-oauth uses MiniMax's consumer-facing portal with OAuth login — no billing setup required. The minimax and minimax-cn providers use MINIMAX_API_KEY / MINIMAX_CN_API_KEY — for programmatic access. See the MiniMax OAuth guide for a full walkthrough.

NVIDIA NIM​

Nemotron and other open source models via build.nvidia.com (free API key) or a local NIM endpoint.

# Cloud (build.nvidia.com)
nastech chat --provider nvidia --model nvidia/nemotron-3-super-120b-a12b
# Requires: NVIDIA_API_KEY in ~/.nastech/.env

# Local NIM endpoint — override base URL
NVIDIA_BASE_URL=http://localhost:8000/v1 nastech chat --provider nvidia --model nvidia/nemotron-3-super-120b-a12b

Or set it permanently in config.yaml:

model:
provider: "nvidia"
default: "nvidia/nemotron-3-super-120b-a12b"
Local NIM

For on-prem deployments (DGX Spark, local GPU), set NVIDIA_BASE_URL=http://localhost:8000/v1. NIM exposes the same OpenAI-compatible chat completions API as build.nvidia.com, so switching between cloud and local is a one-line env-var change.

Nastech automatically attaches the NIM billing-origin header on every request to build.nvidia.com — no configuration needed. This routes consumption against the correct origin in NVIDIA's billing dashboard.

GMI Cloud​

Open and reasoning models via GMI Cloud — OpenAI-compatible API, API key authentication.

# GMI Cloud
nastech chat --provider gmi --model deepseek-ai/DeepSeek-V3.2
# Requires: GMI_API_KEY in ~/.nastech/.env

Or set it permanently in config.yaml:

model:
provider: "gmi"
default: "deepseek-ai/DeepSeek-V3.2"

The base URL can be overridden with GMI_BASE_URL (default: https://api.gmi-serving.com/v1).

Actual Computer​

Your own hardware as a private inference cluster via Actual Computer. Two serving modes, both using Chat Completions so reasoning and final content are returned together:

  • Hosted relay — https://api.actual.inc, end-to-end encrypted, routes to your cluster. Authenticate with an ac_ inference key from actual.inc/user/keys.
  • Local daemon — on-device at http://127.0.0.1:8080, fully offline. No API key needed: Nastech detects the loopback base URL and authenticates with an internal placeholder automatically.
# Hosted relay (ACTUAL_API_KEY in ~/.nastech/.env)
nastech chat --provider actual --model <model-id-from-your-cluster>

# Local daemon (model.base_url in ~/.nastech/config.yaml, no key)
nastech chat --provider actual --model <installed-model-name>

Store provider settings in ~/.nastech/config.yaml; only the hosted API key belongs in .env:

model:
provider: "actual"
default: "<model-id>"
base_url: "http://127.0.0.1:8080/v1" # Omit for the hosted relay.

Notes:

  • Model IDs come from your cluster's GET /v1/models — discover with nastech model or curl -s https://api.actual.inc/v1/models -H "Authorization: Bearer $ACTUAL_API_KEY".
  • Bare hosts in model.base_url are normalized: http://127.0.0.1:8080 becomes http://127.0.0.1:8080/v1 automatically. The legacy ACTUAL_BASE_URL environment variable is a fallback when no Actual URL is configured in YAML.
  • Actual uses /v1/chat/completions for chat, compaction, title generation, and every other auxiliary task. This also applies to custom providers targeting api.actual.inc, model switches, and fallbacks. Legacy Responses settings in the main model, custom provider, or auxiliary task configuration are overridden automatically.
  • Reasoning effort is clamped to Actual's supported range (none/low/medium/high/max) — a global xhigh/ultra setting will not 400 requests.
  • Small local models: Nastech' full default toolset plus the system prompt can exceed a 32k context window, producing an empty-stream error from llama.cpp-family servers. Restrict the toolset (-t file,web) or load the model with a larger context. The optional actual-setup skill (nastech skills install official/devops/actual-setup) covers setup and troubleshooting in detail.
  • Aliases: actual-computer, actualcomputer, aci.

StepFun​

Step-series models via StepFun — OpenAI-compatible API, API key authentication.

# StepFun
nastech chat --provider stepfun --model step-3.5-flash
# Requires: STEPFUN_API_KEY in ~/.nastech/.env

Or set it permanently in config.yaml:

model:
provider: "stepfun"
default: "step-3.5-flash"

The base URL can be overridden with STEPFUN_BASE_URL (default: https://api.stepfun.com/v1).

Hugging Face Inference Providers​

Hugging Face Inference Providers routes to 20+ open models through a unified OpenAI-compatible endpoint (router.huggingface.co/v1). Requests are automatically routed to the fastest available backend (Groq, Together, SambaNova, etc.) with automatic failover.

# Use any available model
nastech chat --provider huggingface --model Qwen/Qwen3.5-397B-A17B
# Requires: HF_TOKEN in ~/.nastech/.env

# Short alias
nastech chat --provider hf --model deepseek-ai/DeepSeek-V3.2

Or set it permanently in config.yaml:

model:
provider: "huggingface"
default: "Qwen/Qwen3.5-397B-A17B"

Get your token at huggingface.co/settings/tokens — make sure to enable the "Make calls to Inference Providers" permission. Free tier included ($0.10/month credit, no markup on provider rates).

You can append routing suffixes to model names: :fastest (default), :cheapest, or :provider_name to force a specific backend.

The base URL can be overridden with HF_BASE_URL.

Custom & Self-Hosted LLM Providers​

Nastech Agent works with any OpenAI-compatible API endpoint. If a server implements /v1/chat/completions, you can point Nastech at it. This means you can use local models, GPU inference servers, multi-provider routers, or any third-party API.

General Setup​

Three ways to configure a custom endpoint:

Interactive setup (recommended):

nastech model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter: API base URL, API key, Model name

Manual config (config.yaml):

# In ~/.nastech/config.yaml
model:
default: your-model-name
provider: custom
base_url: http://localhost:8000/v1
api_key: your-key-or-leave-empty-for-local
Legacy env vars

LLM_MODEL in .env is removed — config.yaml is the single source of truth for model and endpoint configuration. OPENAI_BASE_URL is still honored, but only for the openai-api provider (it overrides the OpenAI endpoint for direct API-key access). For other providers and custom endpoints, use nastech model or set model.base_url in config.yaml directly. If you have stale entries in your .env, they are automatically cleared on the next nastech setup or config migration.

Both approaches persist to config.yaml, which is the source of truth for model, provider, and base URL.

Switching Models with /model​

nastech model vs /model

nastech model (run from your terminal, outside any chat session) is the full provider setup wizard. Use it to add new providers, run OAuth flows, enter API keys, and configure custom endpoints.

/model (typed inside an active Nastech chat session) can only switch between providers and models you've already set up. It cannot add new providers, run OAuth, or prompt for API keys. If you've only configured one provider (e.g. OpenRouter), /model will only show models for that provider.

To add a new provider: Exit your session (Ctrl+C or /quit), run nastech model, set up the new provider, then start a new session.

Once you have at least one custom endpoint configured, you can switch models mid-session:

/model custom:qwen-2.5 # Switch to a model on your custom endpoint
/model custom # Auto-detect the model from the endpoint
/model openrouter:claude-sonnet-4 # Switch back to a cloud provider

If you have named custom providers configured (see below), use the triple syntax:

/model custom:local:qwen-2.5 # Use the "local" custom provider with model qwen-2.5
/model custom:work:llama3 # Use the "work" custom provider with llama3

When switching providers, Nastech persists the base URL and provider to config so the change survives restarts. When switching away from a custom endpoint to a built-in provider, the stale base URL is automatically cleared.

tip

/model custom (bare, no model name) queries your endpoint's /models API and auto-selects the model if exactly one is loaded. Useful for local servers running a single model.

Everything below follows this same pattern — just change the URL, key, and model name.


Ollama — Local Models, Zero Config​

Ollama runs open-weight models locally with one command. Best for: quick local experimentation, privacy-sensitive work, offline use. Supports tool calling via the OpenAI-compatible API.

# Install and run a model
ollama pull qwen2.5-coder:32b
ollama serve # Starts on port 11434

Then configure Nastech:

nastech model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:11434/v1
# Skip API key (Ollama doesn't need one)
# Enter model name (e.g. qwen2.5-coder:32b)

Or configure config.yaml directly:

model:
default: qwen2.5-coder:32b
provider: custom
base_url: http://localhost:11434/v1
context_length: 64000 # See warning below
Ollama defaults to very low context lengths

Ollama does not use your model's full context window by default. Depending on your VRAM, the default is:

Available VRAMDefault context
Less than 24 GB4,096 tokens
24–48 GB32,768 tokens
48+ GB256,000 tokens

Nastech Agent requires at least 64,000 tokens of context for agent use with tools. Smaller windows are rejected at startup because the system prompt, tool schemas, and working conversation state need enough room for reliable multi-step workflows.

How to increase it (pick one):

# Option 1: Set server-wide via environment variable (recommended)
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

# Option 2: For systemd-managed Ollama
sudo systemctl edit ollama.service
# Add: Environment="OLLAMA_CONTEXT_LENGTH=64000"
# Then: sudo systemctl daemon-reload && sudo systemctl restart ollama

# Option 3: Bake it into a custom model (persistent per-model)
echo -e "FROM qwen2.5-coder:32b\nPARAMETER num_ctx 64000" > Modelfile
ollama create qwen2.5-coder-64k -f Modelfile

You cannot set context length through the OpenAI-compatible API (/v1/chat/completions). It must be configured server-side or via a Modelfile. This is the #1 source of confusion when integrating Ollama with tools like Nastech.

Verify your context is set correctly:

ollama ps
# Look at the CONTEXT column — it should show your configured value
tip

List available models with ollama list. Pull any model from the Ollama library with ollama pull <model>. Ollama handles GPU offloading automatically — no configuration needed for most setups.


vLLM — High-Performance GPU Inference​

vLLM is the standard for production LLM serving. Best for: maximum throughput on GPU hardware, serving large models, continuous batching.

pip install vllm
vllm serve meta-llama/Llama-3.1-70B-Instruct \
--port 8000 \
--max-model-len 65536 \
--tensor-parallel-size 2 \
--enable-auto-tool-choice \
--tool-call-parser nastech

Then configure Nastech:

nastech model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:8000/v1
# Skip API key (or enter one if you configured vLLM with --api-key)
# Enter model name: meta-llama/Llama-3.1-70B-Instruct

Context length: vLLM reads the model's max_position_embeddings by default. If that exceeds your GPU memory, it errors and asks you to set --max-model-len lower. You can also use --max-model-len auto to automatically find the maximum that fits. Set --gpu-memory-utilization 0.95 (default 0.9) to squeeze more context into VRAM.

Tool calling requires explicit flags:

FlagPurpose
--enable-auto-tool-choiceRequired for tool_choice: "auto" (the default in Nastech)
--tool-call-parser <name>Parser for the model's tool call format

Supported parsers: nastech (Qwen 2.5, Nastech 2/3), llama3_json (Llama 3.x), mistral, deepseek_v3, deepseek_v31, xlam, pythonic. Without these flags, tool calls won't work — the model will output tool calls as text.

Qwen reasoning parsers: Nastech preserves structured reasoning metadata such as reasoning, reasoning_content, and streamed reasoning deltas when OpenAI-compatible servers return them. That metadata is treated as reasoning/thinking trace data, not as a replacement for the assistant's visible answer. For Qwen reasoning models served by vLLM, make sure the final user-visible response still appears in content. If --reasoning-parser qwen3 leaves content empty in your deployment, either disable that parser or pass a server-supported request option such as chat_template_kwargs.enable_thinking: false through extra_body.

tip

vLLM supports human-readable sizes: --max-model-len 64k (lowercase k = 1000, uppercase K = 1024).


SGLang — Fast Serving with RadixAttention​

SGLang is an alternative to vLLM with RadixAttention for KV cache reuse. Best for: multi-turn conversations (prefix caching), constrained decoding, structured output.

pip install "sglang[all]"
python -m sglang.launch_server \
--model meta-llama/Llama-3.1-70B-Instruct \
--port 30000 \
--context-length 65536 \
--tp 2 \
--tool-call-parser qwen

Then configure Nastech:

nastech model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:30000/v1
# Enter model name: meta-llama/Llama-3.1-70B-Instruct

Context length: SGLang reads from the model's config by default. Use --context-length to override. If you need to exceed the model's declared maximum, set SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1.

Tool calling: Use --tool-call-parser with the appropriate parser for your model family: qwen (Qwen 2.5), llama3, llama4, deepseekv3, mistral, glm. Without this flag, tool calls come back as plain text.

SGLang defaults to 128 max output tokens

If responses seem truncated, check the server's generation default and configure it on the server (for example SGLang's --default-max-tokens). Nastech does not expose an output-token cap setting.


llama.cpp / llama-server — CPU & Metal Inference​

llama.cpp runs quantized models on CPU, Apple Silicon (Metal), and consumer GPUs. Best for: running models without a datacenter GPU, Mac users, edge deployment.

# Build and start llama-server
cmake -B build && cmake --build build --config Release
./build/bin/llama-server \
--jinja -fa \
-c 64000 \
-ngl 99 \
-m models/qwen2.5-coder-32b-instruct-Q4_K_M.gguf \
--port 8080 --host 0.0.0.0

Context length (-c): Recent builds default to 0 which reads the model's training context from the GGUF metadata. For models with 128k+ training context, this can OOM trying to allocate the full KV cache. Set -c explicitly to at least 64,000 tokens for Nastech. If using parallel slots (-np), the total context is divided among slots — with -c 64000 -np 4, each slot only gets 16k, which is below Nastech' minimum per active session.

Then configure Nastech to point at it:

nastech model
# Select "Custom endpoint (self-hosted / VLLM / etc.)"
# Enter URL: http://localhost:8080/v1
# Skip API key (local servers don't need one)
# Enter model name — or leave blank to auto-detect if only one model is loaded

This saves the endpoint to config.yaml so it persists across sessions.

--jinja is required for tool calling

Without --jinja, llama-server ignores the tools parameter entirely. The model will try to call tools by writing JSON in its response text, but Nastech won't recognize it as a tool call — you'll see raw JSON like {"name": "web_search", ...} printed as a message instead of an actual search.

Native tool calling support (best performance): Llama 3.x, Qwen 2.5 (including Coder), Nastech 2/3, Mistral, DeepSeek, Functionary. All other models use a generic handler that works but may be less efficient. See the llama.cpp function calling docs for the full list.

You can verify tool support is active by checking http://localhost:8080/props — the chat_template field should be present.

tip

Download GGUF models from Hugging Face. Q4_K_M quantization offers the best balance of quality vs. memory usage.


LM Studio — Desktop App with Local Models​

LM Studio is a desktop app for running local models with a GUI. Best for: users who prefer a visual interface, quick model testing, developers on macOS/Windows/Linux.

Start the server from the LM Studio app (Developer tab → Start Server), or use the CLI:

lms server start # Starts on port 1234
lms load qwen2.5-coder --context-length 64000

Then configure Nastech:

nastech model
# Select "LM Studio"
# Press Enter to use http://localhost:1234/v1
# Pick one of the discovered models
# If LM Studio server auth is enabled, enter LM_API_KEY when prompted

Nastech preserves the context of an already-loaded LM Studio instance. For an unloaded model in the default explicit mode, Nastech omits context_length unless you configured one in Nastech, so LM Studio can apply its own model setting. Nastech then uses only the context length LM Studio reports after loading.

To change context length in LM Studio:

  1. Click the gear icon next to the model picker
  2. Set "Context Length" to at least 64000 for a smooth experience
  3. Reload the model for the change to take effect
  4. If your machine cannot fit 64000, consider using a smaller model with larger context lengths.

Alternatively, use the CLI: lms load model-name --context-length 64000

You can use the CLI to estimate if the model will fit: lms load model-name --context-length 64000 --estimate-only

To set persistent per-model defaults: My Models tab → gear icon on the model → set context size. :::

If you use LM Studio's Just-In-Time loading / Auto-Evict feature and want LM Studio to manage model loading and eviction from normal chat requests, skip Nastech' explicit preload step:

nastech config set model.lmstudio_load_mode jit

Set it back to the default explicit preload behavior with:

nastech config set model.lmstudio_load_mode explicit

Tool calling: Supported since LM Studio 0.3.6. Models with native tool-calling training (Qwen 2.5, Llama 3.x, Mistral, Nastech) are auto-detected and shown with a tool badge. Other models use a generic fallback that may be less reliable.


WSL2 Networking (Windows Users)​

Since Nastech Agent requires a Unix environment, Windows users run it inside WSL2. If your model server (Ollama, LM Studio, etc.) runs on the Windows host, you need to bridge the network gap — WSL2 uses a virtual network adapter with its own subnet, so localhost inside WSL2 refers to the Linux VM, not the Windows host.

Both in WSL2? No problem.

If your model server also runs inside WSL2 (common for vLLM, SGLang, and llama-server), localhost works as expected — they share the same network namespace. Skip this section.

Available on Windows 11 22H2+, mirrored mode makes localhost work bidirectionally between Windows and WSL2 — the simplest fix.

  1. Create or edit %USERPROFILE%\.wslconfig (e.g., C:\Users\YourName\.wslconfig):

    [wsl2]
    networkingMode=mirrored
  2. Restart WSL from PowerShell:

    wsl --shutdown
  3. Reopen your WSL2 terminal. localhost now reaches Windows services:

    curl http://localhost:11434/v1/models # Ollama on Windows — works
Hyper-V Firewall

On some Windows 11 builds, the Hyper-V firewall blocks mirrored connections by default. If localhost still doesn't work after enabling mirrored mode, run this in an Admin PowerShell:

Set-NetFirewallHyperVVMSetting -Name '{40E0AC32-46A5-438A-A0B2-2B479E8F2E90}' -DefaultInboundAction Allow

Option 2: Use the Windows Host IP (Windows 10 / older builds)​

If you can't use mirrored mode, find the Windows host IP from inside WSL2 and use that instead of localhost:

# Get the Windows host IP (the default gateway of WSL2's virtual network)
ip route show | grep -i default | awk '{ print $3 }'
# Example output: 172.29.192.1

Use that IP in your Nastech config:

model:
default: qwen2.5-coder:32b
provider: custom
base_url: http://172.29.192.1:11434/v1 # Windows host IP, not localhost
Dynamic helper

The host IP can change on WSL2 restart. You can grab it dynamically in your shell:

export WSL_HOST=$(ip route show | grep -i default | awk '{ print $3 }')
echo "Windows host at: $WSL_HOST"
curl http://$WSL_HOST:11434/v1/models # Test Ollama

Or use your machine's mDNS name (requires libnss-mdns in WSL2):

sudo apt install libnss-mdns
curl http://$(hostname).local:11434/v1/models

Server Bind Address (Required for NAT Mode)​

If you're using Option 2 (NAT mode with the host IP), the model server on Windows must accept connections from outside 127.0.0.1. By default, most servers only listen on localhost — WSL2 connections in NAT mode come from a different virtual subnet and will be refused. In mirrored mode, localhost maps directly so the default 127.0.0.1 binding works fine.

ServerDefault bindHow to fix
Ollama127.0.0.1Set OLLAMA_HOST=0.0.0.0 environment variable before starting Ollama (System Settings → Environment Variables on Windows, or edit the Ollama service)
LM Studio127.0.0.1Enable "Serve on Network" in the Developer tab → Server settings
llama-server127.0.0.1Add --host 0.0.0.0 to the startup command
vLLM0.0.0.0Already binds to all interfaces by default
SGLang127.0.0.1Add --host 0.0.0.0 to the startup command

Ollama on Windows (detailed): Ollama runs as a Windows service. To set OLLAMA_HOST:

  1. Open System Properties → Environment Variables
  2. Add a new System variable: OLLAMA_HOST = 0.0.0.0
  3. Restart the Ollama service (or reboot)

Windows Firewall​

Windows Firewall treats WSL2 as a separate network (in both NAT and mirrored mode). If connections still fail after the steps above, add a firewall rule for your model server's port:

# Run in Admin PowerShell — replace PORT with your server's port
New-NetFirewallRule -DisplayName "Allow WSL2 to Model Server" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 11434

Common ports: Ollama 11434, vLLM 8000, SGLang 30000, llama-server 8080, LM Studio 1234.

Quick Verification​

From inside WSL2, test that you can reach your model server:

# Replace URL with your server's address and port
curl http://localhost:11434/v1/models # Mirrored mode
curl http://172.29.192.1:11434/v1/models # NAT mode (use your actual host IP)

If you get a JSON response listing your models, you're good. Use that same URL as the base_url in your Nastech config.


Troubleshooting Local Models​

These issues affect all local inference servers when used with Nastech.

"Connection refused" from WSL2 to a Windows-hosted model server​

If you're running Nastech inside WSL2 and your model server on the Windows host, http://localhost:<port> won't work in WSL2's default NAT networking mode. See WSL2 Networking above for the fix.

Tool calls appear as text instead of executing​

The model outputs something like {"name": "web_search", "arguments": {...}} as a message instead of actually calling the tool.

Cause: Your server doesn't have tool calling enabled, or the model doesn't support it through the server's tool calling implementation.

ServerFix
llama.cppAdd --jinja to the startup command
vLLMAdd --enable-auto-tool-choice --tool-call-parser nastech
SGLangAdd --tool-call-parser qwen (or appropriate parser)
OllamaTool calling is enabled by default — make sure your model supports it (check with ollama show model-name)
LM StudioUpdate to 0.3.6+ and use a model with native tool support

Model seems to forget context or give incoherent responses​

Cause: Context window is too small. When the conversation exceeds the context limit, most servers silently drop older messages. Nastech's system prompt + tool schemas alone can use 4k–8k tokens.

Diagnosis:

# Check what Nastech thinks the context is
# Look at startup line: "Context limit: X tokens"

# Check your server's actual context
# Ollama: ollama ps (CONTEXT column)
# llama.cpp: curl http://localhost:8080/props | jq '.default_generation_settings.n_ctx'
# vLLM: check --max-model-len in startup args

Fix: Set context to at least 64,000 tokens for agent use. See each server's section above for the specific flag.

"Context limit: 2048 tokens" at startup​

Nastech auto-detects context length from your server's /v1/models endpoint. If the server reports a low value (or doesn't report one at all), Nastech uses the model's declared limit which may be wrong.

Fix: Set it explicitly in config.yaml:

model:
default: your-model
provider: custom
base_url: http://localhost:11434/v1
context_length: 64000

Responses get cut off mid-sentence​

Possible causes:

  1. Low output limit on the server — configure the server's generation default (for example SGLang's --default-max-tokens). Nastech does not expose an output-token cap setting. Response length is distinct from the conversation's context window (context_length).
  2. Context exhaustion — The model filled its context window. Increase model.context_length or enable context compression in Nastech.

LiteLLM Proxy — Multi-Provider Gateway​

LiteLLM is an OpenAI-compatible proxy that unifies 100+ LLM providers behind a single API. Best for: switching between providers without config changes, load balancing, fallback chains, budget controls.

# Install and start
pip install "litellm[proxy]"
litellm --model anthropic/claude-sonnet-4 --port 4000

# Or with a config file for multiple models:
litellm --config litellm_config.yaml --port 4000

Then configure Nastech with nastech model → Custom endpoint → http://localhost:4000/v1.

Example litellm_config.yaml with fallback:

model_list:
- model_name: "best"
litellm_params:
model: anthropic/claude-sonnet-4
api_key: sk-ant-...
- model_name: "best"
litellm_params:
model: openai/gpt-4o
api_key: sk-...
router_settings:
routing_strategy: "latency-based-routing"

ClawRouter — Cost-Optimized Routing​

ClawRouter by BlockRunAI is a local routing proxy that auto-selects models based on query complexity. It classifies requests across 14 dimensions and routes to the cheapest model that can handle the task. Payment is via USDC cryptocurrency (no API keys).

# Install and start
npx @blockrun/clawrouter # Starts on port 8402

Then configure Nastech with nastech model → Custom endpoint → http://localhost:8402/v1 → model name blockrun/auto.

Routing profiles:

ProfileStrategySavings
blockrun/autoBalanced quality/cost74-100%
blockrun/ecoCheapest possible95-100%
blockrun/premiumBest quality models0%
blockrun/freeFree models only100%
blockrun/agenticOptimized for tool usevaries
note

ClawRouter requires a USDC-funded wallet on Base or Solana for payment. All requests route through BlockRun's backend API. Run npx @blockrun/clawrouter doctor to check wallet status.


Other Compatible Providers​

Any service with an OpenAI-compatible API works. Some popular options:

ProviderBase URLNotes
Together AIhttps://api.together.xyz/v1Cloud-hosted open models
Groqhttps://api.groq.com/openai/v1Ultra-fast inference
DeepSeekhttps://api.deepseek.com/v1DeepSeek models
Fireworks AIhttps://api.fireworks.ai/inference/v1Fast open model hosting
GMI Cloudhttps://api.gmi-serving.com/v1Managed OpenAI-compatible inference
Actual Computerhttps://api.actual.inc/v1Private relay to your own cluster; local daemon at http://127.0.0.1:8080/v1
Cerebrashttps://api.cerebras.ai/v1Wafer-scale chip inference
Mistral AIhttps://api.mistral.ai/v1Mistral models
OpenAIhttps://api.openai.com/v1Direct OpenAI access
Azure OpenAIhttps://YOUR.openai.azure.com/Enterprise OpenAI
LocalAIhttp://localhost:8080/v1Self-hosted, multi-model
Janhttp://localhost:1337/v1Desktop app with local models

Configure any of these with nastech model → Custom endpoint, or in config.yaml:

model:
default: meta-llama/Llama-3.1-70B-Instruct-Turbo
provider: custom
base_url: https://api.together.xyz/v1
api_key: your-together-key

Context Length Detection​

Context windows and output limits are different

context_length is the total context window — the combined budget for input and output tokens (e.g. 200,000 for Claude Opus 4.6). Nastech uses this to decide when to compress history and to validate API requests.

Output limits govern a single generated response, not the conversation history. Nastech no longer reads model.max_tokens, NASTECH_MAX_TOKENS, provider output-cap settings, or model_overrides.*.*.max_output_tokens. Remove these legacy settings. Custom OpenAI-compatible endpoints receive no automatic catalog-sized output cap. Their server defaults apply; these can be lower than the model maximum.

Native Anthropic Messages (including the native Anthropic Bedrock path) requires max_tokens, so Nastech supplies an internal value. Bedrock Converse is a separate protocol: its optional inferenceConfig.maxTokens is omitted by default, which AWS documents as the model maximum. Internal bounded tasks and provider-specific protocol requirements remain implementation details. Omission does not universally select a model's maximum output.

Set context_length when auto-detection gets the window size wrong.

Nastech uses a multi-source resolution chain to detect the correct context window for your model and provider:

  1. Config override — model.context_length in config.yaml (highest priority)
  2. Custom provider per-model — providers.<name>.models.<id>.context_length
  3. Persistent cache — previously discovered values (survives restarts)
  4. Endpoint /models — queries your server's API (local/custom endpoints)
  5. Anthropic /v1/models — queries Anthropic's API for max_input_tokens (API-key users only)
  6. OpenRouter API — live model metadata from OpenRouter
  7. Nastech Portal — suffix-matches Nastech model IDs against OpenRouter metadata
  8. models.dev — community-maintained registry with provider-specific context lengths for 3800+ models across 100+ providers
  9. Fallback defaults — broad model family patterns (128K default)

For most setups this works out of the box. The system is provider-aware — the same model can have different context limits depending on who serves it (e.g., claude-opus-4.6 is 1M on Anthropic direct but 128K on GitHub Copilot).

To set the context length explicitly, add context_length to your model config:

model:
default: "qwen3.5:9b"
base_url: "http://localhost:8080/v1"
context_length: 131072 # tokens

For custom endpoints, you can also set context length per model:

providers:
my-local-llm:
api: "http://localhost:11434/v1"
models:
qwen3.5:27b:
context_length: 64000
deepseek-r1:70b:
context_length: 65536

nastech model will prompt for context length when configuring a custom endpoint. Leave it blank for auto-detection.

When to set this manually
  • You're using Ollama with a custom num_ctx that's lower than the model's maximum
  • You want to limit context below the model's maximum (e.g., 8k on a 128k model to save VRAM)
  • You're running behind a proxy that doesn't expose /v1/models

Named Custom Providers​

If you work with multiple custom endpoints (e.g., a local dev server and a remote GPU server), you can define them as named custom providers under the providers: dict in config.yaml, keyed by provider name:

providers:
local:
api: http://localhost:8080/v1
# api_key omitted — Nastech uses "no-key-required" for keyless local servers
work:
api: https://gpu-server.internal.corp/v1
key_env: CORP_API_KEY
transport: chat_completions # set explicitly by `nastech model` → Custom Endpoint wizard; auto-detection still happens as a fallback
anthropic-proxy:
api: https://proxy.example.com/anthropic
key_env: ANTHROPIC_PROXY_KEY
transport: anthropic_messages # for Anthropic-compatible proxies

Each entry accepts: api (the endpoint base URL — base_url/url are accepted aliases), name (optional display name; defaults to the dict key), key_env or inline api_key or key_cmd (see below), transport (chat_completions / anthropic_messages / codex_responses), default_model, models, context_length, discover_models, extra_body, extra_headers, ssl_ca_cert / ssl_verify, catalog_provider (see below), and enabled: false to hide an entry without deleting it.

Command-minted credentials (key_cmd)​

Vision, thinking, and native local-model capability probes materialize the same callable credential used by chat before building authentication headers. They reuse the command token cache without replacing the chat client's callable. If a command cannot mint a string token, these best-effort probes send no bearer rather than an object representation or a lower-priority configured credential. Native local-model probes remove inherited Authorization on a failed explicit callable while retaining unrelated configured headers. Chat retains its normal error handling.

Enterprise gateways often issue short-lived bearer tokens (SSO/OIDC brokers, cloud IAM, internal auth proxies) rather than static API keys, so a token copied into .env goes stale mid-session and requests start returning 401. key_cmd names a command that prints a token; Nastech runs it and caches the result until shortly before expiry, so long sessions keep working with no restart:

providers:
my-gateway:
base_url: "https://gateway.internal.example.com/v1"
api_mode: chat_completions
key_cmd: "my-auth-cli print-token --profile prod"

Works with any helper that prints a token — databricks auth token, gcloud auth print-access-token, az account get-access-token, vault read, or Claude Code-style apiKeyHelper scripts.

The command must print only the token on stdout: either bare, or as JSON with an access_token field (expires_in is honored; absolute expiry/expiresOn ISO timestamps too). Multi-line output is rejected rather than guessed at. If no expiry is advertised, the token is re-minted on a bounded window.

Precedence: an explicit --api-key flag still wins; otherwise key_cmd beats a static api_key/key_env on the same entry. The minted credential applies to the main agent turn and to auxiliary tasks (title generation, compression, vision, embedding) alike.

Model discovery also honors key_cmd for both providers: and legacy custom_providers: entries, including nastech model setup. Helpers run only when an authenticated live catalog probe is needed: disabled discovery and warm catalog cache reads do not mint tokens. Catalogs are scoped to the command identity, so rotating a bearer does not invalidate the catalog. Probe helpers use their own short-lived token source, not the inference client's token cache; minted bearers are never saved to config.yaml. If a helper fails, discovery falls back to the configured model without exposing the helper's output.

Not to be confused with secrets.command, which runs a helper once at startup to populate env vars process-wide. Use that for a vault/keychain helper handing back many secrets; use key_cmd when one provider's credential must be re-minted during a session.

Legacy format

Older configs used a top-level custom_providers: list instead. It still works — Nastech reads both — and nastech update auto-migrates it to the providers: dict (config v12). Field names differ slightly in the dict format: legacy model is default_model, and legacy api_mode is transport.

Reasoning effort on custom endpoints. The configured reasoning_effort (/reasoning max, agent.reasoning_effort) reaches a custom endpoint unchanged on both the chat_completions and the codex_responses transport — up to max; only the Nastech-internal ultra is clamped to max. Two exceptions follow the host rather than the entry: a custom entry pointed at api.openai.com keeps OpenAI's per-model ladder (max is a gpt-5.6-only level there), and an entry pointed at a provider whose profile publishes a per-model vocabulary (Ramp Router) is clamped to that catalog. An endpoint that rejects the level answers with an HTTP 400 instead of Nastech silently downgrading it.

Some OpenAI-compatible endpoints need provider-specific request body fields. Add an extra_body map to the matching custom provider and Nastech will merge it into each chat-completions request for that endpoint:

providers:
gemma-local:
api: http://localhost:8080/v1
default_model: google/gemma-4-31b-it
extra_body:
enable_thinking: true
reasoning_effort: high

Use the shape your server documents. For example, vLLM Gemma deployments and some NVIDIA NIM endpoints expect enable_thinking under chat_template_kwargs instead of as a top-level extra_body field:

extra_body:
chat_template_kwargs:
enable_thinking: true

For Qwen reasoning models served by vLLM, this same shape can be used to disable thinking when a reasoning parser separates all generated text into reasoning fields and leaves the assistant content empty:

extra_body:
chat_template_kwargs:
enable_thinking: false

The configured extra_body follows the provider everywhere: it is merged at agent construction, survives every gateway turn (including turns where /fast layers service_tier/speed overrides on top — those merge over your extra_body rather than replacing it), and is re-derived on /model switches — switching to a named custom provider applies its extra_body, and switching away clears it so it never leaks to another provider.

The nastech model → Custom Endpoint wizard now prompts for the API mode explicitly and persists your answer to config.yaml (as transport on the provider entry). URL-based auto-detection (e.g. /anthropic paths → anthropic_messages) still happens as a fallback when the field is left blank.

Native vision for custom-provider models. If your custom endpoint serves a vision-capable model that isn't in models.dev, set model.supports_vision: true so Nastech routes attached images natively (as image_url parts) instead of pre-processing them through vision_analyze. Single knob — no need to also set agent.image_input_mode: native.

model:
provider: custom
base_url: http://localhost:8080/v1
default: qwen3.6-35b-a3b
supports_vision: true # send images natively; otherwise vision_analyze pre-describes them

The same key is honored on per-named-provider models (providers.<name>.models.<id>.supports_vision) and accepts standard YAML booleans (true/false/yes/no/on/off/1/0).

A model_overrides entry that only corrects metadata (for example context_window) for a model the catalog does not know leaves vision and reasoning capability and the output-token limit unknown — vision_analyze, video_analyze and the reasoning-effort picker stay available. Only an explicit supports_vision: false / supports_reasoning: false in the override marks the model as text-only or non-reasoning.

Inheriting a catalogued vendor's metadata (catalog_provider). When a named custom provider (a gateway, proxy or reseller) serves models that Nastech already knows under a built-in provider, point the entry at that vendor and its models inherit the catalogued context window, output limit, vision and reasoning flags — no model_overrides needed:

providers:
my-gateway:
api: https://gateway.example.com/v1
key_env: GATEWAY_API_KEY
catalog_provider: deepseek # metadata lookups use DeepSeek's catalog entries

catalog_provider accepts a Nastech provider id (deepseek, anthropic, openai, …) or a models.dev id. It affects metadata lookups only — requests still go to your api URL with your credentials — and an explicit model_overrides entry for the same model still wins.

Switch between them mid-session with the triple syntax:

/model custom:local:qwen-2.5 # Use the "local" endpoint with qwen-2.5
/model custom:work:llama3-70b # Use the "work" endpoint with llama3-70b
/model custom:anthropic-proxy:claude-sonnet-4 # Use the proxy

You can also select named custom providers from the interactive nastech model menu.


Cookbook: Together AI, Groq, Perplexity​

The cloud providers listed in Other Compatible Providers all speak OpenAI's REST dialect, so they wire up the same way under the providers: dict. Three worked recipes follow. Each drops into ~/.nastech/config.yaml and the matching API key goes in ~/.nastech/.env.

Together AI​

Hosts open-weight models (Llama, MiniMax, Gemma, DeepSeek, Qwen) at prices significantly below first-party APIs. Good default for multi-model fleets.

# ~/.nastech/config.yaml
providers:
together:
api: https://api.together.xyz/v1
key_env: TOGETHER_API_KEY
# transport: chat_completions # default — no need to set

model:
default: MiniMaxAI/MiniMax-M2.7 # or any model from together.ai/models
provider: custom:together
# ~/.nastech/.env
TOGETHER_API_KEY=your-together-key

Switch models mid-session:

/model custom:together:meta-llama/Llama-3.3-70B-Instruct-Turbo
/model custom:together:google/gemma-4-31b-it
/model custom:together:deepseek-ai/DeepSeek-V3

Together's /v1/models endpoint works, so nastech model can auto-discover available models.

Groq​

Ultra-fast inference (~500 tok/s on Llama-3.3-70B). Small catalog but strong for latency-sensitive interactive use.

# ~/.nastech/config.yaml
providers:
groq:
api: https://api.groq.com/openai/v1
key_env: GROQ_API_KEY

model:
default: llama-3.3-70b-versatile
provider: custom:groq
# ~/.nastech/.env
GROQ_API_KEY=your-groq-key

Perplexity​

Useful when you want a model that does live web search and citation automatically. Strict about which models are available — check perplexity.ai/settings/api for the current list.

# ~/.nastech/config.yaml
providers:
perplexity:
api: https://api.perplexity.ai
key_env: PERPLEXITY_API_KEY

model:
default: sonar
provider: custom:perplexity
# ~/.nastech/.env
PERPLEXITY_API_KEY=your-perplexity-key

Perplexity's Agent API (api: https://api.perplexity.ai/v1 with api_mode: codex_responses) reserves the function names web_search, search_files, fetch_url, people_search and finance_search for its own built-in tools. Nastech renames its client tools of the same name to nastech_<name> on the wire and maps them back before dispatch, for the main agent loop and auxiliary calls (title generation, compression, MoA aggregation) alike — the same treatment OpenCode's /v1/responses endpoints get.

Multiple providers in one config​

The three recipes compose — use all of them together and switch per turn with /model custom:<name>:<model>:

providers:
together:
api: https://api.together.xyz/v1
key_env: TOGETHER_API_KEY
groq:
api: https://api.groq.com/openai/v1
key_env: GROQ_API_KEY
perplexity:
api: https://api.perplexity.ai
key_env: PERPLEXITY_API_KEY

model:
default: MiniMaxAI/MiniMax-M2.7
provider: custom:together # boot to Together; switch freely after
Troubleshooting
  • nastech doctor should print no Unknown provider warnings for any of these names after the CLI validator fixes in #15083.
  • If a provider's /v1/models endpoint is unreachable (Perplexity is the common one), nastech model will persist the model with a warning rather than hard-reject — see #15136.
  • To skip named providers entirely and use bare provider: custom with CUSTOM_BASE_URL env var, see #15103.

Choosing the Right Setup​

Use CaseRecommended
Just want it to workOpenRouter (default) or Nastech Portal
Local models, easy setupOllama
Production GPU servingvLLM or SGLang
Mac / no GPUOllama or llama.cpp
Multi-provider routingLiteLLM Proxy or OpenRouter
Cost optimizationClawRouter or OpenRouter with sort: "price"
Maximum privacyOllama, vLLM, or llama.cpp (fully local)
Enterprise / AzureAzure OpenAI with custom endpoint
Chinese AI modelsz.ai (GLM), Kimi/Moonshot (kimi-coding or kimi-coding-cn), MiniMax, Xiaomi MiMo, or Tencent TokenHub (first-class providers)
tip

You can switch between providers at any time with nastech model — no restart required. Your conversation history, memory, and skills carry over regardless of which provider you use.

Optional API Keys​

FeatureProviderEnv Variable
Web scrapingFirecrawlFIRECRAWL_API_KEY, FIRECRAWL_API_URL
Browser automationBrowserbaseBROWSERBASE_API_KEY, BROWSERBASE_PROJECT_ID
Image generationFALFAL_KEY
Premium TTS voicesElevenLabsELEVENLABS_API_KEY
OpenAI TTS + voice transcriptionOpenAIVOICE_TOOLS_OPENAI_KEY
Mistral TTS + voice transcriptionMistralMISTRAL_API_KEY
Cross-session user modelingHonchoHONCHO_API_KEY
Semantic long-term memorySupermemorySUPERMEMORY_API_KEY

Self-Hosting Firecrawl​

By default, Nastech uses the Firecrawl cloud API for web search and scraping. If you prefer to run Firecrawl locally, you can point Nastech at a self-hosted instance instead. See Firecrawl's SELF_HOST.md for complete setup instructions.

What you get: No API key required, no rate limits, no per-page costs, full data sovereignty.

What you lose: The cloud version uses Firecrawl's proprietary "Fire-engine" for advanced anti-bot bypassing (Cloudflare, CAPTCHAs, IP rotation). Self-hosted uses basic fetch + Playwright, so some protected sites may fail. Search uses DuckDuckGo instead of Google.

Setup:

  1. Clone and start the Firecrawl Docker stack (5 containers: API, Playwright, Redis, RabbitMQ, PostgreSQL — requires ~4-8 GB RAM):

    git clone https://github.com/firecrawl/firecrawl
    cd firecrawl
    # In .env, set: USE_DB_AUTHENTICATION=false, HOST=0.0.0.0, PORT=3002
    docker compose up -d
  2. Point Nastech at your instance (no API key needed):

    nastech config set FIRECRAWL_API_URL http://localhost:3002

You can also set both FIRECRAWL_API_KEY and FIRECRAWL_API_URL if your self-hosted instance has authentication enabled.

OpenRouter Provider Routing​

When using OpenRouter, you can control how requests are routed across providers. Add a provider_routing section to ~/.nastech/config.yaml:

provider_routing:
sort: "throughput" # "price" (default), "throughput", or "latency"
# only: ["anthropic"] # Only use these providers
# ignore: ["deepinfra"] # Skip these providers
# order: ["anthropic", "google"] # Try providers in this order
# require_parameters: true # Only use providers that support all request params
# data_collection: "deny" # Exclude providers that may store/train on data
# models: # Per-model pins (same keys; unset keys fall through)
# "openai/gpt-6-astra": {only: ["openai"]}
# "anthropic/claude-fable-5.1": {only: ["anthropic"]}

Shortcuts: Append :nitro to any model name for throughput sorting (e.g., anthropic/claude-sonnet-4:nitro), or :floor for price sorting. Per-model details: Provider Routing.

OpenRouter Pareto Code Router​

OpenRouter ships an experimental coding-model router at openrouter/pareto-code that auto-routes requests to the cheapest model meeting a coding-quality bar (ranked by Artificial Analysis). Pick this model and tune the min_coding_score knob in ~/.nastech/config.yaml:

model:
provider: openrouter
model: openrouter/pareto-code

openrouter:
min_coding_score: 0.65 # 0.0–1.0; higher = stronger (more expensive) coders. Default 0.65.

Notes:

  • min_coding_score is only sent when model.model is openrouter/pareto-code. On any other model the value is a no-op.
  • Set to empty string (or remove the line) to let OpenRouter pick the strongest available coder — its documented behavior when the plugins block is omitted.
  • Selection is deterministic per score on a given day, but the actual model chosen can shift as the Pareto frontier moves (new models, benchmark updates).
  • See OpenRouter's Pareto Router docs for the full router behavior.
  • To use the Pareto Code router for a specific auxiliary task (compression, vision, etc.) instead of the main agent, set extra_body.plugins under that task — see Auxiliary Models → OpenRouter routing & Pareto Code for auxiliary tasks.

Fallback Providers​

Configure a chain of backup providers Nastech tries in order when the primary model fails (rate limits, server errors, auth failures). The canonical format is a top-level fallback_providers: list:

fallback_providers:
- provider: openrouter
model: anthropic/claude-sonnet-4
- provider: anthropic
model: claude-sonnet-4
# base_url: http://localhost:8000/v1 # optional, for custom endpoints
# api_mode: chat_completions # optional override

The legacy single-pair fallback_model: dict is still accepted for back-compat:

fallback_model:
provider: openrouter
model: anthropic/claude-sonnet-4

When activated, the fallback swaps the model and provider mid-session without losing your conversation. The chain is tried entry-by-entry; activation is one-shot per session.

Supported providers: openrouter, nastech, novita, openai-codex, copilot, copilot-acp, anthropic, gemini, qwen-oauth, huggingface, zai, kimi-coding, kimi-coding-cn, minimax, minimax-cn, minimax-oauth, deepseek, nvidia, xai, xai-oauth, ollama-cloud, bedrock, ai-gateway, azure-foundry, opencode-zen, opencode-go, commandcode, commandcode-anthropic, kilocode, xiaomi, arcee, gmi, actual, stepfun, lmstudio, alibaba, alibaba-coding-plan, tencent-tokenhub, tencent-tokenplan, nebius-token-factory, router, custom.

tip

Fallback is configured exclusively through config.yaml — or interactively via nastech fallback. For full details on when it triggers, how the chain advances, and how it interacts with auxiliary tasks and delegation, see Fallback Providers.


See Also​

  • Configuration — General configuration (directory structure, config precedence, terminal backends, memory, compression, and more)
  • Environment Variables — Complete reference of all environment variables