🔌 API Endpoints
Authentication: Pass your API key via
Authorization: Bearer YOUR_API_KEY or
X-Api-Key: YOUR_API_KEY header.
🍎 Xcode Setup (Apple Intelligence / Custom LLM)
Xcode 26+ lets you point Apple Intelligence at a custom OpenAI-compatible endpoint. Configure it as follows:
- Open Xcode → Settings → Generative AI (or Features → Generative AI depending on your version).
-
Set Base URL to:
https://ai.techxartisan.com/api
⚠️ Do NOT include/v1— Xcode appends it automatically. Using/api/v1here will result in/api/v1/v1/models(404). -
Paste your API Key into the API Key field.
(Get yours from the API Keys page.) -
Pick a Model from the dropdown.
The list is fetched fromGET /api/v1/models. - Click Test Connection — if everything is correct, you'll get a green checkmark and the model list will populate.
Common pitfall: If Xcode reports "Cannot fetch models" or "Provider is not valid":
- Base URL is
https://ai.techxartisan.com/api— nothttps://ai.techxartisan.com/api/v1 - API key is in the right field (no extra whitespace)
- Your account is active (check the Dashboard)
💻 Quick Start — curl Examples
OpenAI format — works with any OpenAI SDK, LangChain, Cursor, etc.
# Non-streaming curl https://ai.techxartisan.com/api/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "qwen3.7-plus", "messages": [{"role": "user", "content": "Hello!"}], "stream": false }' # Streaming curl https://ai.techxartisan.com/api/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "qwen3.7-plus", "messages": [{"role": "user", "content": "Hello!"}], "stream": true }'
Anthropic format — works with Claude CLI, Claude Code, Anthropic SDK.
curl https://ai.techxartisan.com/api/v1/messages/ \ -H "x-api-key: YOUR_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-4-20250514", "max_tokens": 1024, "messages": [{"role": "user", "content": "Hello!"}] }'
🤖 Available Models
Pass any of these names as the model field in your request.
The router will pick the best upstream supplier automatically.
PaddlePaddle/ERNIE-4.5-300B-A47B-PT
—
PaddlePaddle/ERNIE-4.5-VL-28B-A3B-PT
—
Qwen/Qwen3-4B
—
Qwen/Qwen3-Embedding-0.6B
—
Qwen/Qwen3-Embedding-4B
—
Qwen/Qwen3-Embedding-8B
—
Qwen3.8-Flash
—
Qwen3.8-Max
—
deepseek-ai/DeepSeek-V4-Flash-0731
—
deepseek-ai/DeepSeek-V4-Pro-0813
—
deepseek-v4-flash-vision-exp
—
👁
glm-5.3-flash
—
👁
💭
qwen3-coder-plus
—
qwen3-max-2026-01-23
—
test/model-xyz
—
GLM-5.2
flagship
MedAIBase/AntAngelMed
flagship
MiniMax/MiniMax-M1-80k
flagship
MusePublic/Qwen-Image-Edit
flagship
OpenGVLab/InternVL3_5-241B-A28B
flagship
Qwen/Qwen-Image-Edit
flagship
Qwen/Qwen3-235B-A22B
flagship
Qwen/Qwen3-235B-A22B-Instruct-2507
flagship
Qwen/Qwen3-235B-A22B-Thinking-2507
flagship
ZhipuAI/GLM-5.2
flagship
deepseek-ai/DeepSeek-V4-Pro
flagship
💭
glm-5
flagship
glm-5.1
flagship
mistralai/Mistral-Large-Instruct-2407
flagship
nvidia/nemotron-3-super-120b-a12b
flagship
nvidia/nemotron-3-ultra-550b-a55b
flagship
opencompass/CompassJudger-1-32B-Instruct
flagship
qwen3.7-plus
flagship
👁
💭
deepseek-v4-flash
standard
💭
glm-4.7
standard
google/diffusiongemma-26b-a4b-it
standard
kimi-k2.5
standard
👁
meituan-longcat/LongCat-Flash-Lite
standard
meta/llama-3.2-11b-vision-instruct
standard
👁
meta/muse-glimmer-30b
standard
nvidia/ising-calibration-1.5-31b
standard
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning
standard
nvidia/nemotron-3.5-content-safety
standard
nvidia/nemotron-3.5-lightning-30b-a3b
standard
nvidia/nemotron-parse-2.0
standard
nvidia/riva-translate-4b-instruct-v1.1
standard
nvidia/riva-translate-4b-instruct-v2
standard
openai/gpt-oss-20b
standard
poolside/laguna-xs-2.1
standard
qwen3.5-plus
standard
👁
qwen3.6-plus
standard
✨ Features
- Smart routing — requests are automatically load-balanced across multiple suppliers with failover.
- Cross-model fallback — if the requested model is unavailable, the router falls back to a model in the same tier (flagship / standard / free).
- Format conversion — send requests in either OpenAI or Anthropic format; the router translates as needed.
- Token compression — RTK (input) and Caveman (output) compression can be enabled per-client in the admin panel.
- Streaming — SSE streaming is supported on both OpenAI and Anthropic endpoints.
- Thinking / extended thinking — supported where the upstream model allows it (look for 💭 in the model list).