🔌 API Endpoints

Authentication: Pass your API key via Authorization: Bearer YOUR_API_KEY or X-Api-Key: YOUR_API_KEY header.

🍎 Xcode Setup (Apple Intelligence / Custom LLM)

Xcode 26+ lets you point Apple Intelligence at a custom OpenAI-compatible endpoint. Configure it as follows:

  1. Open Xcode → Settings → Generative AI (or Features → Generative AI depending on your version).
  2. Set Base URL to:
    https://ai.techxartisan.com/api
    ⚠️ Do NOT include /v1 — Xcode appends it automatically. Using /api/v1 here will result in /api/v1/v1/models (404).
  3. Paste your API Key into the API Key field.
    (Get yours from the API Keys page.)
  4. Pick a Model from the dropdown.
    The list is fetched from GET /api/v1/models.
  5. Click Test Connection — if everything is correct, you'll get a green checkmark and the model list will populate.
Common pitfall: If Xcode reports "Cannot fetch models" or "Provider is not valid":
  • Base URL is https://ai.techxartisan.com/api — not https://ai.techxartisan.com/api/v1
  • API key is in the right field (no extra whitespace)
  • Your account is active (check the Dashboard)

💻 Quick Start — curl Examples

OpenAI format — works with any OpenAI SDK, LangChain, Cursor, etc.

# Non-streaming
curl https://ai.techxartisan.com/api/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.7-plus",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": false
  }'

# Streaming
curl https://ai.techxartisan.com/api/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.7-plus",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'

Anthropic format — works with Claude CLI, Claude Code, Anthropic SDK.

curl https://ai.techxartisan.com/api/v1/messages/ \
  -H "x-api-key: YOUR_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-20250514",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

🤖 Available Models

Pass any of these names as the model field in your request. The router will pick the best upstream supplier automatically.

PaddlePaddle/ERNIE-4.5-300B-A47B-PT —
PaddlePaddle/ERNIE-4.5-VL-28B-A3B-PT —
Qwen/Qwen3-4B —
Qwen/Qwen3-Embedding-0.6B —
Qwen/Qwen3-Embedding-4B —
Qwen/Qwen3-Embedding-8B —
Qwen3.8-Flash —
Qwen3.8-Max —
deepseek-ai/DeepSeek-V4-Flash-0731 —
deepseek-ai/DeepSeek-V4-Pro-0813 —
deepseek-v4-flash-vision-exp — 👁
glm-5.3-flash — 👁 💭
qwen3-coder-plus —
qwen3-max-2026-01-23 —
test/model-xyz —
GLM-5.2 flagship
MedAIBase/AntAngelMed flagship
MiniMax/MiniMax-M1-80k flagship
MusePublic/Qwen-Image-Edit flagship
OpenGVLab/InternVL3_5-241B-A28B flagship
Qwen/Qwen-Image-Edit flagship
Qwen/Qwen3-235B-A22B flagship
Qwen/Qwen3-235B-A22B-Instruct-2507 flagship
Qwen/Qwen3-235B-A22B-Thinking-2507 flagship
ZhipuAI/GLM-5.2 flagship
deepseek-ai/DeepSeek-V4-Pro flagship 💭
glm-5 flagship
glm-5.1 flagship
mistralai/Mistral-Large-Instruct-2407 flagship
nvidia/nemotron-3-super-120b-a12b flagship
nvidia/nemotron-3-ultra-550b-a55b flagship
opencompass/CompassJudger-1-32B-Instruct flagship
qwen3.7-plus flagship 👁 💭
deepseek-v4-flash standard 💭
glm-4.7 standard
google/diffusiongemma-26b-a4b-it standard
kimi-k2.5 standard 👁
meituan-longcat/LongCat-Flash-Lite standard
meta/llama-3.2-11b-vision-instruct standard 👁
meta/muse-glimmer-30b standard
nvidia/ising-calibration-1.5-31b standard
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning standard
nvidia/nemotron-3.5-content-safety standard
nvidia/nemotron-3.5-lightning-30b-a3b standard
nvidia/nemotron-parse-2.0 standard
nvidia/riva-translate-4b-instruct-v1.1 standard
nvidia/riva-translate-4b-instruct-v2 standard
openai/gpt-oss-20b standard
poolside/laguna-xs-2.1 standard
qwen3.5-plus standard 👁
qwen3.6-plus standard

✨ Features

  • Smart routing — requests are automatically load-balanced across multiple suppliers with failover.
  • Cross-model fallback — if the requested model is unavailable, the router falls back to a model in the same tier (flagship / standard / free).
  • Format conversion — send requests in either OpenAI or Anthropic format; the router translates as needed.
  • Token compression — RTK (input) and Caveman (output) compression can be enabled per-client in the admin panel.
  • Streaming — SSE streaming is supported on both OpenAI and Anthropic endpoints.
  • Thinking / extended thinking — supported where the upstream model allows it (look for 💭 in the model list).